BEGIN:VCALENDAR
VERSION:2.0
PRODID:icalendar-ruby
CALSCALE:GREGORIAN
X-WR-CALNAME:ESE Seminar: Hongyi Wang
X-WR-TIMEZONE:Central Time (US & Canada)
BEGIN:VEVENT
DTSTAMP:20260814T090710Z
UID:tag:localist.com\,2008:EventInstance_45940602383210
DTSTART:20240326T160000Z
DTEND:20240326T170000Z
DESCRIPTION:A Step Further Toward Scalable and Automatic Distributed Large 
 Language Model Pre-training\n\nAbstract: Large Language Models (LLMs) are 
 at the forefront of advances in the field of AI. Nonetheless\, training th
 ese LLMs is computationally daunting\, which necessitates distributed trai
 ning. Distributed training generally suffers from bottlenecks\, including 
 heavy communication costs and the need for extensive performance tuning. D
 istributed training with hybrid parallelism\, which combines data and mode
 l parallelism\, is essential for LLM pre-training. Designing effective hyb
 rid parallelism strategies requires a substantial tuning effort and specia
 lized expertise. In this talk\, I will first discuss how to automatically 
 and efficiently design high-throughput hybrid-parallelism strategies using
  system cost models. Then\, I will demonstrate the use of these automatica
 lly designed hybrid parallelism strategies to train state-of-the-art LLMs 
 from scratch. Finally\, I will introduce a low-rank training framework to 
 enhance communication efficiency in data parallelism. This proposed framew
 ork achieves near-ideal scalability without sacrificing model quality by l
 everaging a full-rank to low-rank training strategy and a layer-wise adapt
 ive rank selection mechanism.
GEO:38.648794;-90.30145
LOCATION:Preston M. Green Hall\, Rodin Auditorium\, L0120
SUMMARY:ESE Seminar: Hongyi Wang
URL;VALUE=URI:https://happenings.washu.edu/event/ese-seminar-hongyi-wang
CATEGORIES:Seminar/Colloquia
END:VEVENT
END:VCALENDAR
