Pretraining LLM: data, compute, scaling laws และต้นทุน
ฝึก 20 โมเดลบน GPU แล็ปท็อปเพื่อวัด scaling law และตรวจสอบประมาณการ compute 6ND กับตัวนับ FLOP จริง
ในหน้านี้
บทที่ 9 จบลงด้วย transformer block ที่ train ได้ เอาหลาย block มาซ้อนกัน ชี้ next-token loss จาก บทที่ 8 ไปที่ output แล้วก็ไม่มีอะไรเหลือให้คิดค้นอีก สิ่งที่เหลือทั้งหมดคือการซื้อ
นี่เป็นการเปลี่ยนแปลงที่ใหญ่กว่าที่ฟังดูเป็นมาก ทุกบทก่อนหน้านี้ถามว่า มันเรียนรู้ไหม? — คำถามแบบใช่หรือไม่ใช่ที่แล็ปท็อปตอบได้ในสิบ分钟 บทนี้ถามคำถามที่มีเงินอยู่ในนั้น: เมื่อมีปริมาณเลขคณิตคงที่ โมเดลที่ดีที่สุดที่ฉันซื้อได้คืออะไร? คำตอบคือสูตรหนึ่ง และในปี 2018 ไม่มีใครเห็นว่ามันชัดเจนเลย
นี่คือคำถามนั้นที่ตอบด้วยการวัด บน GPU แล็ปท็อปเครื่องเดียว โมเดล 20 ตัว ตั้งแต่ 98,624 ถึง 15 ล้าน parameters ถูก train จากศูนย์บน Wikipedia 174 ล้าน tokens — vocabulary แบบ BPE ขนาด 2,048-token ที่ train แบบเดียวกับ บทที่ 7, transformer ของบทที่ 9 แต่ละ run ได้ compute budget หนึ่งในสามชุดแบบ เป๊ะ ๆ และไม่มากกว่านั้นแม้แต่ operation เดียว ดังนั้นโมเดลที่ใหญ่กว่าจำเป็นต้องอ่าน text น้อยกว่า loss บน held-out ที่ดีที่สุดในแต่ละ budget คือ:
budget C (FLOPs) best loss reached by a model of
1.00e13 5.3531 98,624 params
3.16e13 4.8638 98,624 params
1.00e14 4.3383 295,808 params
fitted: L = (Cc / C)^0.0913 over one decade of computeเลขคณิตมากขึ้นสิบเท่าลด loss ลง 19% และสามจุดอยู่บนเส้นตรงใน log-log ไม่มีอะไรในเก้าบทแรกทำนายสิ่งนี้ได้ ไม่มี theorem รองรับ — มันคือความสม่ำเสมอเชิงประจักษ์ ซึ่งคงอยู่ด้วย exponent ต่างออกไปตลอดสิบ order of magnitude ระหว่างแล็ปท็อปเครื่องนี้กับ datacentre และเป็นข้อสังเกตเดียวที่โน้มน้าวทั้งอุตสาหกรรมให้ใช้เงินเท่า GDP ของประเทศเล็ก ๆ กับ GPUs
Pretraining คืออะไร และมีอะไรใหม่เกี่ยวกับมัน
ลิงก์ไปยังส่วน: Pretraining คืออะไร และมีอะไรใหม่เกี่ยวกับมันไม่มีอะไรใน objective เปลี่ยนไป โมเดลยังทำนาย token ถัดไป loss ยังเป็น cross-entropy จาก บทที่ 4 ที่นำไปใช้กับ factorisation ของบทที่ 8 optimiser ยังเป็น AdamW จาก บทที่ 6 Pretraining ไม่ใช่ algorithm ใหม่; มันคือ algorithm เดิมที่ run บน corpus ใหญ่พอจนต้องทำ budget ให้ run นั้น สองสิ่งทำให้เป็นไปได้: labels ฟรี เพราะ target สำหรับตำแหน่ง คือ token ที่ และมีอยู่ใน text อยู่แล้ว; และส่วนสุดท้ายของบทที่ 6 เอาข้อคัดค้านออกไป เพราะโมเดลที่มี parameters มากกว่ากฎคลาสสิกอนุญาตอย่างมากไม่ได้พัง แต่มันดีขึ้น สิ่งที่ได้ออกมาคือ base model — บางสิ่งที่ต่อ text ไม่ใช่ตอบคำถาม
นับ compute ก่อนใช้เงิน: 6ND
ลิงก์ไปยังส่วน: นับ compute ก่อนใช้เงิน: 6NDก่อนจะทำ budget ให้สิ่งนี้ได้ ต้องนับมันก่อน และวงการนับด้วยสูตรเดียว:
โดยที่ คือจำนวน parameters, คือ training tokens และ คือ floating-point operations ทั้งหมด Kaplan et al. อนุมานมันเป็นสองขั้นตอน1 Forward: 2 FLOPs ต่อ parameter ต่อ token เพราะทุก parameter ใน matrix multiply ถูกใช้หนึ่งครั้งต่อ token ในการคูณหนึ่งครั้งและบวกหนึ่งครั้ง Backward: สองเท่าของ forward เพราะ backward pass จาก บทที่ 5 คำนวณ gradients สองชุดในแต่ละ layer — เทียบกับ inputs ของ layer เพื่อให้ signal เดินทางต่อไป และเทียบกับ weights ของมัน — แต่ละชุดเป็น matrix multiply ขนาดเท่า forward หนึ่งครั้ง ดังนั้น
นั่นคือการอนุมานทั้งหมด และควรตรวจสอบมากกว่าเชื่อ PyTorch มีตัวนับ FLOP จริงคือ torch.utils.flop_counter.FlopCounterMode ซึ่ง intercept ทุก operation ที่โมเดล dispatch แล้วรวมงานจริงทั้งหมด ลอง run มันข้ามสี่ order of magnitude โดยตัวใหญ่สุดบน device meta ซึ่ง allocate shapes แต่ไม่ใช้ memory:
from torch.utils.flop_counter import FlopCounterMode
counter = FlopCounterMode(display=False)
with counter:
loss = model(x, targets)[1]
loss.backward()
measured = counter.get_total_flops()
print(measured / (6 * n_params * n_tokens))| configuration | ไม่รวม embeddings | ทั้งหมด | วัดได้, fwd+bwd | ÷ ( ทั้งหมด) | ÷ (ไม่รวม emb.) | fwd+bwd ÷ fwd |
|---|---|---|---|---|---|---|
| 128, 4 layers, 256 | 788,736 | 7,254,400 | 4.60e10 | 1.031 | 9.485 | 3.000 |
| 512, 8 layers, 256 | 25,183,232 | 51,045,888 | 3.26e11 | 1.038 | 2.104 | 3.000 |
| 768, 12 layers, 1024 | 84,973,056 | 124,356,864 | 1.75e12 | 1.145 | 1.676 | 3.000 |
| 1600, 48 layers, 1024 | 1,474,870,400 | 1,556,920,000 | 2.10e13 | 1.100 | 1.161 | 3.000 |
| 4096, 32 layers, 2048 | 6,442,983,424 | 6,582,444,032 | 8.74e13 | 1.080 | 1.104 | 3.000 |
| 8192, 80 layers, 8192 | 64,427,147,264 | 65,544,929,280 | 3.75e15 | 1.163 | 1.183 | 3.000 |
อัตราส่วน forward+backward ต่อ forward คือ 3.000 แบบเป๊ะ ๆ ในทุก scale: ไม่ใช่ approximation ที่บังเอิญดี แต่คือเอกลักษณ์ทางเลขคณิตข้างต้นที่ถูกตัวนับซึ่งไม่รู้อะไรเกี่ยวกับการอนุมานนั้นคืนกลับมาเป็นเลขกลม
จากนั้นค่ารวมที่วัดได้อยู่สูงกว่า ระหว่าง 3% ถึง 17% เมื่อ นับ embedding matrices — และ clause นั้นสำคัญ เพราะ paper ผู้ก่อตั้งสองฉบับนับ ต่างกัน Kaplan ไม่รวม “all vocabulary and positional embeddings” เพราะการทำเช่นนั้น “produces significantly cleaner scaling laws” (§1.3); Appendix F ของ Chinchilla บอกว่า “we also count embeddings matrices in the total parameter count”2 สำหรับ vocabulary กว้างและ hidden dimension แคบ ทั้งสองต่างกันเป็น factor เก้าอย่างที่แถวแรกแสดง
ช่องว่างที่เหลือคือสิ่งที่ จงใจละไว้: attention scores Eq. (2.2) ของ Kaplan เขียนต้นทุน forward เป็น และตัด term ที่สองทิ้งเพราะ — ปลอดภัยในปี 2020 ปลอดภัยน้อยลงตอนนี้ และเป็นเหตุผลที่อัตราส่วน drift สูงขึ้นเมื่อ โตขึ้น — ซึ่งเป็นเหตุผลที่สองแถวตรงนี้มี เท่ากันที่ 1,024 และอัตราส่วน ลดลง จาก 1.145 เป็น 1.100 เมื่อ เพิ่มจาก 768 เป็น 1,600 มันคือ cost ที่บทที่ 9 แนะนำ และ บทที่ 16 จะเปลี่ยนให้เป็นราคา
Memory: สิ่งที่ต้อง fit จริง ๆ
ลิงก์ไปยังส่วน: Memory: สิ่งที่ต้อง fit จริง ๆCompute ตัดสินว่า run จะใช้เวลานานแค่ไหน; memory ตัดสินว่ามันเริ่มได้ไหม Train ด้วย fp32 AdamW ธรรมดา แล้วทุก parameter แบกตัวเลขสี่ตัว: weight, gradient ของมัน, และ running mean กับ variance ของ Adam — ค่าเฉลี่ยสองตัวที่สร้างด้วยมือในบทที่ 6 ตัวเลขสี่ตัวที่ตัวละสี่ bytes คือ 16 bytes ต่อ parameter ก่อนมี activation แม้แต่ตัวเดียว วัดบน GPU แล็ปท็อป 8 GB โดยดู resident allocation ณ จุดใน step ที่ไม่มี graph มีชีวิตอยู่:
| model | vocabulary | batch | คาดการณ์ | resident ที่วัดได้ | peak ในหนึ่ง step | ส่วนต่าง | |
|---|---|---|---|---|---|---|---|
| 512, 8 layers | 50,257 | 8 | 51,045,888 | 779 MB | 801 MB | 2,500 MB | 1,699 MB |
| 512, 8 layers | 4,096 | 8 | 27,411,456 | 418 MB | 426 MB | 1,043 MB | 617 MB |
| 256, 6 layers | 4,096 | 8 | 5,839,360 | 89 MB | 89 MB | 382 MB | 293 MB |
| 256, 6 layers | 4,096 | 32 | 5,839,360 | 89 MB | 89 MB | 1,259 MB | 1,170 MB |
| 256, 6 layers | 4,096 | 128 | 5,839,360 | 89 MB | 89 MB | 4,771 MB | 4,681 MB |
การคาดการณ์และการวัดตรงกันภายใน 3% สิ่งที่น่าประหลาดใจคือคอลัมน์สุดท้าย: activations ใหญ่กว่าโมเดลอย่างท่วมท้น โมเดล 5.8-million-parameter ตัวเดิมที่ต้องใช้ persistent state 89 MB ต้องใช้ activations 4,681 MB ที่ batch 128 — ห้าสิบสองเท่าของโมเดล — และส่วนมากของสิ่งนั้นไม่ใช่ transformer เลย มันคือ logits, vector ขนาด vocabulary หนึ่งตัวต่อ token ที่ entry ละสี่ bytes: 512 MB ในแถวสุดท้าย, 393 MB ในแถวแรก ขนาด vocabulary ถูกเลือกในบทที่ 7 และมันยังคงตัดสินว่าอะไร fit บนการ์ด
term ไหนครองขึ้นกับรูปร่างของ run ซึ่งเป็นเหตุผลที่ Micikevicius et al. บอกว่า memory “is dominated by activations”3 ขณะที่ ZeRO บอกว่าโมเดล 1.5-billion-parameter ต้องการ model states อย่างเดียว “at least 24 GB”4 ZeRO มาถึง 16 bytes เดิมผ่านอีกเส้นทาง — สำหรับ fp16 weights, สำหรับ fp16 gradients, อย่างละตัวสำหรับ fp32 master weights และ moments สองตัวของ Adam — ซึ่งสำหรับ 70 พันล้าน parameters คือ 1.12 terabytes เท่ากับ GPUs 80 GB สิบสี่ตัวก่อนมี activation แม้แต่ตัวเดียว
Parallelism ในหนึ่งย่อหน้าและหนึ่งการมอบหมายต่อ
ลิงก์ไปยังส่วน: Parallelism ในหนึ่งย่อหน้าและหนึ่งการมอบหมายต่อที่ frontier scale ไม่มีสิ่งใด fit บนอุปกรณ์เดียว ดังนั้น run จึงถูกแบ่งพร้อมกันสี่ทาง Data parallelism วางสำเนาโมเดลไว้บนทุก GPU แล้วเฉลี่ย gradients — ค่า default และเป็นแบบที่ ZeRO ปรับปรุงด้วยการไม่เก็บสำเนาซ้ำซ้อนของ optimiser state Tensor parallelism แยก matrices แต่ละตัวข้าม devices Pipeline parallelism ให้แต่ละ device รับกลุ่ม layers ที่ต่อเนื่องกัน Context parallelism แยก sequence เอง ซึ่งจำเป็นก็ต่อเมื่อ ยาวพอให้ attention term ครอง Table 4 ของ Llama 3 แสดงทั้งสี่พร้อมกัน: tensor 8, context สูงสุด 16, pipeline 16, data สูงสุด 128 ข้าม 16,384 H100 GPUs5 คอร์สนี้จะพูดถึงมันเท่านี้; engineering ของ distributed training เป็นอีกหนึ่ง semester เต็ม ๆ และ Stanford CS336 คือ semester นั้น lectures 5 ถึง 8 พร้อม code6 สิ่งที่เหลือหลังการมอบหมายต่อคือเลขเดียว model FLOPs utilisation — สัดส่วนของ peak arithmetic ของ GPU ที่ run จริงทำได้ — ซึ่งเป็นสิ่งที่เปลี่ยน ที่เรียบร้อยให้เป็น wall-clock time และดังนั้นให้เป็นเงิน
Kaplan และการเดิมพันที่อุตสาหกรรมวาง
ลิงก์ไปยังส่วน: Kaplan และการเดิมพันที่อุตสาหกรรมวางในเดือนมกราคม 2020 Kaplan et al. train grid ของ transformers และพบว่า test loss ตาม power law ในทรัพยากรทั้งสามตลอดมากกว่าหก orders of magnitude1 §1.2 ของพวกเขาให้ fitted laws สามชุด:
พร้อม companion สำหรับ data และ สำหรับ optimally allocated compute ค่าคงที่ไม่ใช่ universal และ paper ก็บอกเช่นนั้น: “the precise numerical values of , and depend on the vocabulary size and tokenization and hence do not have a fundamental meaning.”
exponents เล็กมาก: parameters มากขึ้นสิบเท่าซื้อ factor ออกจาก loss ที่เหลือ ฟังเหมือนไม่มีอะไร และนี่คือข้อเท็จจริงสำคัญที่สุดตรงนี้ — ผลตอบแทนแย่มากและมันไม่เคยหยุด power law ที่มี exponent เล็กสัญญาว่า order of magnitude ถัดไปจะช่วย น้อยกว่าครั้งก่อน แต่ช่วยไปตลอด การซื้อ compute หยุดเป็นการพนันและกลายเป็นการซื้อที่มีอัตราแลกเปลี่ยนตีพิมพ์ไว้ ซึ่งตรงกับ argument ที่ปลดล็อกเงินทุน
จากนั้น prescription ก็ตามมา และนี่คือจุดที่ paper ผิดในแบบที่ทำให้อุตสาหกรรมเสียเงินจำนวนมาก Table 6 ของ Kaplan ให้ และ : compute มากขึ้นสิบเท่าหมายถึงโมเดลใหญ่ขึ้น 5.4 เท่า แต่ป้อน text มากขึ้นแค่ 1.9 เท่า abstract พูดชัด — “optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.” วงการทำตามนั้นจริง: GPT-3 มี 175 พันล้าน parameters บน 300 พันล้าน tokens,7 Gopher 280 พันล้านบน 300 พันล้าน, Megatron-Turing NLG 530 พันล้านบน 270 พันล้าน2 ครึ่ง token ถึงสอง tokens ต่อ parameter ทุกที่เหมือนกันหมด
Chinchilla และสิ่งที่ sweep ข้างบนกำลังวัด
ลิงก์ไปยังส่วน: Chinchilla และสิ่งที่ sweep ข้างบนกำลังวัดในเดือนมีนาคม 2022 Hoffmann et al. train โมเดลกว่า 400 ตัวตั้งแต่ 70 ล้านถึง 16 พันล้าน parameters และมาถึงข้อสรุปตรงกันข้ามผ่านสามเส้นทางอิสระ2 Table 2 ของพวกเขารายงาน exponent ใน เป็น 0.50, 0.49 และ 0.46 เทียบกับ 0.73 ของ Kaplan พูดง่าย ๆ: model size และ training data ควรโตในสัดส่วนเท่ากัน
วิธีที่สองของพวกเขาคือวิธีที่ sweep ตอนต้นบทนี้ทำซ้ำ ใน scale หนึ่งในล้าน: fix budget, train หลาย sizes ที่ budget นั้นแบบเป๊ะ ๆ, plot final loss เทียบกับ model size
| parameters | |||
|---|---|---|---|
| 98,624 | 5.3531 (171) | 4.8638 (542) | — |
| 150,320 | 5.4636 (74) | — | — |
| 194,208 | 5.5041 (44) | 4.8730 (140) | 4.4040 (442) |
| 295,808 | 5.5550 (19) | 4.9029 (60) | 4.3383 (190) |
| 665,280 | 5.7849 (3.8) | 5.1254 (12) | 4.4192 (38) |
| 1,280,768 | 5.8174 (1.0) | 5.1751 (3.2) | 4.5003 (10) |
| 3,101,568 | — | 5.4686 (0.5) | 4.7768 (1.7) |
| 5,315,072 | — | 5.5894 (0.2) | 4.8514 (0.6) |
| 15,053,568 | — | — | 5.3534 (0.1) |
Held-out loss เป็น nats ต่อ token, tokens ต่อ parameter ในวงเล็บ, ตัวหนาคือโมเดลที่ดีที่สุดในแต่ละ budget; dash คือจุดที่ไม่ได้ run เพราะ budget ต้องการ text มากกว่าที่ corpus มี หรือ size อยู่นอกช่วงที่ sweep ตรงนั้น
อ่านลงคอลัมน์: loss ลดลง แตะ bottom แล้วไต่ขึ้นอีก โมเดลใหญ่เกินไปสำหรับ budget ได้ง่ายพอ ๆ กับเล็กเกินไป — ที่ โทษของการเลือก 665,280 parameters แทน 295,808 คือ 0.08 nats ซึ่งบน envelope ที่ fit ข้างบนคือ loss ที่โมเดลขนาดถูกต้องจะไปถึงด้วย compute น้อยกว่า 18% การเลือก shape ผิดทิ้ง budget ไปหนึ่งในห้า นั่นคือ Figure 3 ของ Chinchilla ในบ่ายเดียวบน GPU หนึ่งตัว แทนที่จะใช้โมเดลสี่ร้อยตัว
ตอนนี้อ่านข้ามแถว ที่ โมเดลที่ดีที่สุดคือตัวเล็กสุดที่ sweep; ที่ คือ 295,808 parameters ซึ่งมีจุดประกบทั้งสองด้าน optimum เลื่อนไปทางขวาเมื่อ budget โตขึ้น ซึ่งคือเนื้อหาทั้งหมดของการแก้ไข Fit วิธีที่สามของ paper — surface เหนือทุก run — แล้ว minimise ภายใต้ :
L(N, D) = 24.7 / N^0.195 + 46.8 / D^0.169 (E fits to ~0; see below)
implied N_opt ∝ C^0.464
compare Chinchilla 0.46-0.50 · Besiroglu 0.513 · Kaplan 0.730.46 จากแล็ปท็อป เทียบกับ 0.73 ของ Kaplan การตรงกันสามหลักจาก fit สาม budget คือโชค; การตรงกันหลักแรกไม่ใช่ exponent เคลื่อนที่ได้ — constant ไม่ได้ เพราะ token-to-parameter ratio ที่ optima เหล่านี้คือ 170 ถึง 540 ไม่ใช่ 20 เหตุผลสามข้อ ล้วนให้บทเรียน fit ไปที่ศูนย์เพราะที่ loss สูงกว่า 4 nats run ยังห่างจาก entropy floor ที่ครอง fit ของ Chinchilla มาก Batch size และ learning rate ถูก fix แทนที่จะ tune ต่อจุด ซึ่งทำให้ runs ที่ได้ steps น้อยที่สุดเสียเปรียบ — และนั่นคือโมเดลใหญ่: ที่ FLOPs โมเดล 1.28-million-parameter ได้ optimiser steps รวม 159 ครั้ง ต่ำกว่าหลายพันครั้งที่ term ของ Kaplan บอกว่าโมเดลใด ๆ ต้องการมาก scaling law ถูก fit ภายใน regime หนึ่ง และอันนี้อยู่ต่ำกว่า Chinchilla หก orders of magnitude
ดังนั้น abstract ของ paper: “current large language models are significantly undertrained” Chinchilla คือการสาธิต — 70 พันล้าน parameters บน 1.4 ล้านล้าน tokens, compute รวมเท่าเดิม กับ Gopher 280 พันล้านบน 300 พันล้าน เอาชนะมันใน 51 จาก 57 tasks ของ MMLU, 67.5% ต่อ 60%2 เล็กลงสี่เท่า text มากขึ้นสี่เท่าครึ่ง เงินเท่าเดิม โมเดลดีกว่า
ข้อควรระวังสองข้อเกี่ยวกับ ratio ดังกล่าว “ยี่สิบ tokens ต่อ parameter” ไม่ใช่ประโยคใน paper ซึ่งบอกเพียงว่า “for every doubling of model size the number of training tokens should also be doubled”; เลข 20 เป็น inference จาก Table 3 และจาก Chinchilla เองที่ 70 B บน 1.4 T และความแม่นยำของมันแย่กว่าที่ตีพิมพ์: Besiroglu et al. refit จาก digitisation ของ Figure 4 พบว่า parameters เดิม “fit the reconstructed data poorly” โดยมี intervals “implausibly tight given the number of data points” และวางช่วงที่ซื่อตรงไว้ที่ “between 4 and 40” tokens ต่อ parameter8
รายละเอียดหนึ่งของวิธีของ Chinchilla ชำระสัญญาที่ บทที่ 1 เคยให้ไว้เกี่ยวกับ learning-rate schedules cosine schedule ต้อง match กับ token budget โมเดลที่จะเห็น 10 ล้าน tokens ต้อง decay learning rate ไปศูนย์ที่ 10 ล้าน tokens; ให้ schedule ที่ออกแบบสำหรับ 100 ล้าน แล้วหยุดก่อนกำหนด คุณกำลังอ่าน loss กลางทางลงที่ rate สูงเกินไป Chinchilla train แต่ละโมเดลที่ cycle lengths สี่แบบเพื่อ control สิ่งนี้โดยเฉพาะ; sweep ข้างบนตั้ง schedule จาก budget ด้วยเหตุผลเดียวกัน
สิ่งที่ scaling laws ไม่ได้สัญญา
ลิงก์ไปยังส่วน: สิ่งที่ scaling laws ไม่ได้สัญญามันเป็นผลเชิงประจักษ์ที่มีประโยชน์ที่สุดในวงการ และถูกขายเกินจริงเป็นประจำ ข้อจำกัดสี่ข้อ
มัน predict loss ไม่ใช่ capability ฝั่งซ้ายคือ cross-entropy บน held-out text ไม่มีอะไรใน papers เหล่านี้อนุญาตให้ claim ว่าโมเดลจะเขียน SQL ถูกต้อง ปฏิเสธ request อันตราย หรือใช้ tool ได้ นี่คือบทเรียนของบทที่ 5 อีกครั้ง: prediction ของ loss ไม่ใช่ prediction ของ behaviour ที่คุณกำลังจ่ายเงินซื้อ
มันถูก fit ไม่ได้ถูก derive ไม่มีทฤษฎีใดผลิต ค่าคงที่เคลื่อนตาม tokenizer — ซึ่งเป็นเหตุผลที่ perplexity comparison ข้าม tokenizers สองตัวไม่มีความหมาย อย่างที่บทที่ 8 อธิบาย — และเคลื่อนตาม data mixture, architecture และ optimiser ทุก law ที่ตีพิมพ์คือ law ของ setup ที่ผลิตมัน ซึ่งเป็นเหตุผลที่ Meta refit ของตัวเองก่อน Llama 35
มัน assume token สดใหม่สำหรับทุก step ซึ่งแอบ assume corpus อนันต์ Muennighoff et al. วัดว่าจะเกิดอะไรเมื่อมันหมด: data ซ้ำถึงสี่ epochs แทบไม่มีค่าใช้จ่าย — โมเดล 8.7-billion-parameter บน 44 พันล้าน unique tokens ที่เห็นสี่ครั้งจบด้วย “only 0.5% higher validation loss” กว่าโมเดลเดียวกันบน unique tokens 178 พันล้าน — ขณะที่เลยประมาณสิบหก epochs ไปแล้ว compute เพิ่มไม่ซื้ออะไรเลย9
และไม่มีใคร train แบบ compute-optimal อีกแล้ว Chinchilla minimise cost ของ training; โมเดลที่ deploy แล้วจ่ายประมาณ FLOPs ต่อ generated token ไปตลอด LLaMA 1 พูดตรง ๆ: “given a target level of performance, the preferred model is not the fastest to train but the fastest at inference”10 Sardana et al. formalise มันด้วยการ minimise แทน และพบว่าใครก็ตามที่คาดหวัง requests หนึ่งพันล้านควร train “smaller and longer than Chinchilla-optimal”11 §9.1 ของ Llama 3 เห็นด้วย: โมเดลเล็กของมัน train “far beyond the point of compute optimal training, effectively trading training compute for inference efficiency”5 ratio ไม่ได้ล้าสมัย; มันตอบคำถามที่ไม่ใช่คำถามที่กำลังถูกถามอีกต่อไป
Emergent abilities และข้อถกเถียงว่ามันจริงไหม
ลิงก์ไปยังส่วน: Emergent abilities และข้อถกเถียงว่ามันจริงไหมLoss ลดลงอย่างราบรื่น Benchmark scores บางครั้งไม่เป็นแบบนั้น Wei et al. รวบรวมกรณีที่ task อยู่ระดับเดาสุ่มตลอดหลาย orders of magnitude ของ training compute แล้วกระโดด — arithmetic สามหลักปรากฏใน GPT-3 ที่ประมาณ FLOPs, MMLU สูงกว่า guessing ระหว่าง และ — และตั้งชื่อ pattern นี้ว่า: “an ability is emergent if it is not present in smaller models but is present in larger models”12 หากนั่นเป็น property จริง การ extrapolate จากการทดลองราคาถูกก็ไม่ปลอดภัย เพราะ capability ที่คุณซื้ออาจไม่มีอยู่ใน scale ใดที่คุณจ่ายเพื่อทดสอบไหว
Schaeffer, Miranda และ Koyejo โต้แย้งว่าส่วนใหญ่มันเป็น artefact ของการวัด และ mechanism คือเลขคณิต13 Per-token loss ลดลงอย่างราบรื่น ดังนั้นความน่าจะเป็นที่ token หนึ่งจะถูกต้อง ดีขึ้นทีละน้อย ให้คะแนนโมเดลด้วย exact string match บนคำตอบยาว -token แล้วคุณยกความน่าจะเป็นนั้นเป็นกำลัง — curve ราบรื่นที่ถูกยกกำลังมาก ๆ ดูเหมือนหน้าผา เปลี่ยนเป็น metric ที่นับ tokens แทนที่จะเรียกร้องให้ถูกทั้งหมด บน outputs เดิม แล้ว “the family's performance smoothly, continuously and predictably improves with increasing scale”
ตัวเลขจาก audit ของพวกเขาคือสิ่งที่ควรจำ — “of the 39 preferred metrics in BIG-Bench, at most 5 display emergence” โดยมี metrics ไม่ต่อเนื่องสองตัวคิดเป็นกว่า 92% ของกรณีที่ claim — และคำเตือนของพวกเขาก็ควรจำเช่นกัน: “nothing in this paper should be interpreted as claiming that large language models cannot display emergent abilities” การกระโดดใน chart คือหลักฐานเกี่ยวกับ metric จนกว่าจะพิสูจน์เป็นอย่างอื่น บทที่ 29 คือจุดที่เรื่องนี้กลายเป็นปัญหาของคุณ เพราะการเลือก metric แบบ hard-cutoff เป็นการตัดสินใจที่คุณจะทำโดยไม่รู้ตัว
Data มาจากไหน
ลิงก์ไปยังส่วน: Data มาจากไหนCorpus คือส่วนของ pretraining run ที่ไม่มีสมการกำกับ และเป็นที่ที่การตัดสินใจสำคัญที่สุดจำนวนมากอยู่ วัตถุดิบคือ web crawl: archive เดือนสิงหาคม 2026 ของ Common Crawl มี “2.14 billion web pages or 360 TiB of uncompressed content” หนึ่งเดือนของมัน ดาวน์โหลดฟรี14 แทบไม่มีส่วนไหนใช้ได้ในสภาพเดิม Paper T5 บอกว่า crawl “largely comprises gibberish or boiler-plate text like menus, error messages, or duplicate text” และ pipeline C4 ที่มันแนะนำคือรายการ heuristics แบบทื่อ ๆ — เก็บเฉพาะบรรทัดที่ลงท้ายด้วย terminal punctuation, ทิ้ง pages ที่มีน้อยกว่าสามประโยค, ทิ้ง page ใด ๆ ที่มีวงเล็บปีกกาหรือคำจากรายการสาธารณะของคำหยาบ — เปลี่ยน text รายเดือนยี่สิบ terabytes ให้เหลือประมาณ 750 GB15
ทื่อคือคำที่ถูกต้อง Dodge et al. audit ว่า filters เหล่านั้นลบอะไรออกไป และพบว่า obscenity blocklist ลบ 42% ของ documents ใน African-American English และ 32% ใน Hispanic-aligned English เทียบกับ 6.2% ของ White-aligned English เหลือ corpus ที่ 97.8% เป็นหมวดสุดท้าย16 กฎที่ไม่มีความเห็นเกี่ยวกับ dialect กลับมีความเห็น
จากนั้น deduplication ซึ่งไม่ใช่งาน housekeeping: Lee et al. พบประโยค 61 คำที่ซ้ำ 61,036 ครั้งใน C4 และแสดงว่า deduplicating ลดอัตราที่โมเดล “emit memorized text” ลงสิบเท่า จาก 1.9% ของ generated tokens เป็น 0.19%17 แต่ more is not better — ทีม FineWeb deduplicate globally ข้าม 96 crawls ได้ 4 ล้านล้าน tokens และไม่มี gain ที่วัดได้ จากนั้น deduplicate แต่ละ crawl แยกกัน ได้ 20 ล้านล้าน และ match กับ corpus ที่ดีที่สุดที่มีอยู่18
จากนั้น contamination Llama 3 วัดของตัวเองและเผยแพร่: 98% ของ AGIEval, 95% ของ BIG-Bench Hard และ 85% ของ HellaSwag overlap กับ training set ด้วย 8-grams และสำหรับ MMLU overlap สูงจน “it is impossible to get a good performance gain estimate”5 §4 ของ GPT-3 รายงาน bug ใน filtering ที่ปล่อย benchmarks ไว้ใน data โดยไม่มีทางย้อนกลับ: “because of cost considerations it was infeasible to retrain the model”7
Provenance คือส่วนที่ยังไม่คลี่คลาย The Pile ส่ง component 100.96 GiB ชื่อ Books3 — 12% ของ corpus และตาม consent table ของ paper เอง คือหนังสือจาก private torrent tracker;19 มันถูกนำ offline ในเดือนสิงหาคม 2023 หลังมี copyright complaint สถานะทางกฎหมาย ณ กันยายน 2026 ยังไม่ settled และคำวินิจฉัยสหรัฐฯ สามคดีที่ถูกอ้างว่าเป็น trend ก็ไม่เห็นตรงกัน Alsup พบว่าการ train บนหนังสือที่ได้มาอย่างถูกกฎหมาย “exceedingly transformative” ขณะที่ถือว่า library ที่สร้างจากสำเนาละเมิดลิขสิทธิ์ไม่ใช่ และ Anthropic settled ครึ่งนั้นเป็นเงิน $1.5 billion ครอบคลุมงาน 482,460 ชิ้น หรือประมาณ $3,000 ต่อชิ้น อนุมัติ 20 กรกฎาคม 202620 Chhabria ให้ summary judgment แก่ Meta พร้อมเขียนว่าคำวินิจฉัยของเขา “does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful” เพียงแต่ว่า “these plaintiffs made the wrong arguments”21 Bibas ซึ่งตัดสิน against Ross Intelligence ระบุว่า “only non-generative AI is before me today”22 ยังไม่มี US appellate court ใดตัดสินคำถามนี้
คนทำส่วนที่ loss ทำไม่ได้ TIME รายงานในเดือนมกราคม 2023 ว่าคนงานที่ label toxic text ให้ OpenAI ผ่านบริษัท Sama รับกลับบ้าน “between around $1.32 and $2 per hour” เพื่ออ่านข้อความที่บรรยาย child sexual abuse, torture และ self-harm ขณะที่ OpenAI จ่าย Sama $12.50 ต่อชั่วโมงสำหรับงานนี้; Sama โต้แย้งทั้งช่วงค่าจ้างและ quota23 นั่นคือ filtering รอบ pretraining ไม่ใช่ pretraining เอง — แต่มันอยู่ใน invoice เดียวกัน และเป็นจุดที่มีคนคนหนึ่งนั่งอยู่
ไฟฟ้าเป็นของจริงและมักถูก quote ผิด ตัวเลขตีพิมพ์ที่ระมัดระวังที่สุดคือของ BLOOM: 1,082,990 GPU-hours, 433 MWh และ 24.7 tonnes ของ CO₂ equivalent สำหรับ run, 50.5 เมื่อนับ manufacturing และ idle nodes;24 Patterson et al. วาง GPT-3 ที่ 1,287 MWh และ 552 tonnes25 ข้อควรระวังสองข้อ ข้อได้เปรียบของ BLOOM คือ grid nuclear ฝรั่งเศสที่ 57 g CO₂ ต่อ kWh ไม่ใช่ efficiency — มันใช้พลังงาน มากกว่า OPT-175B และตัวเลข emissions ที่ถูก quote มากที่สุดของวงการ คือ 626,155 lb ของ Strubell et al. สำหรับ neural architecture search ภายหลังถูกแสดงว่า สูงเกินจริง 88 เท่า เพราะ assume ว่า search run ที่ full model size ทั้งที่ run บน proxy26 framing ของ LBNL คือแบบที่ปกป้องได้: US data centres ใช้ 192 TWh ในปี 2024, 4.7% ของไฟฟ้าทั้งประเทศ — ตัวเลขที่ผูกกับอุตสาหกรรม ไม่ใช่กับ run ใด run หนึ่ง27
เส้นที่ลากผ่านทั้งหมดคือสิ่งที่ Bender et al. ตั้งชื่อว่า documentation debt: “putting ourselves in a situation where the datasets are both undocumented and too large to document post hoc”28 ทุกข้อเท็จจริงข้างต้นมีอยู่เพราะมีใครบางคนไปดู สำหรับ corpora เบื้องหลังโมเดลที่คนส่วนใหญ่ใช้ ไม่มีใครทำได้
จริง ๆ แล้วมันราคาเท่าไร
ลิงก์ไปยังส่วน: จริง ๆ แล้วมันราคาเท่าไรตอนนี้คือเลขคณิตที่ทุกคนอยากได้ จาก inputs ที่ cite ไว้สี่รายการ เพื่อให้เมื่อมัน stale จะเห็นชัดว่าต้อง replace อะไร
Peak throughput
ลิงก์ไปยังส่วน: Peak throughputหน้า H100 ของ NVIDIA ระบุ BF16 tensor-core throughput 1,979 teraFLOPS ใต้ footnote ที่เขียนว่า “with sparsity”29 ไม่มี pretraining run ใดใช้ structured sparsity ดังนั้นตัวเลข dense คือครึ่งหนึ่งของมัน: 989.5 TFLOP/s
Utilisation
ลิงก์ไปยังส่วน: UtilisationTable 4 ของ Llama 3 รายงาน BF16 model FLOPs utilisation 38–43% ใช้ 40%: useful arithmetic 395.8 TFLOP/s ต่อ GPU5
ราคา on-demand ของ Lambda สำหรับ node 8×H100 SXM เข้าถึงเมื่อ 2026-09-06: $3.99 ต่อ GPU-hour ดังนั้น $31.92 ต่อชั่วโมงสำหรับ node30
ratio ของ Chinchilla, , ให้ และดังนั้น
| budget | H100-hours | FLOPs | compute-optimal params | tokens | บน node 8×H100 หนึ่งตัว | GPUs เพื่อจบใน 90 วัน |
|---|---|---|---|---|---|---|
| $100 | 25 | 3.6e19 | 546 M | 10.9 B | 3.1 h | 1 |
| $1,000 | 251 | 3.6e20 | 1.73 B | 34.5 B | 31.3 h | 1 |
| $10,000 | 2,506 | 3.6e21 | 5.46 B | 109 B | 13 days | 2 |
| $100,000 | 25,063 | 3.6e22 | 17.3 B | 345 B | 131 days | 12 |
| $1,000,000 | 250,627 | 3.6e23 | 54.6 B | 1.09 T | 4 years | 116 |
| $10,000,000 | 2,506,266 | 3.6e24 | 173 B | 3.45 T | 36 years | 1,160 |
| $100,000,000 | 25,062,657 | 3.6e25 | 546 B | 10.9 T | 358 years | 11,603 |
อ่านสองคอลัมน์สุดท้ายคู่กัน ที่ $10,000 คุณได้โมเดล 5-billion-parameter บน node เช่าหนึ่งตัวในสองสัปดาห์ ที่ $100,000,000 เลขคณิตบอกว่า 546 พันล้าน parameters — และ H100 สิบสองพันตัวที่ต่อสายเข้าด้วยกันเป็นเวลาสามเดือน ซึ่งไม่ใช่สิ่งที่คุณเช่าด้วยบัตรเครดิต เลยประมาณ $100,000 ไปแล้ว binding constraint หยุดเป็นเงินและกลายเป็น cluster
ก่อนเชื่อตารางแบบนั้น ให้ทดสอบกับ runs ที่มี real cost เผยแพร่ — llm.c reproduce GPT-2 124M ใน “~90 minutes” บน node 8×A100 “for about $20” และ GPT-2 1.6B ใน 24 ชั่วโมงบน node 8×H100 ในราคา $67231
$672, against what $672 actually bought (llm.c GPT-2 1.6B, one 8xH100 node, 24 h)
this table predicts: 168 H100-hours N = 1.41 B params D = 28.3 B tokens
what was actually run: 192 H100-hours N = 1.558 B params D = 33.6 B tokens
Llama 3 405B, against Meta's own published GPU-hours
from the paper's 3.8e25 FLOPs at 40 % MFU: 26.67 M H100-hours
published in Meta's Llama 3.1 model card: 30.84 M H100-hours ratio 0.86ทั้งสองอยู่ภายในประมาณ 15% ซึ่งเป็นความแม่นยำที่ประมาณการแบบนี้ควรได้รับโดยประมาณ และดีกว่าความแม่นยำที่มันมักถูก quote อย่างมาก
การเปรียบเทียบ headline พร้อมทั้งสองนิยามบนโต๊ะ
ลิงก์ไปยังส่วน: การเปรียบเทียบ headline พร้อมทั้งสองนิยามบนโต๊ะตัวเลขที่ถูกพูดซ้ำมากที่สุดในเรื่องนี้คือโมเดลระดับ GPT-2 ที่มีต้นทุนประมาณ $43,000 ในปี 2019 สามารถ reproduce ได้วันนี้ในราคาไม่กี่สิบดอลลาร์ ครึ่งสมัยใหม่มีเอกสารดี; ครึ่งประวัติศาสตร์ไม่มี
วันนี้ README ของ nanochat โดย Karpathy: “you can train your own GPT-2 capability LLM ... for only $48 (~2 hours of 8XH100 GPU node) ... On a spot instance, the total cost can be closer to ~$15.”32 “GPT-2 capability” ตรงนี้มีความหมาย precise และเผยแพร่แล้ว — เอาชนะ CORE score 0.256525 ของ GPT-2 — บน leaderboard ที่ entry ดีที่สุด ณ 14 มีนาคม 2026 คือ 1.65 ชั่วโมง ค่า $48 assume $3 ต่อ GPU-hour ต่ำกว่า list ของ Lambda ที่ $3.99; ที่ list จะใกล้ $64
ในปี 2019 ไม่มี primary source: OpenAI ไม่เคยเผยแพร่ duration หรือ cost chain คือ The Register, กุมภาพันธ์ 2019 รายงาน “256 Google TPU3 cores” โดยไม่มีราคาและไม่มี duration; จากนั้น Synced, มิถุนายน 2019 ระบุว่า hardware cost $256 ต่อชั่วโมงบน Google Cloud และเขียนชัดว่า “OpenAI didn't specify the training duration” $43,008 คือ $256 ต่อชั่วโมงคูณ 168 ชั่วโมงที่ assume ขึ้นมาโดยไม่มีใครเคย source
ดังนั้น headline ที่ซื่อตรงคือ: โมเดลที่ match benchmark score ที่ตีพิมพ์ของ GPT-2 สามารถ train ได้วันนี้ในราคาต่ำกว่า $100 มากบน hardware เช่า เทียบกับต้นทุนปี 2019 ที่ไม่เคยถูกเผยแพร่ และ estimate อันโด่งดังของมันตั้งอยู่บนการเดา duration แบบไม่มี source การพังทลายของต้นทุนเป็นเรื่องจริง และครึ่งสมัยใหม่ reproduce ได้โดยใครก็ตามที่มีบัตรเครดิต; ratio คือเลขคณิตบนตัวเลขที่ไม่มีอยู่ นั่นคือสภาพของ published training costs โดยทั่วไป Paper GPT-3 ไม่มีจำนวนเงินดอลลาร์เลย มีเพียง FLOPs ใน Table D.1;7 paper Llama 3 ก็ไม่มีเช่นกัน5 training cost ทุกตัวที่คุณเคยอ่านคือ estimate จาก FLOP count, hardware assumption และ price assumption — ควรถามเสมอว่าเป็นของใคร
Base model รู้อะไร และมันหยุดรู้เมื่อไหร่
ลิงก์ไปยังส่วน: Base model รู้อะไร และมันหยุดรู้เมื่อไหร่สิ่งที่ออกมาได้เห็น corpus คงที่ที่ประกอบขึ้น ณ เวลาคงที่ และมี properties สองอย่างตามมา
อย่างแรกคือ knowledge cutoff หลังวันที่เก็บรวบรวม โมเดลไม่รู้อะไรเลย — ไม่ใช่ “ไม่แน่ใจ” แต่ ไม่มีอะไร — และมันจะ confabulate ได้อย่างลื่นไหลแทนที่จะบอกเช่นนั้น เพราะการบอกเช่นนั้นไม่เคยเป็น behaviour ที่มันถูก train ให้ทำ model card ของ Llama 3.1 ให้เดือนธันวาคม 2023;33 ทุกโมเดลมีหนึ่งค่า และมันเป็น property ของ training data ไม่ใช่ deployment การ workaround เรื่องนี้เป็น retrieval problem ซึ่งคือ บทที่ 19
อย่างที่สองคือ base model เติมต่อมากกว่าตอบ ให้ “เมืองหลวงของฝรั่งเศสคืออะไร?” กับมัน แล้ว continuation ที่เป็นไปได้คือคำถามอีกข้อ เพราะใน corpus string นั้นมักปรากฏในรายการแบบฝึกหัด
จากนี้ไปที่ไหน
ลิงก์ไปยังส่วน: จากนี้ไปที่ไหนtext completer ไม่ใช่ assistant มันไม่ทำตาม instructions เพราะไม่มีอะไรใน corpus บอกว่าคำขอควรถูกเชื่อฟังแทนที่จะถูกต่อ มันไม่มี notion ของ conversation ที่มีผู้เข้าร่วมสองฝ่าย มันจะสร้าง continuation ที่น่าจะเป็นที่สุดของ prompt อันตรายอย่างเต็มใจ เพราะ probable คือสิ่งเดียวที่มันเคยถูก optimise เพื่อให้ได้
การเปลี่ยนมันเป็นสิ่งที่ตอบได้ต้องใช้ stage ที่สองซึ่ง cost เพียงเศษเสี้ยวของหนึ่งเปอร์เซ็นต์ของ stage แรก และประกอบเกือบทั้งหมดจากการแสดง examples ของ behaviour ที่คุณต้องการ แล้วเปรียบเทียบ outputs ของมันเองเป็นคู่ ๆ stage นั้นคือที่มาของ instruction following, chat templates, refusals และ — สิ่งนี้ทำให้คนแปลกใจ — ความสามารถในการ call tool ทั้งหมด บทที่ 11 คือ stage นั้น: supervised fine-tuning, RLHF, DPO และ GRPO, และคำถามว่า “aligned” หมายถึงอะไรและใครเป็นคนตัดสิน
Sources and method
ลิงก์ไปยังส่วน: Sources and methodควรอ่านคู่กับบทนี้ด้วย: build-nanogpt ของ Karpathy และ video ประกอบ ซึ่งพา reproduce GPT-2 แบบครบ end to end ด้วย pace ที่บทนี้ทำไม่ได้; และ Stanford CS324, Large Language Models, ซึ่ง lectures เรื่อง data และ environmental impact ลงลึกกว่าส่วนข้างบนใน material ที่คอร์สนี้แตะครั้งเดียวแล้วมอบหมายต่อ
รายการอ้างอิง
ลิงก์ไปยังส่วน: รายการอ้างอิง-
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D. Scaling Laws for Neural Language Models. arXiv:2001.08361 (2020). power laws สามชุดคือ Eqs. (1.1)–(1.3) ใน §1.2 และ constants ครบอยู่ใน Appendix A, Table 5; การอนุมาน คือ §2.1; exponents ของ compute-allocation อยู่ใน Table 6 โปรดสังเกตว่ามี compute laws สองชุดคือ ที่ fixed batch size และ ที่ optimal batch size; paper บอกว่าชุดหลัง “should be used to make predictions”. ↩ ↩2
-
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E. et al. Training Compute-Optimal Large Language Models. arXiv:2203.15556 (2022). Exponents ใน Table 2, budgets ที่ project ใน Table 3, การเปรียบเทียบ Gopher ใน §4, convention การนับ parameters ใน Appendix F prose ใต้ Table 3 ไม่ตรงกับ Table 3 เองสำหรับแถว 175 B และ 280 B; table คือเวอร์ชันที่ควร quote. ↩ ↩2 ↩3 ↩4
-
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D. et al. Mixed Precision Training. arXiv:1710.03740 (2017), ICLR 2018. FP32 master weights ใน §3.1, loss scaling ใน §3.2. ↩ ↩2
-
Rajbhandari, S., Rajbhandari, S., Ruwase, O. and He, Y. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 (2019), SC20. บัญชี คือ §3.1; ตัวเลข residual-state สำหรับ activations คือ §3.2. ↩
-
Grattafiori, A. et al. (Llama Team, AI @ Meta). The Llama 3 Herd of Models. arXiv:2407.21783 (2024). Compute budget และ token count ใน §1, scaling law ที่ refit ใน §3.2.1, configuration ของ parallelism และ MFU ใน Table 4, contamination analysis ใน §5.1.4, statement เรื่อง over-training ใน §9.1 paper ไม่มี dollar figures และไม่มี emissions table. ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
Stanford CS336, Language Modeling from Scratch. Lecture 2 ครอบคลุม resource accounting, lectures 5–8 ครอบคลุม GPUs, kernels และ parallelism, lectures 9 และ 11 ครอบคลุม scaling, lectures 13–14 ครอบคลุม data นี่คือคอร์สที่บทนี้มอบหมาย engineering ของมันต่อไป และเป็นสาธารณะ. ↩
-
Brown, T. B. et al. Language Models are Few-Shot Learners. arXiv:2005.14165 (2020). Compute ใน Appendix D, Table D.1 — ซึ่งมีคอลัมน์ที่ขึ้นหัวตรงตัวว่า “flops per param per token” และค่าของทุกแถว GPT-3 คือ 6 Contamination analysis ใน §4. ↩ ↩2 ↩3
-
Besiroglu, T., Erdil, E., Barnett, M. and You, J. Chinchilla Scaling: A replication attempt. arXiv:2404.10102 (2024). สร้าง data ของ Chinchilla ใหม่ด้วยการ digitise Figure 4, refit และรายงาน exponents ที่แก้แล้วกับ intervals ที่กว้างขึ้นมาก. ↩
-
Muennighoff, N., Rush, A. M., Barak, B., Le Scao, T., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T. and Raffel, C. Scaling Data-Constrained Language Models. arXiv:2305.16264 (2023), NeurIPS 2023. ผลลัพธ์ four-epoch อยู่ใน §6; half-life สิบหก-epoch คือ fitted . ↩
-
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T. et al. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 (2023). §1 ระบุ inference-cost argument against Chinchilla-optimal training. ↩
-
Sardana, N., Portes, J., Doubov, S. and Frankle, J. Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws. arXiv:2401.00448 (2023), ICML 2024. §5 ของพวกเขายังมี counterweight: โมเดลที่ train ที่ token ratios สุดโต่งยังดีขึ้น แต่ “more slowly than scaling laws predict”. ↩
-
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S. et al. Emergent Abilities of Large Language Models. arXiv:2206.07682 (2022), TMLR. นิยามอยู่ใน §2, examples และ compute thresholds ใน §3–4 และ Table 1. ↩
-
Schaeffer, R., Miranda, B. and Koyejo, S. Are Emergent Abilities of Large Language Models a Mirage? arXiv:2304.15004 (2023), NeurIPS 2023 outstanding paper. argument เรื่อง metric คือ §2, meta-analysis ของ BIG-Bench คือ §4, vision example ที่สร้างขึ้นคือ §5. ↩
-
Common Crawl, August 2026 Crawl Archive Now Available (CC-MAIN-2026-34), เผยแพร่ 24 August 2026, เข้าถึง 2026-09-06 หน้าแรกของมันเอง claim ว่า “over 300 billion pages spanning 15 years”, “totalling more than 10 petabytes” — เป็นตัวเลขของ archive ทั้งหมด ไม่ใช่ monthly crawl ที่ตีราคาไว้ตรงนี้. ↩
-
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W. and Liu, P. J. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 (2019), JMLR 21(140). filters ของ C4 อยู่ใน §2.2 paper ให้ขนาดเป็น bytes ไม่ใช่ tokens; ตัวเลข 156-billion-token ที่มักถูก attribute ให้มันมาจาก Dodge et al. ข้างล่าง. ↩
-
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M. and Gardner, M. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758 (2021), EMNLP 2021. อัตราการลบ dialect อยู่ใน §5.3; benchmark contamination ใน C4 อยู่ใน §4.2. ↩
-
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C. and Carlini, N. Deduplicating Training Data Makes Language Models Better. arXiv:2107.06499 (2021), ACL 2022. การซ้ำ 61,036 ครั้งอยู่ใน footnote 1; ตัวเลข memorisation อยู่ใน §6.2, Table 4 และเป็นเปอร์เซ็นต์ของ generated tokens ภายใต้ criterion exact-match 50-token. ↩
-
Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L. and Wolf, T. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557 (2024), NeurIPS 2024 Datasets and Benchmarks. ผล deduplication อยู่ใน §3.4 dataset ที่ release โตเกิน 15 ล้านล้าน tokens ของ paper ไปแล้ว. ↩
-
Gao, L., Biderman, S., Black, S. et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv:2101.00027 (2020). Books3 อยู่ใน §2.3 และ Table 1; consent table คือ Table 5 corpus มี 825.18 GiB ดังนั้นแม้แต่ title ก็เป็นการปัดลง. ↩
-
Bartz v. Anthropic, No. 4:24-cv-05417 (N.D. Cal.). คำสั่ง fair-use 23 June 2025 (Dkt. 231); class certification 17 July 2025; final approval and judgment 20 July 2026 (Dkt. 680). Settlement ปล่อยเฉพาะ past inputs ไม่ใช่ outputs และไม่ใช่ future conduct. ↩
-
Kadrey v. Meta, No. 3:23-cv-03417-VC (N.D. Cal.), summary judgment 25 June 2025 (Dkt. 598). โปรดสังเกตว่า distribution claim เรื่อง torrenting ยังไม่ได้ตัดสินและยัง live อยู่. ↩
-
Thomson Reuters v. ROSS Intelligence, No. 1:20-cv-00613-SB (D. Del.), revised opinion 11 February 2025 (Dkt. 770), Bibas J. อยู่ระหว่าง interlocutory appeal ต่อ Third Circuit (No. 25-2153), argued 11 June 2026, ยังไม่ตัดสิน ณ เวลาที่เขียน. ↩
-
Perrigo, B. Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. TIME, 18 January 2023. $2 เป็น ceiling สำหรับ senior reviewers ที่ทำได้ทุก target; junior labellers ซึ่งเป็นส่วนใหญ่รับกลับบ้าน $1.32 คำโต้แย้งของ Sama ที่ quote ใน article เดียวกันให้ $1.46–$3.74 และ quota ต่ำกว่า. ↩
-
Luccioni, A. S., Viguier, S. and Ligozat, A.-L. Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. arXiv:2211.02001 (2022), JMLR 24(253). Tables 1 และ 3. ↩
-
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M. and Dean, J. Carbon Emissions and Large Neural Network Training. arXiv:2104.10350 (2021). ตัวเลขของ GPT-3 คือ Table 4; การแก้ estimate ของ NAS คือ §4.1. ↩
-
Strubell, E., Ganesh, A. and McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. arXiv:1906.02243 (2019), ACL 2019. ควรอ่านอย่างละเอียดเพราะสิ่งที่เกิดกับตัวเลขที่ถูก quote มากที่สุดของมัน: paper ระมัดระวัง ระบุ extrapolation ของตัวเอง และยังผิดสอง orders of magnitude ในบรรทัดเดียวที่ทุกคนพูดซ้ำ. ↩
-
Smith, S. J., Hubbard, A., Newkirk, A., Ganeshalingam, M., Holecek, B., Sartor, D., Mills, M. and Shehabi, A. United States Data Center Energy Usage Report: 2025 Update. LBNL-2001758 (18 June 2026). สิ่งนี้ revise historical series ที่ถูก cite กว้างขวางในรายงาน 2024 ลง; ถ้าคุณ quote ตัวเลข 176 TWh สำหรับปี 2023 คุณกำลัง quote edition ที่ถูกแทนที่แล้ว. ↩
-
Bender, E. M., Gebru, T., McMillan-Major, A. and Shmitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? FAccT '21, pp. 610–623. DOI 10.1145/3442188.3445922. “Documentation debt” คือ §4.4 โปรดสังเกตว่า carbon figures ของ paper เอง cite จาก Strubell et al. และรับ correction ข้างต้นต่อมา — ซึ่งเป็น illustration ของ argument ของมัน มากกว่าจะเป็น refutation. ↩
-
NVIDIA. NVIDIA H100 Tensor Core GPU product page,
nvidia.com/en-us/data-center/h100/(เข้าถึง 2026-09-06). ทุกแถว tensor-core บนหน้านั้นยกเว้น FP64 มี footnote “with sparsity”; ตัวเลข dense BF16 ที่ใช้ที่นี่คือครึ่งหนึ่งของ 1,979 TFLOPS ที่เผยแพร่. ↩ -
Lambda. GPU Cloud pricing,
lambda.ai/pricing(เข้าถึง 2026-09-06). On-demand, ต่อ GPU ต่อชั่วโมง, ก่อนภาษี ราคาในส่วนนี้จะ stale เร็วกว่าอย่างอื่นในคอร์สนี้; เลขคณิตรอบ ๆ มันจะไม่ stale. ↩ -
Karpathy, A.
karpathy/llm.c, discussion #481, Reproducing GPT-2 (124M) in llm.c in 90 minutes for $20 (28 May 2024), และ discussion #677, Let's reproduce GPT-2 (1.6B): one 8XH100 node, 24 hours, $672, in llm.c (11 July 2024). ↩ -
Karpathy, A.
karpathy/nanochat, README และ leaderboard “time to GPT-2” (เข้าถึง 2026-09-06). ตัวเลข $48 และนิยาม CORE-score ของ “GPT-2 capability” อยู่ใน README ทั้งคู่;speedrun.shของ repository เองบอกว่า “approximately 1.5 hours” ดังนั้นให้ถือเลขสองชั่วโมงเป็นการปัดเศษ. ↩ -
Meta. Llama 3.1 model card,
models/llama3_1/MODEL_CARD.mdในmeta-llama/llama-models(เข้าถึง 2026-09-06). แหล่งของ 30.84 M H100-hours สำหรับโมเดล 405 B, total 39.3 M, ตัวเลข location-based 11,390 tCO2eq และ data cutoff เดือนธันวาคม 2023. ↩