ข้ามไปยังเนื้อหา
10/30บทที่ 10 จาก 30

Pretraining LLM: data, compute, scaling laws และต้นทุน

ฝึก 20 โมเดลบน GPU แล็ปท็อปเพื่อวัด scaling law และตรวจสอบประมาณการ compute 6ND กับตัวนับ FLOP จริง

ในหน้านี้

บทที่ 9 จบลงด้วย transformer block ที่ train ได้ เอาหลาย block มาซ้อนกัน ชี้ next-token loss จาก บทที่ 8 ไปที่ output แล้วก็ไม่มีอะไรเหลือให้คิดค้นอีก สิ่งที่เหลือทั้งหมดคือการซื้อ

นี่เป็นการเปลี่ยนแปลงที่ใหญ่กว่าที่ฟังดูเป็นมาก ทุกบทก่อนหน้านี้ถามว่า มันเรียนรู้ไหม? — คำถามแบบใช่หรือไม่ใช่ที่แล็ปท็อปตอบได้ในสิบ分钟 บทนี้ถามคำถามที่มีเงินอยู่ในนั้น: เมื่อมีปริมาณเลขคณิตคงที่ โมเดลที่ดีที่สุดที่ฉันซื้อได้คืออะไร? คำตอบคือสูตรหนึ่ง และในปี 2018 ไม่มีใครเห็นว่ามันชัดเจนเลย

นี่คือคำถามนั้นที่ตอบด้วยการวัด บน GPU แล็ปท็อปเครื่องเดียว โมเดล 20 ตัว ตั้งแต่ 98,624 ถึง 15 ล้าน parameters ถูก train จากศูนย์บน Wikipedia 174 ล้าน tokens — vocabulary แบบ BPE ขนาด 2,048-token ที่ train แบบเดียวกับ บทที่ 7, transformer ของบทที่ 9 แต่ละ run ได้ compute budget หนึ่งในสามชุดแบบ เป๊ะ ๆ และไม่มากกว่านั้นแม้แต่ operation เดียว ดังนั้นโมเดลที่ใหญ่กว่าจำเป็นต้องอ่าน text น้อยกว่า loss บน held-out ที่ดีที่สุดในแต่ละ budget คือ:

TEXT
budget C (FLOPs)   best loss   reached by a model of
       1.00e13       5.3531           98,624 params
       3.16e13       4.8638           98,624 params
       1.00e14       4.3383          295,808 params

fitted:  L = (Cc / C)^0.0913     over one decade of compute

เลขคณิตมากขึ้นสิบเท่าลด loss ลง 19% และสามจุดอยู่บนเส้นตรงใน log-log ไม่มีอะไรในเก้าบทแรกทำนายสิ่งนี้ได้ ไม่มี theorem รองรับ — มันคือความสม่ำเสมอเชิงประจักษ์ ซึ่งคงอยู่ด้วย exponent ต่างออกไปตลอดสิบ order of magnitude ระหว่างแล็ปท็อปเครื่องนี้กับ datacentre และเป็นข้อสังเกตเดียวที่โน้มน้าวทั้งอุตสาหกรรมให้ใช้เงินเท่า GDP ของประเทศเล็ก ๆ กับ GPUs

Pretraining คืออะไร และมีอะไรใหม่เกี่ยวกับมัน

ลิงก์ไปยังส่วน: Pretraining คืออะไร และมีอะไรใหม่เกี่ยวกับมัน

ไม่มีอะไรใน objective เปลี่ยนไป โมเดลยังทำนาย token ถัดไป loss ยังเป็น cross-entropy จาก บทที่ 4 ที่นำไปใช้กับ factorisation ของบทที่ 8 optimiser ยังเป็น AdamW จาก บทที่ 6 Pretraining ไม่ใช่ algorithm ใหม่; มันคือ algorithm เดิมที่ run บน corpus ใหญ่พอจนต้องทำ budget ให้ run นั้น สองสิ่งทำให้เป็นไปได้: labels ฟรี เพราะ target สำหรับตำแหน่ง tt คือ token ที่ t+1t+1 และมีอยู่ใน text อยู่แล้ว; และส่วนสุดท้ายของบทที่ 6 เอาข้อคัดค้านออกไป เพราะโมเดลที่มี parameters มากกว่ากฎคลาสสิกอนุญาตอย่างมากไม่ได้พัง แต่มันดีขึ้น สิ่งที่ได้ออกมาคือ base model — บางสิ่งที่ต่อ text ไม่ใช่ตอบคำถาม

ก่อนจะทำ budget ให้สิ่งนี้ได้ ต้องนับมันก่อน และวงการนับด้วยสูตรเดียว:

C6NDC \approx 6ND

โดยที่ NN คือจำนวน parameters, DD คือ training tokens และ CC คือ floating-point operations ทั้งหมด Kaplan et al. อนุมานมันเป็นสองขั้นตอน1 Forward: 2 FLOPs ต่อ parameter ต่อ token เพราะทุก parameter ใน matrix multiply ถูกใช้หนึ่งครั้งต่อ token ในการคูณหนึ่งครั้งและบวกหนึ่งครั้ง Backward: สองเท่าของ forward เพราะ backward pass จาก บทที่ 5 คำนวณ gradients สองชุดในแต่ละ layer — เทียบกับ inputs ของ layer เพื่อให้ signal เดินทางต่อไป และเทียบกับ weights ของมัน — แต่ละชุดเป็น matrix multiply ขนาดเท่า forward หนึ่งครั้ง ดังนั้น 4N4N

นั่นคือการอนุมานทั้งหมด และควรตรวจสอบมากกว่าเชื่อ PyTorch มีตัวนับ FLOP จริงคือ torch.utils.flop_counter.FlopCounterMode ซึ่ง intercept ทุก operation ที่โมเดล dispatch แล้วรวมงานจริงทั้งหมด ลอง run มันข้ามสี่ order of magnitude โดยตัวใหญ่สุดบน device meta ซึ่ง allocate shapes แต่ไม่ใช้ memory:

flops.pyPYTHON
from torch.utils.flop_counter import FlopCounterMode

counter = FlopCounterMode(display=False)
with counter:                       
    loss = model(x, targets)[1]     
    loss.backward()                 
measured = counter.get_total_flops()
print(measured / (6 * n_params * n_tokens))
configurationNN ไม่รวม embeddingsNN ทั้งหมดวัดได้, fwd+bwd÷ 6ND6ND (NN ทั้งหมด)÷ 6ND6ND (ไม่รวม emb.)fwd+bwd ÷ fwd
dd 128, 4 layers, TT 256788,7367,254,4004.60e101.0319.4853.000
dd 512, 8 layers, TT 25625,183,23251,045,8883.26e111.0382.1043.000
dd 768, 12 layers, TT 102484,973,056124,356,8641.75e121.1451.6763.000
dd 1600, 48 layers, TT 10241,474,870,4001,556,920,0002.10e131.1001.1613.000
dd 4096, 32 layers, TT 20486,442,983,4246,582,444,0328.74e131.0801.1043.000
dd 8192, 80 layers, TT 819264,427,147,26465,544,929,2803.75e151.1631.1833.000

อัตราส่วน forward+backward ต่อ forward คือ 3.000 แบบเป๊ะ ๆ ในทุก scale: ไม่ใช่ approximation ที่บังเอิญดี แต่คือเอกลักษณ์ทางเลขคณิตข้างต้นที่ถูกตัวนับซึ่งไม่รู้อะไรเกี่ยวกับการอนุมานนั้นคืนกลับมาเป็นเลขกลม

จากนั้นค่ารวมที่วัดได้อยู่สูงกว่า 6ND6ND ระหว่าง 3% ถึง 17% เมื่อ NN นับ embedding matrices — และ clause นั้นสำคัญ เพราะ paper ผู้ก่อตั้งสองฉบับนับ NN ต่างกัน Kaplan ไม่รวม “all vocabulary and positional embeddings” เพราะการทำเช่นนั้น “produces significantly cleaner scaling laws” (§1.3); Appendix F ของ Chinchilla บอกว่า “we also count embeddings matrices in the total parameter count”2 สำหรับ vocabulary กว้างและ hidden dimension แคบ ทั้งสองต่างกันเป็น factor เก้าอย่างที่แถวแรกแสดง

ช่องว่างที่เหลือคือสิ่งที่ 6ND6ND จงใจละไว้: attention scores Eq. (2.2) ของ Kaplan เขียนต้นทุน forward เป็น 2N+2nlayernctxdmodel2N + 2\,n_{\text{layer}} n_{\text{ctx}} d_{\text{model}} และตัด term ที่สองทิ้งเพราะ dmodelnctx/12d_{\text{model}} \gg n_{\text{ctx}}/12 — ปลอดภัยในปี 2020 ปลอดภัยน้อยลงตอนนี้ และเป็นเหตุผลที่อัตราส่วน drift สูงขึ้นเมื่อ T/dT/d โตขึ้น — ซึ่งเป็นเหตุผลที่สองแถวตรงนี้มี TT เท่ากันที่ 1,024 และอัตราส่วน ลดลง จาก 1.145 เป็น 1.100 เมื่อ dd เพิ่มจาก 768 เป็น 1,600 มันคือ cost O(T2)O(T^2) ที่บทที่ 9 แนะนำ และ บทที่ 16 จะเปลี่ยนให้เป็นราคา

Compute ตัดสินว่า run จะใช้เวลานานแค่ไหน; memory ตัดสินว่ามันเริ่มได้ไหม Train ด้วย fp32 AdamW ธรรมดา แล้วทุก parameter แบกตัวเลขสี่ตัว: weight, gradient ของมัน, และ running mean mm กับ variance vv ของ Adam — ค่าเฉลี่ยสองตัวที่สร้างด้วยมือในบทที่ 6 ตัวเลขสี่ตัวที่ตัวละสี่ bytes คือ 16 bytes ต่อ parameter ก่อนมี activation แม้แต่ตัวเดียว วัดบน GPU แล็ปท็อป 8 GB โดยดู resident allocation ณ จุดใน step ที่ไม่มี graph มีชีวิตอยู่:

modelvocabularybatchNN16N16N คาดการณ์resident ที่วัดได้peak ในหนึ่ง stepส่วนต่าง
dd 512, 8 layers50,257851,045,888779 MB801 MB2,500 MB1,699 MB
dd 512, 8 layers4,096827,411,456418 MB426 MB1,043 MB617 MB
dd 256, 6 layers4,09685,839,36089 MB89 MB382 MB293 MB
dd 256, 6 layers4,096325,839,36089 MB89 MB1,259 MB1,170 MB
dd 256, 6 layers4,0961285,839,36089 MB89 MB4,771 MB4,681 MB

การคาดการณ์และการวัดตรงกันภายใน 3% สิ่งที่น่าประหลาดใจคือคอลัมน์สุดท้าย: activations ใหญ่กว่าโมเดลอย่างท่วมท้น โมเดล 5.8-million-parameter ตัวเดิมที่ต้องใช้ persistent state 89 MB ต้องใช้ activations 4,681 MB ที่ batch 128 — ห้าสิบสองเท่าของโมเดล — และส่วนมากของสิ่งนั้นไม่ใช่ transformer เลย มันคือ logits, vector ขนาด vocabulary หนึ่งตัวต่อ token ที่ entry ละสี่ bytes: 512 MB ในแถวสุดท้าย, 393 MB ในแถวแรก ขนาด vocabulary ถูกเลือกในบทที่ 7 และมันยังคงตัดสินว่าอะไร fit บนการ์ด

term ไหนครองขึ้นกับรูปร่างของ run ซึ่งเป็นเหตุผลที่ Micikevicius et al. บอกว่า memory “is dominated by activations”3 ขณะที่ ZeRO บอกว่าโมเดล 1.5-billion-parameter ต้องการ model states อย่างเดียว “at least 24 GB”4 ZeRO มาถึง 16 bytes เดิมผ่านอีกเส้นทาง — 2Ψ2\Psi สำหรับ fp16 weights, 2Ψ2\Psi สำหรับ fp16 gradients, 4Ψ4\Psi อย่างละตัวสำหรับ fp32 master weights และ moments สองตัวของ Adam — ซึ่งสำหรับ 70 พันล้าน parameters คือ 1.12 terabytes เท่ากับ GPUs 80 GB สิบสี่ตัวก่อนมี activation แม้แต่ตัวเดียว

Parallelism ในหนึ่งย่อหน้าและหนึ่งการมอบหมายต่อ

ลิงก์ไปยังส่วน: Parallelism ในหนึ่งย่อหน้าและหนึ่งการมอบหมายต่อ

ที่ frontier scale ไม่มีสิ่งใด fit บนอุปกรณ์เดียว ดังนั้น run จึงถูกแบ่งพร้อมกันสี่ทาง Data parallelism วางสำเนาโมเดลไว้บนทุก GPU แล้วเฉลี่ย gradients — ค่า default และเป็นแบบที่ ZeRO ปรับปรุงด้วยการไม่เก็บสำเนาซ้ำซ้อนของ optimiser state Tensor parallelism แยก matrices แต่ละตัวข้าม devices Pipeline parallelism ให้แต่ละ device รับกลุ่ม layers ที่ต่อเนื่องกัน Context parallelism แยก sequence เอง ซึ่งจำเป็นก็ต่อเมื่อ TT ยาวพอให้ attention term ครอง Table 4 ของ Llama 3 แสดงทั้งสี่พร้อมกัน: tensor 8, context สูงสุด 16, pipeline 16, data สูงสุด 128 ข้าม 16,384 H100 GPUs5 คอร์สนี้จะพูดถึงมันเท่านี้; engineering ของ distributed training เป็นอีกหนึ่ง semester เต็ม ๆ และ Stanford CS336 คือ semester นั้น lectures 5 ถึง 8 พร้อม code6 สิ่งที่เหลือหลังการมอบหมายต่อคือเลขเดียว model FLOPs utilisation — สัดส่วนของ peak arithmetic ของ GPU ที่ run จริงทำได้ — ซึ่งเป็นสิ่งที่เปลี่ยน 6ND6ND ที่เรียบร้อยให้เป็น wall-clock time และดังนั้นให้เป็นเงิน

Kaplan และการเดิมพันที่อุตสาหกรรมวาง

ลิงก์ไปยังส่วน: Kaplan และการเดิมพันที่อุตสาหกรรมวาง

ในเดือนมกราคม 2020 Kaplan et al. train grid ของ transformers และพบว่า test loss ตาม power law ในทรัพยากรทั้งสามตลอดมากกว่าหก orders of magnitude1 §1.2 ของพวกเขาให้ fitted laws สามชุด:

L(N)=(NcN)αN,αN0.076,Nc8.8×1013L(N) = \left(\frac{N_c}{N}\right)^{\alpha_N}, \qquad \alpha_N \approx 0.076, \qquad N_c \approx 8.8 \times 10^{13}

พร้อม companion αD0.095\alpha_D \approx 0.095 สำหรับ data และ αCmin0.050\alpha_C^{\min} \approx 0.050 สำหรับ optimally allocated compute ค่าคงที่ไม่ใช่ universal และ paper ก็บอกเช่นนั้น: “the precise numerical values of NcN_c, CcminC_c^{\min} and DcD_c depend on the vocabulary size and tokenization and hence do not have a fundamental meaning.”

exponents เล็กมาก: parameters มากขึ้นสิบเท่าซื้อ factor 100.0761.1910^{0.076} \approx 1.19 ออกจาก loss ที่เหลือ ฟังเหมือนไม่มีอะไร และนี่คือข้อเท็จจริงสำคัญที่สุดตรงนี้ — ผลตอบแทนแย่มากและมันไม่เคยหยุด power law ที่มี exponent เล็กสัญญาว่า order of magnitude ถัดไปจะช่วย น้อยกว่าครั้งก่อน แต่ช่วยไปตลอด การซื้อ compute หยุดเป็นการพนันและกลายเป็นการซื้อที่มีอัตราแลกเปลี่ยนตีพิมพ์ไว้ ซึ่งตรงกับ argument ที่ปลดล็อกเงินทุน

จากนั้น prescription ก็ตามมา และนี่คือจุดที่ paper ผิดในแบบที่ทำให้อุตสาหกรรมเสียเงินจำนวนมาก Table 6 ของ Kaplan ให้ NoptC0.73N_{\text{opt}} \propto C^{0.73} และ DoptC0.27D_{\text{opt}} \propto C^{0.27}: compute มากขึ้นสิบเท่าหมายถึงโมเดลใหญ่ขึ้น 5.4 เท่า แต่ป้อน text มากขึ้นแค่ 1.9 เท่า abstract พูดชัด — “optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.” วงการทำตามนั้นจริง: GPT-3 มี 175 พันล้าน parameters บน 300 พันล้าน tokens,7 Gopher 280 พันล้านบน 300 พันล้าน, Megatron-Turing NLG 530 พันล้านบน 270 พันล้าน2 ครึ่ง token ถึงสอง tokens ต่อ parameter ทุกที่เหมือนกันหมด

Chinchilla และสิ่งที่ sweep ข้างบนกำลังวัด

ลิงก์ไปยังส่วน: Chinchilla และสิ่งที่ sweep ข้างบนกำลังวัด

ในเดือนมีนาคม 2022 Hoffmann et al. train โมเดลกว่า 400 ตัวตั้งแต่ 70 ล้านถึง 16 พันล้าน parameters และมาถึงข้อสรุปตรงกันข้ามผ่านสามเส้นทางอิสระ2 Table 2 ของพวกเขารายงาน exponent aa ใน NoptCaN_{\text{opt}} \propto C^{a} เป็น 0.50, 0.49 และ 0.46 เทียบกับ 0.73 ของ Kaplan พูดง่าย ๆ: model size และ training data ควรโตในสัดส่วนเท่ากัน

วิธีที่สองของพวกเขาคือวิธีที่ sweep ตอนต้นบทนี้ทำซ้ำ ใน scale หนึ่งในล้าน: fix budget, train หลาย sizes ที่ budget นั้นแบบเป๊ะ ๆ, plot final loss เทียบกับ model size

parametersC=1013C = 10^{13}C=3.16×1013C = 3.16 \times 10^{13}C=1014C = 10^{14}
98,6245.3531 (171)4.8638 (542)
150,3205.4636 (74)
194,2085.5041 (44)4.8730 (140)4.4040 (442)
295,8085.5550 (19)4.9029 (60)4.3383 (190)
665,2805.7849 (3.8)5.1254 (12)4.4192 (38)
1,280,7685.8174 (1.0)5.1751 (3.2)4.5003 (10)
3,101,5685.4686 (0.5)4.7768 (1.7)
5,315,0725.5894 (0.2)4.8514 (0.6)
15,053,5685.3534 (0.1)

Held-out loss เป็น nats ต่อ token, tokens ต่อ parameter ในวงเล็บ, ตัวหนาคือโมเดลที่ดีที่สุดในแต่ละ budget; dash คือจุดที่ไม่ได้ run เพราะ budget ต้องการ text มากกว่าที่ corpus มี หรือ size อยู่นอกช่วงที่ sweep ตรงนั้น

อ่านลงคอลัมน์: loss ลดลง แตะ bottom แล้วไต่ขึ้นอีก โมเดลใหญ่เกินไปสำหรับ budget ได้ง่ายพอ ๆ กับเล็กเกินไป — ที่ 101410^{14} โทษของการเลือก 665,280 parameters แทน 295,808 คือ 0.08 nats ซึ่งบน envelope ที่ fit ข้างบนคือ loss ที่โมเดลขนาดถูกต้องจะไปถึงด้วย compute น้อยกว่า 18% การเลือก shape ผิดทิ้ง budget ไปหนึ่งในห้า นั่นคือ Figure 3 ของ Chinchilla ในบ่ายเดียวบน GPU หนึ่งตัว แทนที่จะใช้โมเดลสี่ร้อยตัว

ตอนนี้อ่านข้ามแถว ที่ 101310^{13} โมเดลที่ดีที่สุดคือตัวเล็กสุดที่ sweep; ที่ 101410^{14} คือ 295,808 parameters ซึ่งมีจุดประกบทั้งสองด้าน optimum เลื่อนไปทางขวาเมื่อ budget โตขึ้น ซึ่งคือเนื้อหาทั้งหมดของการแก้ไข Fit วิธีที่สามของ paper — surface L(N,D)=E+A/Nα+B/DβL(N,D) = E + A/N^{\alpha} + B/D^{\beta} เหนือทุก run — แล้ว minimise ภายใต้ C=6NDC = 6ND:

TEXT
L(N, D) = 24.7 / N^0.195 + 46.8 / D^0.169       (E fits to ~0; see below)
implied   N_opt ∝ C^0.464
  compare   Chinchilla 0.46-0.50 · Besiroglu 0.513 · Kaplan 0.73

0.46 จากแล็ปท็อป เทียบกับ 0.73 ของ Kaplan การตรงกันสามหลักจาก fit สาม budget คือโชค; การตรงกันหลักแรกไม่ใช่ exponent เคลื่อนที่ได้ — constant ไม่ได้ เพราะ token-to-parameter ratio ที่ optima เหล่านี้คือ 170 ถึง 540 ไม่ใช่ 20 เหตุผลสามข้อ ล้วนให้บทเรียน EE fit ไปที่ศูนย์เพราะที่ loss สูงกว่า 4 nats run ยังห่างจาก entropy floor ที่ครอง fit ของ Chinchilla มาก Batch size และ learning rate ถูก fix แทนที่จะ tune ต่อจุด ซึ่งทำให้ runs ที่ได้ steps น้อยที่สุดเสียเปรียบ — และนั่นคือโมเดลใหญ่: ที่ 101310^{13} FLOPs โมเดล 1.28-million-parameter ได้ optimiser steps รวม 159 ครั้ง ต่ำกว่าหลายพันครั้งที่ term SminS_{\min} ของ Kaplan บอกว่าโมเดลใด ๆ ต้องการมาก scaling law ถูก fit ภายใน regime หนึ่ง และอันนี้อยู่ต่ำกว่า Chinchilla หก orders of magnitude

ดังนั้น abstract ของ paper: “current large language models are significantly undertrained” Chinchilla คือการสาธิต — 70 พันล้าน parameters บน 1.4 ล้านล้าน tokens, compute รวมเท่าเดิม กับ Gopher 280 พันล้านบน 300 พันล้าน เอาชนะมันใน 51 จาก 57 tasks ของ MMLU, 67.5% ต่อ 60%2 เล็กลงสี่เท่า text มากขึ้นสี่เท่าครึ่ง เงินเท่าเดิม โมเดลดีกว่า

ข้อควรระวังสองข้อเกี่ยวกับ ratio ดังกล่าว “ยี่สิบ tokens ต่อ parameter” ไม่ใช่ประโยคใน paper ซึ่งบอกเพียงว่า “for every doubling of model size the number of training tokens should also be doubled”; เลข 20 เป็น inference จาก Table 3 และจาก Chinchilla เองที่ 70 B บน 1.4 T และความแม่นยำของมันแย่กว่าที่ตีพิมพ์: Besiroglu et al. refit จาก digitisation ของ Figure 4 พบว่า parameters เดิม “fit the reconstructed data poorly” โดยมี intervals “implausibly tight given the number of data points” และวางช่วงที่ซื่อตรงไว้ที่ “between 4 and 40” tokens ต่อ parameter8

รายละเอียดหนึ่งของวิธีของ Chinchilla ชำระสัญญาที่ บทที่ 1 เคยให้ไว้เกี่ยวกับ learning-rate schedules cosine schedule ต้อง match กับ token budget โมเดลที่จะเห็น 10 ล้าน tokens ต้อง decay learning rate ไปศูนย์ที่ 10 ล้าน tokens; ให้ schedule ที่ออกแบบสำหรับ 100 ล้าน แล้วหยุดก่อนกำหนด คุณกำลังอ่าน loss กลางทางลงที่ rate สูงเกินไป Chinchilla train แต่ละโมเดลที่ cycle lengths สี่แบบเพื่อ control สิ่งนี้โดยเฉพาะ; sweep ข้างบนตั้ง schedule จาก budget ด้วยเหตุผลเดียวกัน

มันเป็นผลเชิงประจักษ์ที่มีประโยชน์ที่สุดในวงการ และถูกขายเกินจริงเป็นประจำ ข้อจำกัดสี่ข้อ

มัน predict loss ไม่ใช่ capability ฝั่งซ้ายคือ cross-entropy บน held-out text ไม่มีอะไรใน papers เหล่านี้อนุญาตให้ claim ว่าโมเดลจะเขียน SQL ถูกต้อง ปฏิเสธ request อันตราย หรือใช้ tool ได้ นี่คือบทเรียนของบทที่ 5 อีกครั้ง: prediction ของ loss ไม่ใช่ prediction ของ behaviour ที่คุณกำลังจ่ายเงินซื้อ

มันถูก fit ไม่ได้ถูก derive ไม่มีทฤษฎีใดผลิต αN=0.076\alpha_N = 0.076 ค่าคงที่เคลื่อนตาม tokenizer — ซึ่งเป็นเหตุผลที่ perplexity comparison ข้าม tokenizers สองตัวไม่มีความหมาย อย่างที่บทที่ 8 อธิบาย — และเคลื่อนตาม data mixture, architecture และ optimiser ทุก law ที่ตีพิมพ์คือ law ของ setup ที่ผลิตมัน ซึ่งเป็นเหตุผลที่ Meta refit ของตัวเองก่อน Llama 35

มัน assume token สดใหม่สำหรับทุก step ซึ่งแอบ assume corpus อนันต์ Muennighoff et al. วัดว่าจะเกิดอะไรเมื่อมันหมด: data ซ้ำถึงสี่ epochs แทบไม่มีค่าใช้จ่าย — โมเดล 8.7-billion-parameter บน 44 พันล้าน unique tokens ที่เห็นสี่ครั้งจบด้วย “only 0.5% higher validation loss” กว่าโมเดลเดียวกันบน unique tokens 178 พันล้าน — ขณะที่เลยประมาณสิบหก epochs ไปแล้ว compute เพิ่มไม่ซื้ออะไรเลย9

และไม่มีใคร train แบบ compute-optimal อีกแล้ว Chinchilla minimise cost ของ training; โมเดลที่ deploy แล้วจ่ายประมาณ 2N2N FLOPs ต่อ generated token ไปตลอด LLaMA 1 พูดตรง ๆ: “given a target level of performance, the preferred model is not the fastest to train but the fastest at inference”10 Sardana et al. formalise มันด้วยการ minimise 6NDtrain+2NDinference6ND_{\text{train}} + 2ND_{\text{inference}} แทน และพบว่าใครก็ตามที่คาดหวัง requests หนึ่งพันล้านควร train “smaller and longer than Chinchilla-optimal”11 §9.1 ของ Llama 3 เห็นด้วย: โมเดลเล็กของมัน train “far beyond the point of compute optimal training, effectively trading training compute for inference efficiency”5 ratio ไม่ได้ล้าสมัย; มันตอบคำถามที่ไม่ใช่คำถามที่กำลังถูกถามอีกต่อไป

Emergent abilities และข้อถกเถียงว่ามันจริงไหม

ลิงก์ไปยังส่วน: Emergent abilities และข้อถกเถียงว่ามันจริงไหม

Loss ลดลงอย่างราบรื่น Benchmark scores บางครั้งไม่เป็นแบบนั้น Wei et al. รวบรวมกรณีที่ task อยู่ระดับเดาสุ่มตลอดหลาย orders of magnitude ของ training compute แล้วกระโดด — arithmetic สามหลักปรากฏใน GPT-3 ที่ประมาณ 2×10222 \times 10^{22} FLOPs, MMLU สูงกว่า guessing ระหว่าง 33 และ 5×10235 \times 10^{23} — และตั้งชื่อ pattern นี้ว่า: “an ability is emergent if it is not present in smaller models but is present in larger models”12 หากนั่นเป็น property จริง การ extrapolate จากการทดลองราคาถูกก็ไม่ปลอดภัย เพราะ capability ที่คุณซื้ออาจไม่มีอยู่ใน scale ใดที่คุณจ่ายเพื่อทดสอบไหว

Schaeffer, Miranda และ Koyejo โต้แย้งว่าส่วนใหญ่มันเป็น artefact ของการวัด และ mechanism คือเลขคณิต13 Per-token loss ลดลงอย่างราบรื่น ดังนั้นความน่าจะเป็นที่ token หนึ่งจะถูกต้อง exp(L)\exp(-\mathcal{L}) ดีขึ้นทีละน้อย ให้คะแนนโมเดลด้วย exact string match บนคำตอบยาว LL-token แล้วคุณยกความน่าจะเป็นนั้นเป็นกำลัง LL — curve ราบรื่นที่ถูกยกกำลังมาก ๆ ดูเหมือนหน้าผา เปลี่ยนเป็น metric ที่นับ tokens แทนที่จะเรียกร้องให้ถูกทั้งหมด บน outputs เดิม แล้ว “the family's performance smoothly, continuously and predictably improves with increasing scale”

ตัวเลขจาก audit ของพวกเขาคือสิ่งที่ควรจำ — “of the 39 preferred metrics in BIG-Bench, at most 5 display emergence” โดยมี metrics ไม่ต่อเนื่องสองตัวคิดเป็นกว่า 92% ของกรณีที่ claim — และคำเตือนของพวกเขาก็ควรจำเช่นกัน: “nothing in this paper should be interpreted as claiming that large language models cannot display emergent abilities” การกระโดดใน chart คือหลักฐานเกี่ยวกับ metric จนกว่าจะพิสูจน์เป็นอย่างอื่น บทที่ 29 คือจุดที่เรื่องนี้กลายเป็นปัญหาของคุณ เพราะการเลือก metric แบบ hard-cutoff เป็นการตัดสินใจที่คุณจะทำโดยไม่รู้ตัว

Corpus คือส่วนของ pretraining run ที่ไม่มีสมการกำกับ และเป็นที่ที่การตัดสินใจสำคัญที่สุดจำนวนมากอยู่ วัตถุดิบคือ web crawl: archive เดือนสิงหาคม 2026 ของ Common Crawl มี “2.14 billion web pages or 360 TiB of uncompressed content” หนึ่งเดือนของมัน ดาวน์โหลดฟรี14 แทบไม่มีส่วนไหนใช้ได้ในสภาพเดิม Paper T5 บอกว่า crawl “largely comprises gibberish or boiler-plate text like menus, error messages, or duplicate text” และ pipeline C4 ที่มันแนะนำคือรายการ heuristics แบบทื่อ ๆ — เก็บเฉพาะบรรทัดที่ลงท้ายด้วย terminal punctuation, ทิ้ง pages ที่มีน้อยกว่าสามประโยค, ทิ้ง page ใด ๆ ที่มีวงเล็บปีกกาหรือคำจากรายการสาธารณะของคำหยาบ — เปลี่ยน text รายเดือนยี่สิบ terabytes ให้เหลือประมาณ 750 GB15

ทื่อคือคำที่ถูกต้อง Dodge et al. audit ว่า filters เหล่านั้นลบอะไรออกไป และพบว่า obscenity blocklist ลบ 42% ของ documents ใน African-American English และ 32% ใน Hispanic-aligned English เทียบกับ 6.2% ของ White-aligned English เหลือ corpus ที่ 97.8% เป็นหมวดสุดท้าย16 กฎที่ไม่มีความเห็นเกี่ยวกับ dialect กลับมีความเห็น

จากนั้น deduplication ซึ่งไม่ใช่งาน housekeeping: Lee et al. พบประโยค 61 คำที่ซ้ำ 61,036 ครั้งใน C4 และแสดงว่า deduplicating ลดอัตราที่โมเดล “emit memorized text” ลงสิบเท่า จาก 1.9% ของ generated tokens เป็น 0.19%17 แต่ more is not better — ทีม FineWeb deduplicate globally ข้าม 96 crawls ได้ 4 ล้านล้าน tokens และไม่มี gain ที่วัดได้ จากนั้น deduplicate แต่ละ crawl แยกกัน ได้ 20 ล้านล้าน และ match กับ corpus ที่ดีที่สุดที่มีอยู่18

จากนั้น contamination Llama 3 วัดของตัวเองและเผยแพร่: 98% ของ AGIEval, 95% ของ BIG-Bench Hard และ 85% ของ HellaSwag overlap กับ training set ด้วย 8-grams และสำหรับ MMLU overlap สูงจน “it is impossible to get a good performance gain estimate”5 §4 ของ GPT-3 รายงาน bug ใน filtering ที่ปล่อย benchmarks ไว้ใน data โดยไม่มีทางย้อนกลับ: “because of cost considerations it was infeasible to retrain the model”7

Provenance คือส่วนที่ยังไม่คลี่คลาย The Pile ส่ง component 100.96 GiB ชื่อ Books3 — 12% ของ corpus และตาม consent table ของ paper เอง คือหนังสือจาก private torrent tracker;19 มันถูกนำ offline ในเดือนสิงหาคม 2023 หลังมี copyright complaint สถานะทางกฎหมาย ณ กันยายน 2026 ยังไม่ settled และคำวินิจฉัยสหรัฐฯ สามคดีที่ถูกอ้างว่าเป็น trend ก็ไม่เห็นตรงกัน Alsup พบว่าการ train บนหนังสือที่ได้มาอย่างถูกกฎหมาย “exceedingly transformative” ขณะที่ถือว่า library ที่สร้างจากสำเนาละเมิดลิขสิทธิ์ไม่ใช่ และ Anthropic settled ครึ่งนั้นเป็นเงิน $1.5 billion ครอบคลุมงาน 482,460 ชิ้น หรือประมาณ $3,000 ต่อชิ้น อนุมัติ 20 กรกฎาคม 202620 Chhabria ให้ summary judgment แก่ Meta พร้อมเขียนว่าคำวินิจฉัยของเขา “does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful” เพียงแต่ว่า “these plaintiffs made the wrong arguments”21 Bibas ซึ่งตัดสิน against Ross Intelligence ระบุว่า “only non-generative AI is before me today”22 ยังไม่มี US appellate court ใดตัดสินคำถามนี้

คนทำส่วนที่ loss ทำไม่ได้ TIME รายงานในเดือนมกราคม 2023 ว่าคนงานที่ label toxic text ให้ OpenAI ผ่านบริษัท Sama รับกลับบ้าน “between around $1.32 and $2 per hour” เพื่ออ่านข้อความที่บรรยาย child sexual abuse, torture และ self-harm ขณะที่ OpenAI จ่าย Sama $12.50 ต่อชั่วโมงสำหรับงานนี้; Sama โต้แย้งทั้งช่วงค่าจ้างและ quota23 นั่นคือ filtering รอบ pretraining ไม่ใช่ pretraining เอง — แต่มันอยู่ใน invoice เดียวกัน และเป็นจุดที่มีคนคนหนึ่งนั่งอยู่

ไฟฟ้าเป็นของจริงและมักถูก quote ผิด ตัวเลขตีพิมพ์ที่ระมัดระวังที่สุดคือของ BLOOM: 1,082,990 GPU-hours, 433 MWh และ 24.7 tonnes ของ CO₂ equivalent สำหรับ run, 50.5 เมื่อนับ manufacturing และ idle nodes;24 Patterson et al. วาง GPT-3 ที่ 1,287 MWh และ 552 tonnes25 ข้อควรระวังสองข้อ ข้อได้เปรียบของ BLOOM คือ grid nuclear ฝรั่งเศสที่ 57 g CO₂ ต่อ kWh ไม่ใช่ efficiency — มันใช้พลังงาน มากกว่า OPT-175B และตัวเลข emissions ที่ถูก quote มากที่สุดของวงการ คือ 626,155 lb ของ Strubell et al. สำหรับ neural architecture search ภายหลังถูกแสดงว่า สูงเกินจริง 88 เท่า เพราะ assume ว่า search run ที่ full model size ทั้งที่ run บน proxy26 framing ของ LBNL คือแบบที่ปกป้องได้: US data centres ใช้ 192 TWh ในปี 2024, 4.7% ของไฟฟ้าทั้งประเทศ — ตัวเลขที่ผูกกับอุตสาหกรรม ไม่ใช่กับ run ใด run หนึ่ง27

เส้นที่ลากผ่านทั้งหมดคือสิ่งที่ Bender et al. ตั้งชื่อว่า documentation debt: “putting ourselves in a situation where the datasets are both undocumented and too large to document post hoc”28 ทุกข้อเท็จจริงข้างต้นมีอยู่เพราะมีใครบางคนไปดู สำหรับ corpora เบื้องหลังโมเดลที่คนส่วนใหญ่ใช้ ไม่มีใครทำได้

ตอนนี้คือเลขคณิตที่ทุกคนอยากได้ จาก inputs ที่ cite ไว้สี่รายการ เพื่อให้เมื่อมัน stale จะเห็นชัดว่าต้อง replace อะไร

หน้า H100 ของ NVIDIA ระบุ BF16 tensor-core throughput 1,979 teraFLOPS ใต้ footnote ที่เขียนว่า “with sparsity”29 ไม่มี pretraining run ใดใช้ structured sparsity ดังนั้นตัวเลข dense คือครึ่งหนึ่งของมัน: 989.5 TFLOP/s

Table 4 ของ Llama 3 รายงาน BF16 model FLOPs utilisation 38–43% ใช้ 40%: useful arithmetic 395.8 TFLOP/s ต่อ GPU5

ราคา on-demand ของ Lambda สำหรับ node 8×H100 SXM เข้าถึงเมื่อ 2026-09-06: $3.99 ต่อ GPU-hour ดังนั้น $31.92 ต่อชั่วโมงสำหรับ node30

ratio ของ Chinchilla, D=20ND = 20N, ให้ C=6ND=120N2C = 6ND = 120N^2 และดังนั้น N=C/120N = \sqrt{C/120}

budgetH100-hoursFLOPscompute-optimal paramstokensบน node 8×H100 หนึ่งตัวGPUs เพื่อจบใน 90 วัน
$100253.6e19546 M10.9 B3.1 h1
$1,0002513.6e201.73 B34.5 B31.3 h1
$10,0002,5063.6e215.46 B109 B13 days2
$100,00025,0633.6e2217.3 B345 B131 days12
$1,000,000250,6273.6e2354.6 B1.09 T4 years116
$10,000,0002,506,2663.6e24173 B3.45 T36 years1,160
$100,000,00025,062,6573.6e25546 B10.9 T358 years11,603

อ่านสองคอลัมน์สุดท้ายคู่กัน ที่ $10,000 คุณได้โมเดล 5-billion-parameter บน node เช่าหนึ่งตัวในสองสัปดาห์ ที่ $100,000,000 เลขคณิตบอกว่า 546 พันล้าน parameters — และ H100 สิบสองพันตัวที่ต่อสายเข้าด้วยกันเป็นเวลาสามเดือน ซึ่งไม่ใช่สิ่งที่คุณเช่าด้วยบัตรเครดิต เลยประมาณ $100,000 ไปแล้ว binding constraint หยุดเป็นเงินและกลายเป็น cluster

ก่อนเชื่อตารางแบบนั้น ให้ทดสอบกับ runs ที่มี real cost เผยแพร่ — llm.c reproduce GPT-2 124M ใน “~90 minutes” บน node 8×A100 “for about $20” และ GPT-2 1.6B ใน 24 ชั่วโมงบน node 8×H100 ในราคา $67231

TEXT
$672, against what $672 actually bought (llm.c GPT-2 1.6B, one 8xH100 node, 24 h)
  this table predicts:      168 H100-hours   N = 1.41 B params   D = 28.3 B tokens
  what was actually run:    192 H100-hours   N = 1.558 B params  D = 33.6 B tokens

Llama 3 405B, against Meta's own published GPU-hours
  from the paper's 3.8e25 FLOPs at 40 % MFU:   26.67 M H100-hours
  published in Meta's Llama 3.1 model card:    30.84 M H100-hours    ratio 0.86

ทั้งสองอยู่ภายในประมาณ 15% ซึ่งเป็นความแม่นยำที่ประมาณการแบบนี้ควรได้รับโดยประมาณ และดีกว่าความแม่นยำที่มันมักถูก quote อย่างมาก

การเปรียบเทียบ headline พร้อมทั้งสองนิยามบนโต๊ะ

ลิงก์ไปยังส่วน: การเปรียบเทียบ headline พร้อมทั้งสองนิยามบนโต๊ะ

ตัวเลขที่ถูกพูดซ้ำมากที่สุดในเรื่องนี้คือโมเดลระดับ GPT-2 ที่มีต้นทุนประมาณ $43,000 ในปี 2019 สามารถ reproduce ได้วันนี้ในราคาไม่กี่สิบดอลลาร์ ครึ่งสมัยใหม่มีเอกสารดี; ครึ่งประวัติศาสตร์ไม่มี

วันนี้ README ของ nanochat โดย Karpathy: “you can train your own GPT-2 capability LLM ... for only $48 (~2 hours of 8XH100 GPU node) ... On a spot instance, the total cost can be closer to ~$15.”32 “GPT-2 capability” ตรงนี้มีความหมาย precise และเผยแพร่แล้ว — เอาชนะ CORE score 0.256525 ของ GPT-2 — บน leaderboard ที่ entry ดีที่สุด ณ 14 มีนาคม 2026 คือ 1.65 ชั่วโมง ค่า $48 assume $3 ต่อ GPU-hour ต่ำกว่า list ของ Lambda ที่ $3.99; ที่ list จะใกล้ $64

ในปี 2019 ไม่มี primary source: OpenAI ไม่เคยเผยแพร่ duration หรือ cost chain คือ The Register, กุมภาพันธ์ 2019 รายงาน “256 Google TPU3 cores” โดยไม่มีราคาและไม่มี duration; จากนั้น Synced, มิถุนายน 2019 ระบุว่า hardware cost $256 ต่อชั่วโมงบน Google Cloud และเขียนชัดว่า “OpenAI didn't specify the training duration” $43,008 คือ $256 ต่อชั่วโมงคูณ 168 ชั่วโมงที่ assume ขึ้นมาโดยไม่มีใครเคย source

ดังนั้น headline ที่ซื่อตรงคือ: โมเดลที่ match benchmark score ที่ตีพิมพ์ของ GPT-2 สามารถ train ได้วันนี้ในราคาต่ำกว่า $100 มากบน hardware เช่า เทียบกับต้นทุนปี 2019 ที่ไม่เคยถูกเผยแพร่ และ estimate อันโด่งดังของมันตั้งอยู่บนการเดา duration แบบไม่มี source การพังทลายของต้นทุนเป็นเรื่องจริง และครึ่งสมัยใหม่ reproduce ได้โดยใครก็ตามที่มีบัตรเครดิต; ratio คือเลขคณิตบนตัวเลขที่ไม่มีอยู่ นั่นคือสภาพของ published training costs โดยทั่วไป Paper GPT-3 ไม่มีจำนวนเงินดอลลาร์เลย มีเพียง 3.14×10233.14 \times 10^{23} FLOPs ใน Table D.1;7 paper Llama 3 ก็ไม่มีเช่นกัน5 training cost ทุกตัวที่คุณเคยอ่านคือ estimate จาก FLOP count, hardware assumption และ price assumption — ควรถามเสมอว่าเป็นของใคร

Base model รู้อะไร และมันหยุดรู้เมื่อไหร่

ลิงก์ไปยังส่วน: Base model รู้อะไร และมันหยุดรู้เมื่อไหร่

สิ่งที่ออกมาได้เห็น corpus คงที่ที่ประกอบขึ้น ณ เวลาคงที่ และมี properties สองอย่างตามมา

อย่างแรกคือ knowledge cutoff หลังวันที่เก็บรวบรวม โมเดลไม่รู้อะไรเลย — ไม่ใช่ “ไม่แน่ใจ” แต่ ไม่มีอะไร — และมันจะ confabulate ได้อย่างลื่นไหลแทนที่จะบอกเช่นนั้น เพราะการบอกเช่นนั้นไม่เคยเป็น behaviour ที่มันถูก train ให้ทำ model card ของ Llama 3.1 ให้เดือนธันวาคม 2023;33 ทุกโมเดลมีหนึ่งค่า และมันเป็น property ของ training data ไม่ใช่ deployment การ workaround เรื่องนี้เป็น retrieval problem ซึ่งคือ บทที่ 19

อย่างที่สองคือ base model เติมต่อมากกว่าตอบ ให้ “เมืองหลวงของฝรั่งเศสคืออะไร?” กับมัน แล้ว continuation ที่เป็นไปได้คือคำถามอีกข้อ เพราะใน corpus string นั้นมักปรากฏในรายการแบบฝึกหัด

text completer ไม่ใช่ assistant มันไม่ทำตาม instructions เพราะไม่มีอะไรใน corpus บอกว่าคำขอควรถูกเชื่อฟังแทนที่จะถูกต่อ มันไม่มี notion ของ conversation ที่มีผู้เข้าร่วมสองฝ่าย มันจะสร้าง continuation ที่น่าจะเป็นที่สุดของ prompt อันตรายอย่างเต็มใจ เพราะ probable คือสิ่งเดียวที่มันเคยถูก optimise เพื่อให้ได้

การเปลี่ยนมันเป็นสิ่งที่ตอบได้ต้องใช้ stage ที่สองซึ่ง cost เพียงเศษเสี้ยวของหนึ่งเปอร์เซ็นต์ของ stage แรก และประกอบเกือบทั้งหมดจากการแสดง examples ของ behaviour ที่คุณต้องการ แล้วเปรียบเทียบ outputs ของมันเองเป็นคู่ ๆ stage นั้นคือที่มาของ instruction following, chat templates, refusals และ — สิ่งนี้ทำให้คนแปลกใจ — ความสามารถในการ call tool ทั้งหมด บทที่ 11 คือ stage นั้น: supervised fine-tuning, RLHF, DPO และ GRPO, และคำถามว่า “aligned” หมายถึงอะไรและใครเป็นคนตัดสิน


ควรอ่านคู่กับบทนี้ด้วย: build-nanogpt ของ Karpathy และ video ประกอบ ซึ่งพา reproduce GPT-2 แบบครบ end to end ด้วย pace ที่บทนี้ทำไม่ได้; และ Stanford CS324, Large Language Models, ซึ่ง lectures เรื่อง data และ environmental impact ลงลึกกว่าส่วนข้างบนใน material ที่คอร์สนี้แตะครั้งเดียวแล้วมอบหมายต่อ

  1. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D. Scaling Laws for Neural Language Models. arXiv:2001.08361 (2020). power laws สามชุดคือ Eqs. (1.1)–(1.3) ใน §1.2 และ constants ครบอยู่ใน Appendix A, Table 5; การอนุมาน 6N6N คือ §2.1; exponents ของ compute-allocation อยู่ใน Table 6 โปรดสังเกตว่ามี compute laws สองชุดคือ αC=0.057\alpha_C = 0.057 ที่ fixed batch size และ αCmin=0.050\alpha_C^{\min} = 0.050 ที่ optimal batch size; paper บอกว่าชุดหลัง “should be used to make predictions”. 2

  2. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E. et al. Training Compute-Optimal Large Language Models. arXiv:2203.15556 (2022). Exponents ใน Table 2, budgets ที่ project ใน Table 3, การเปรียบเทียบ Gopher ใน §4, convention การนับ parameters ใน Appendix F prose ใต้ Table 3 ไม่ตรงกับ Table 3 เองสำหรับแถว 175 B และ 280 B; table คือเวอร์ชันที่ควร quote. 2 3 4

  3. Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D. et al. Mixed Precision Training. arXiv:1710.03740 (2017), ICLR 2018. FP32 master weights ใน §3.1, loss scaling ใน §3.2. 2

  4. Rajbhandari, S., Rajbhandari, S., Ruwase, O. and He, Y. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 (2019), SC20. บัญชี 16Ψ16\Psi คือ §3.1; ตัวเลข residual-state สำหรับ activations คือ §3.2.

  5. Grattafiori, A. et al. (Llama Team, AI @ Meta). The Llama 3 Herd of Models. arXiv:2407.21783 (2024). Compute budget และ token count ใน §1, scaling law ที่ refit ใน §3.2.1, configuration ของ parallelism และ MFU ใน Table 4, contamination analysis ใน §5.1.4, statement เรื่อง over-training ใน §9.1 paper ไม่มี dollar figures และไม่มี emissions table. 2 3 4 5 6

  6. Stanford CS336, Language Modeling from Scratch. Lecture 2 ครอบคลุม resource accounting, lectures 5–8 ครอบคลุม GPUs, kernels และ parallelism, lectures 9 และ 11 ครอบคลุม scaling, lectures 13–14 ครอบคลุม data นี่คือคอร์สที่บทนี้มอบหมาย engineering ของมันต่อไป และเป็นสาธารณะ.

  7. Brown, T. B. et al. Language Models are Few-Shot Learners. arXiv:2005.14165 (2020). Compute ใน Appendix D, Table D.1 — ซึ่งมีคอลัมน์ที่ขึ้นหัวตรงตัวว่า “flops per param per token” และค่าของทุกแถว GPT-3 คือ 6 Contamination analysis ใน §4. 2 3

  8. Besiroglu, T., Erdil, E., Barnett, M. and You, J. Chinchilla Scaling: A replication attempt. arXiv:2404.10102 (2024). สร้าง data ของ Chinchilla ใหม่ด้วยการ digitise Figure 4, refit และรายงาน exponents ที่แก้แล้วกับ intervals ที่กว้างขึ้นมาก.

  9. Muennighoff, N., Rush, A. M., Barak, B., Le Scao, T., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T. and Raffel, C. Scaling Data-Constrained Language Models. arXiv:2305.16264 (2023), NeurIPS 2023. ผลลัพธ์ four-epoch อยู่ใน §6; half-life สิบหก-epoch คือ fitted RD15R_D^* \approx 15.

  10. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T. et al. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 (2023). §1 ระบุ inference-cost argument against Chinchilla-optimal training.

  11. Sardana, N., Portes, J., Doubov, S. and Frankle, J. Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws. arXiv:2401.00448 (2023), ICML 2024. §5 ของพวกเขายังมี counterweight: โมเดลที่ train ที่ token ratios สุดโต่งยังดีขึ้น แต่ “more slowly than scaling laws predict”.

  12. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S. et al. Emergent Abilities of Large Language Models. arXiv:2206.07682 (2022), TMLR. นิยามอยู่ใน §2, examples และ compute thresholds ใน §3–4 และ Table 1.

  13. Schaeffer, R., Miranda, B. and Koyejo, S. Are Emergent Abilities of Large Language Models a Mirage? arXiv:2304.15004 (2023), NeurIPS 2023 outstanding paper. argument เรื่อง metric คือ §2, meta-analysis ของ BIG-Bench คือ §4, vision example ที่สร้างขึ้นคือ §5.

  14. Common Crawl, August 2026 Crawl Archive Now Available (CC-MAIN-2026-34), เผยแพร่ 24 August 2026, เข้าถึง 2026-09-06 หน้าแรกของมันเอง claim ว่า “over 300 billion pages spanning 15 years”, “totalling more than 10 petabytes” — เป็นตัวเลขของ archive ทั้งหมด ไม่ใช่ monthly crawl ที่ตีราคาไว้ตรงนี้.

  15. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W. and Liu, P. J. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 (2019), JMLR 21(140). filters ของ C4 อยู่ใน §2.2 paper ให้ขนาดเป็น bytes ไม่ใช่ tokens; ตัวเลข 156-billion-token ที่มักถูก attribute ให้มันมาจาก Dodge et al. ข้างล่าง.

  16. Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M. and Gardner, M. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758 (2021), EMNLP 2021. อัตราการลบ dialect อยู่ใน §5.3; benchmark contamination ใน C4 อยู่ใน §4.2.

  17. Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C. and Carlini, N. Deduplicating Training Data Makes Language Models Better. arXiv:2107.06499 (2021), ACL 2022. การซ้ำ 61,036 ครั้งอยู่ใน footnote 1; ตัวเลข memorisation อยู่ใน §6.2, Table 4 และเป็นเปอร์เซ็นต์ของ generated tokens ภายใต้ criterion exact-match 50-token.

  18. Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L. and Wolf, T. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557 (2024), NeurIPS 2024 Datasets and Benchmarks. ผล deduplication อยู่ใน §3.4 dataset ที่ release โตเกิน 15 ล้านล้าน tokens ของ paper ไปแล้ว.

  19. Gao, L., Biderman, S., Black, S. et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv:2101.00027 (2020). Books3 อยู่ใน §2.3 และ Table 1; consent table คือ Table 5 corpus มี 825.18 GiB ดังนั้นแม้แต่ title ก็เป็นการปัดลง.

  20. Bartz v. Anthropic, No. 4:24-cv-05417 (N.D. Cal.). คำสั่ง fair-use 23 June 2025 (Dkt. 231); class certification 17 July 2025; final approval and judgment 20 July 2026 (Dkt. 680). Settlement ปล่อยเฉพาะ past inputs ไม่ใช่ outputs และไม่ใช่ future conduct.

  21. Kadrey v. Meta, No. 3:23-cv-03417-VC (N.D. Cal.), summary judgment 25 June 2025 (Dkt. 598). โปรดสังเกตว่า distribution claim เรื่อง torrenting ยังไม่ได้ตัดสินและยัง live อยู่.

  22. Thomson Reuters v. ROSS Intelligence, No. 1:20-cv-00613-SB (D. Del.), revised opinion 11 February 2025 (Dkt. 770), Bibas J. อยู่ระหว่าง interlocutory appeal ต่อ Third Circuit (No. 25-2153), argued 11 June 2026, ยังไม่ตัดสิน ณ เวลาที่เขียน.

  23. Perrigo, B. Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. TIME, 18 January 2023. $2 เป็น ceiling สำหรับ senior reviewers ที่ทำได้ทุก target; junior labellers ซึ่งเป็นส่วนใหญ่รับกลับบ้าน $1.32 คำโต้แย้งของ Sama ที่ quote ใน article เดียวกันให้ $1.46–$3.74 และ quota ต่ำกว่า.

  24. Luccioni, A. S., Viguier, S. and Ligozat, A.-L. Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. arXiv:2211.02001 (2022), JMLR 24(253). Tables 1 และ 3.

  25. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M. and Dean, J. Carbon Emissions and Large Neural Network Training. arXiv:2104.10350 (2021). ตัวเลขของ GPT-3 คือ Table 4; การแก้ estimate ของ NAS คือ §4.1.

  26. Strubell, E., Ganesh, A. and McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. arXiv:1906.02243 (2019), ACL 2019. ควรอ่านอย่างละเอียดเพราะสิ่งที่เกิดกับตัวเลขที่ถูก quote มากที่สุดของมัน: paper ระมัดระวัง ระบุ extrapolation ของตัวเอง และยังผิดสอง orders of magnitude ในบรรทัดเดียวที่ทุกคนพูดซ้ำ.

  27. Smith, S. J., Hubbard, A., Newkirk, A., Ganeshalingam, M., Holecek, B., Sartor, D., Mills, M. and Shehabi, A. United States Data Center Energy Usage Report: 2025 Update. LBNL-2001758 (18 June 2026). สิ่งนี้ revise historical series ที่ถูก cite กว้างขวางในรายงาน 2024 ลง; ถ้าคุณ quote ตัวเลข 176 TWh สำหรับปี 2023 คุณกำลัง quote edition ที่ถูกแทนที่แล้ว.

  28. Bender, E. M., Gebru, T., McMillan-Major, A. and Shmitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? FAccT '21, pp. 610–623. DOI 10.1145/3442188.3445922. “Documentation debt” คือ §4.4 โปรดสังเกตว่า carbon figures ของ paper เอง cite จาก Strubell et al. และรับ correction ข้างต้นต่อมา — ซึ่งเป็น illustration ของ argument ของมัน มากกว่าจะเป็น refutation.

  29. NVIDIA. NVIDIA H100 Tensor Core GPU product page, nvidia.com/en-us/data-center/h100/ (เข้าถึง 2026-09-06). ทุกแถว tensor-core บนหน้านั้นยกเว้น FP64 มี footnote “with sparsity”; ตัวเลข dense BF16 ที่ใช้ที่นี่คือครึ่งหนึ่งของ 1,979 TFLOPS ที่เผยแพร่.

  30. Lambda. GPU Cloud pricing, lambda.ai/pricing (เข้าถึง 2026-09-06). On-demand, ต่อ GPU ต่อชั่วโมง, ก่อนภาษี ราคาในส่วนนี้จะ stale เร็วกว่าอย่างอื่นในคอร์สนี้; เลขคณิตรอบ ๆ มันจะไม่ stale.

  31. Karpathy, A. karpathy/llm.c, discussion #481, Reproducing GPT-2 (124M) in llm.c in 90 minutes for $20 (28 May 2024), และ discussion #677, Let's reproduce GPT-2 (1.6B): one 8XH100 node, 24 hours, $672, in llm.c (11 July 2024).

  32. Karpathy, A. karpathy/nanochat, README และ leaderboard “time to GPT-2” (เข้าถึง 2026-09-06). ตัวเลข $48 และนิยาม CORE-score ของ “GPT-2 capability” อยู่ใน README ทั้งคู่; speedrun.sh ของ repository เองบอกว่า “approximately 1.5 hours” ดังนั้นให้ถือเลขสองชั่วโมงเป็นการปัดเศษ.

  33. Meta. Llama 3.1 model card, models/llama3_1/MODEL_CARD.md ใน meta-llama/llama-models (เข้าถึง 2026-09-06). แหล่งของ 30.84 M H100-hours สำหรับโมเดล 405 B, total 39.3 M, ตัวเลข location-based 11,390 tCO2eq และ data cutoff เดือนธันวาคม 2023.


สร้างโดย

David Vicente Campos

ผู้ก่อตั้ง NeuraLIA Labs และผู้ร่วมก่อตั้ง MyRealFood

ผมเป็นวิศวกรคอมพิวเตอร์ที่จบจากมหาวิทยาลัยเลออน ผมร่วมก่อตั้ง MyRealFood ที่ที่ผมในฐานะ CTO ได้สร้างแอปซึ่งผู้คนหลายล้านคนใช้เพื่อกินให้ดีขึ้น และผมก่อตั้ง NeuraLIA Labs ที่ที่ผมสร้างผลิตภัณฑ์ AI ที่นี่ผมเขียนถึงสิ่งที่ผมต้องทำความเข้าใจระหว่างทาง ในแบบที่ผมเคยหวังว่าจะมีใครสักคนอธิบายให้ผมฟัง

เพิ่มเติมเกี่ยวกับผู้เขียน

เผยแพร่โดย NeuraLIA Labs

รับโพสต์ใหม่ในกล่องจดหมาย

ข่าว AI คู่มือ และอัปเดตผลิตภัณฑ์ — อีเมลสั้น ๆ เมื่อเรามีสิ่งที่คุ้มเวลาของคุณ

ชอบแบบข้อความมากกว่าไหม รับเนื้อหาเดียวกันได้ที่นี่:คอมมูนิตี้ WhatsApp (เปิดในแท็บใหม่)ช่อง Telegram (เปิดในแท็บใหม่)

ดัชนีคอร์ส

Abstract software decision engine with branching paths, probability nodes, and glowing gates.
jevอ่าน 5 นาที

โมเดล AI Jev สร้างมาเพื่อการตัดสินใจ ไม่ใช่การเขียนความเรียง

Jev ของ TypeSafe AI กำลังได้รับความสนใจ เพราะมองความฉลาดของซอฟต์แวร์เป็นปัญหาความน่าจะเป็น: เลือกกิ่งที่ถูกต้อง แนบความมั่นใจ และหลีกเลี่ยงการจ่ายเงินให้ LLM เขียนข้อความเมื่อโค้ดต้องการการตัดสินใจ

Abstract legal research workspace with documents, search nodes and governance controls.
openaiอ่าน 4 นาที

Astra for Law ของ OpenAI คือระบบ AI ด้านกฎหมาย ไม่ใช่โมเดลใหม่

การเปิดตัวด้านกฎหมายของ OpenAI ไม่ได้เน้นโมเดลฐานรากใหม่เท่ากับระบบที่ล้อมรอบโมเดลนั้น: การค้นคืนเฉพาะโดเมน เครื่องมือที่เชื่อถือได้ สิทธิ์ เบนช์มาร์ก และเส้นทางการตรวจทาน

Abstract agent runtime sorting documents, memory blocks and pointer nodes inside a bounded context frame.
context-engineeringอ่าน 4 นาที

วิศวกรรมบริบทสำหรับเอเจนต์ AI ที่ทำงานระยะยาว

เอเจนต์ที่ทำงานต่อเนื่องไม่ได้ล้มเหลวเพียงเพราะหน้าต่างบริบทเล็กเกินไป แต่ล้มเหลวเมื่อไฟล์ ผลลัพธ์จากเครื่องมือ และประวัติที่ค้างเก่าบดบังงานที่เอเจนต์ควรทำให้เสร็จ

พร้อมให้ LIA เลือกโมเดลให้แล้วหรือยัง?

สร้างงานด้วยโมเดล AI ทุกตัวในที่เดียว เริ่มฟรีวันนี้