LLM کی pretraining: data، compute، scaling laws اور لاگت
ایک laptop GPU پر بیس models سے scaling law ناپی گئی، اور 6ND compute تخمینہ حقیقی FLOP counter سے جانچا گیا۔
اس صفحے پر
باب 9 ایک ایسے transformer block پر ختم ہوا جو train ہوتا ہے۔ ان میں سے چند کو stack کریں، باب 8 کے next-token loss کو output کی طرف موڑ دیں، اور ایجاد کرنے کو کچھ نہیں بچتا۔ جو کچھ باقی رہتا ہے وہ خریداری ہے۔
یہ بات جتنی سنائی دیتی ہے اس سے بڑی تبدیلی ہے۔ اب تک ہر باب نے پوچھا تھا کیا یہ سیکھتا ہے؟ — ایک ہاں یا نہیں والا سوال جسے laptop دس منٹ میں نمٹا دیتا ہے۔ یہ باب ایک ایسا سوال پوچھتا ہے جس میں پیسہ شامل ہے: arithmetic کی ایک fixed مقدار کے ساتھ، میں بہترین model کون سا خرید سکتا ہوں؟ جواب ایک formula ہے، اور 2018 میں یہ کسی کے لیے بھی واضح نہیں تھا۔
یہ رہا وہ سوال، ایک laptop GPU پر measurement سے جواب دیا ہوا۔ 98,624 سے 15 million parameters تک بیس models، Wikipedia کے 174 million tokens پر scratch سے train کیے گئے — 2,048-token BPE vocabulary اسی طرح train کی گئی جیسے باب 7 ایک train کرتا ہے، Chapter 9 کا transformer۔ ہر run کو compute کے تین budgets میں سے بالکل ایک ملا اور ایک operation بھی زیادہ نہیں، اس لیے بڑا model لازماً کم text پڑھتا ہے۔ ہر budget پر پہنچنے والا بہترین held-out loss:
budget C (FLOPs) best loss reached by a model of
1.00e13 5.3531 98,624 params
3.16e13 4.8638 98,624 params
1.00e14 4.3383 295,808 params
fitted: L = (Cc / C)^0.0913 over one decade of computeدس گنا arithmetic loss کو 19 % کم کر دیتی ہے، اور تینوں points log-log میں ایک سیدھی line پر آتے ہیں۔ پہلے نو ابواب میں کوئی چیز اس کی پیش گوئی نہیں کرتی۔ اس کے پیچھے کوئی theorem نہیں — یہ ایک empirical regularity ہے، جو اس laptop اور ایک datacentre کے درمیان magnitude کے دس orders پر ایک مختلف exponent کے ساتھ برقرار رہتی ہے، اور یہی وہ واحد observation ہے جس نے ایک industry کو GPUs پر ایک چھوٹے ملک کے GDP جتنا خرچ کرنے پر آمادہ کیا۔
pretraining کیا ہے، اور اس میں نیا کیا ہے
اس حصے کا لنک: pretraining کیا ہے، اور اس میں نیا کیا ہےObjective میں کچھ نہیں بدلتا۔ model اب بھی اگلا token predict کرتا ہے، loss اب بھی باب 4 کی cross-entropy ہے جو Chapter 8 کی factorisation پر apply ہوتی ہے، optimiser اب بھی باب 6 کا AdamW ہے۔ Pretraining کوئی نیا algorithm نہیں؛ یہ وہی algorithm ہے جو اتنے بڑے corpus پر چلایا جاتا ہے کہ run کا budget بنانا پڑتا ہے۔ دو چیزیں اسے ممکن بناتی ہیں: labels free ہیں، کیونکہ position کے لیے target پر token ہے اور پہلے ہی text میں موجود ہے؛ اور Chapter 6 کے آخری section نے اعتراض ہٹا دیا، کیونکہ classical rules سے کہیں زیادہ parameters والا model ٹوٹتا نہیں، بہتر ہوتا ہے۔ نتیجے میں ایک base model نکلتا ہے — ایسی چیز جو answer دینے کے بجائے text جاری رکھتی ہے۔
خرچ کرنے سے پہلے compute گننا: 6ND
اس حصے کا لنک: خرچ کرنے سے پہلے compute گننا: 6NDاس میں سے کسی چیز کا budget بنانے سے پہلے اسے گننا ضروری ہے، اور field اسے ایک formula سے گنتی ہے:
جہاں parameter count ہے، training tokens ہیں اور کل floating-point operations ہیں۔ Kaplan et al. اسے دو steps میں derive کرتے ہیں۔1 Forward: 2 FLOPs per parameter per token، کیونکہ matrix multiply میں ہر parameter ہر token کے لیے ایک بار استعمال ہوتا ہے، ایک multiply اور ایک add میں۔ Backward: forward کا دو گنا، کیونکہ باب 5 کا backward pass ہر layer پر دو gradients compute کرتا ہے — layer کے inputs کے حوالے سے، تاکہ signal سفر جاری رکھے، اور اس کے weights کے حوالے سے — دونوں forward کے سائز کے matrix multiply ہیں، اس لیے ۔
یہ پوری derivation ہے، اور اسے ماننے کے بجائے check کرنا چاہیے۔ PyTorch ایک حقیقی FLOP counter، torch.utils.flop_counter.FlopCounterMode، کے ساتھ آتا ہے، جو model کی dispatch کی ہوئی ہر operation intercept کرتا ہے اور actual work جمع کرتا ہے۔ اسے magnitude کے چار orders پر چلائیں، سب سے بڑا meta device پر، جو shapes allocate کرتا ہے memory نہیں:
from torch.utils.flop_counter import FlopCounterMode
counter = FlopCounterMode(display=False)
with counter:
loss = model(x, targets)[1]
loss.backward()
measured = counter.get_total_flops()
print(measured / (6 * n_params * n_tokens))| configuration | without embeddings | total | measured, fwd+bwd | ÷ (total ) | ÷ (no emb.) | fwd+bwd ÷ fwd |
|---|---|---|---|---|---|---|
| 128, 4 layers, 256 | 788,736 | 7,254,400 | 4.60e10 | 1.031 | 9.485 | 3.000 |
| 512, 8 layers, 256 | 25,183,232 | 51,045,888 | 3.26e11 | 1.038 | 2.104 | 3.000 |
| 768, 12 layers, 1024 | 84,973,056 | 124,356,864 | 1.75e12 | 1.145 | 1.676 | 3.000 |
| 1600, 48 layers, 1024 | 1,474,870,400 | 1,556,920,000 | 2.10e13 | 1.100 | 1.161 | 3.000 |
| 4096, 32 layers, 2048 | 6,442,983,424 | 6,582,444,032 | 8.74e13 | 1.080 | 1.104 | 3.000 |
| 8192, 80 layers, 8192 | 64,427,147,264 | 65,544,929,280 | 3.75e15 | 1.163 | 1.183 | 3.000 |
forward+backward over forward ratio 3.000 ہے، ہر scale پر بالکل: یہ کوئی approximation نہیں جو اتفاقاً اچھی ہو، بلکہ اوپر والی arithmetic identity ہے جو ایک ایسے counter سے round number کی صورت واپس آئی ہے جسے derivation کا کچھ علم نہیں۔
Measured total پھر سے 3 % سے 17 % اوپر بیٹھتا ہے، جب embedding matrices کو گنتا ہے — اور یہ clause اہم ہے، کیونکہ دو founding papers کو مختلف طرح گنتے ہیں۔ Kaplan "all vocabulary and positional embeddings" کو exclude کرتا ہے کیونکہ ایسا کرنے سے "significantly cleaner scaling laws" بنتے ہیں (§1.3)؛ Chinchilla کا Appendix F کہتا ہے "we also count embeddings matrices in the total parameter count"۔2 wide vocabulary اور narrow hidden dimension کے لیے دونوں میں factor nine کا فرق ہے، جیسا پہلی row دکھاتی ہے۔
باقی gap وہ ہے جسے جان بوجھ کر چھوڑتا ہے: attention scores۔ Kaplan کی Eq. (2.2) forward cost کو لکھتی ہے اور دوسرا term چھوڑ دیتی ہے کیونکہ — 2020 میں safe، اب کم safe، اور یہی وجہ ہے کہ بڑھنے پر ratio اوپر drift کرتا ہے — اسی لیے یہاں دو rows کا 1,024 مشترک ہے اور ratio گرتا ہے، 1.145 سے 1.100 تک، جب 768 سے 1,600 ہوتا ہے۔ یہ وہ cost ہے جسے Chapter 9 نے introduce کیا اور Chapter 16 price میں بدلتا ہے۔
Memory: اصل میں fit کیا ہونا چاہیے
اس حصے کا لنک: Memory: اصل میں fit کیا ہونا چاہیےCompute فیصلہ کرتا ہے کہ run میں کتنا وقت لگے گا؛ memory فیصلہ کرتی ہے کہ یہ start بھی ہو سکتا ہے یا نہیں۔ Plain fp32 AdamW سے train کریں تو ہر parameter چار numbers اٹھاتا ہے: weight، اس کا gradient، اور Adam کا running mean اور variance — وہ دو averages جو Chapter 6 میں ہاتھ سے بنائے گئے۔ چار numbers، ہر ایک چار bytes، یعنی ایک activation سے پہلے 16 bytes per parameter۔ 8 GB laptop GPU پر measured، step کے اس point پر resident allocation لیتے ہوئے جہاں کوئی graph alive نہیں:
| model | vocabulary | batch | predicted | resident measured | peak in a step | the difference | |
|---|---|---|---|---|---|---|---|
| 512, 8 layers | 50,257 | 8 | 51,045,888 | 779 MB | 801 MB | 2,500 MB | 1,699 MB |
| 512, 8 layers | 4,096 | 8 | 27,411,456 | 418 MB | 426 MB | 1,043 MB | 617 MB |
| 256, 6 layers | 4,096 | 8 | 5,839,360 | 89 MB | 89 MB | 382 MB | 293 MB |
| 256, 6 layers | 4,096 | 32 | 5,839,360 | 89 MB | 89 MB | 1,259 MB | 1,170 MB |
| 256, 6 layers | 4,096 | 128 | 5,839,360 | 89 MB | 89 MB | 4,771 MB | 4,681 MB |
Prediction اور measurement 3 % کے اندر agree کرتے ہیں۔ حیرت آخری column ہے: activations model کو dwarf کر دیتے ہیں۔ وہی 5.8-million-parameter model جسے persistent state کے لیے 89 MB چاہیے، 128 کے batch پر activations کے لیے 4,681 MB چاہتا ہے — model سے باون گنا — اور اس کا بڑا حصہ transformer بھی نہیں۔ یہ logits ہیں، ہر token کے لیے vocabulary size کا ایک vector، ہر entry چار bytes: آخری row میں 512 MB، پہلی میں 393 MB۔ Vocabulary size Chapter 7 میں چنی گئی تھی، اور یہ اب بھی فیصلہ کر رہی ہے کہ card پر کیا fit ہوتا ہے۔
کون سا term dominate کرتا ہے، یہ run کی shape پر depend کرتا ہے، اسی لیے Micikevicius et al. کہتے ہیں memory "is dominated by activations"3 جبکہ ZeRO کہتا ہے 1.5-billion-parameter model کو صرف model states کے لیے "at least 24 GB" چاہیے۔4 ZeRO ایک اور راستے سے اسی 16 bytes تک پہنچتا ہے — fp16 weights کے لیے ، fp16 gradients کے لیے ، fp32 master weights اور Adam کے دو moments میں سے ہر ایک کے لیے — جو 70 billion parameters کے لیے 1.12 terabytes ہے، ایک activation سے پہلے fourteen 80 GB GPUs کے برابر۔
Parallelism، ایک paragraph اور ایک delegation میں
اس حصے کا لنک: Parallelism، ایک paragraph اور ایک delegation میںFrontier scale پر اس میں سے کچھ بھی ایک device پر fit نہیں ہوتا، اس لیے run ایک ساتھ چار ways میں split ہوتا ہے۔ Data parallelism ہر GPU پر model کی copy رکھتا ہے اور gradients average کرتا ہے — default، اور وہی جسے ZeRO optimiser state کی redundant copies رکھنے سے انکار کر کے improve کرتا ہے۔ Tensor parallelism individual matrices کو devices کے across split کرتا ہے۔ Pipeline parallelism ہر device کو layers کا contiguous group دیتا ہے۔ Context parallelism sequence کو ہی split کرتا ہے، جو صرف تب ضروری ہے جب اتنا long ہو جائے کہ attention term dominate کرے۔ Llama 3 کی Table 4 چاروں کو ایک ساتھ list کرتی ہے: tensor 8، context up to 16، pipeline 16، data up to 128، 16,384 H100 GPUs کے across۔5 یہ course اس بارے میں بس اتنا ہی کہے گا؛ distributed training engineering اپنے آپ میں ایک semester ہے، اور Stanford کا CS336 وہ semester ہے، lectures 5 سے 8، code کے ساتھ۔6 Delegation کے بعد جو چیز باقی رہتی ہے وہ ایک number ہے، model FLOPs utilisation — GPU کی peak arithmetic کا وہ fraction جو real run achieve کرتا ہے — یہی tidy کو wall-clock time اور پھر money میں بدلتا ہے۔
Kaplan، اور industry کی لگائی ہوئی شرط
اس حصے کا لنک: Kaplan، اور industry کی لگائی ہوئی شرطجنوری 2020 میں، Kaplan et al. نے transformers کا ایک grid train کیا اور پایا کہ test loss تینوں resources میں سے ہر ایک کے ساتھ چھ orders of magnitude سے زیادہ پر power law follow کرتا ہے۔1 ان کا §1.2 تین fitted laws دیتا ہے:
ساتھ میں data کے لیے اور optimally allocated compute کے لیے ۔ Constants universal نہیں، اور paper خود کہتا ہے: "the precise numerical values of , and depend on the vocabulary size and tokenization and hence do not have a fundamental meaning."
Exponents بہت چھوٹے ہیں: دس گنا parameters remaining loss سے factor کم کرواتے ہیں۔ یہ کچھ بھی نہیں لگتا، اور یہی یہاں سب سے اہم fact ہے — returns terrible ہیں اور وہ کبھی رکتی نہیں۔ small exponent والا power law وعدہ کرتا ہے کہ اگلا order of magnitude مدد کرے گا، پچھلے سے کم، ہمیشہ۔ Compute خریدنا gamble نہیں رہتا بلکہ published exchange rate کے ساتھ purchase بن جاتا ہے، اور یہی argument capital کو unlock کرنے والا تھا۔
پھر prescription آئی، اور یہاں paper اس طرح غلط تھا جس نے industry کو بہت پیسہ لگوا دیا۔ Kaplan کی Table 6 اور دیتی ہے: دس گنا compute کا مطلب 5.4 times bigger model جسے صرف 1.9 times اتنا text کھلایا گیا۔ Abstract explicit ہے — "optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence." Field نے بالکل یہی کیا: GPT-3، 300 billion tokens پر 175 billion parameters،7 Gopher 300 billion پر 280 billion، Megatron-Turing NLG 270 billion پر 530 billion۔2 آدھے token سے دو tokens per parameter، across the board۔
Chinchilla، اور اوپر والا sweep کیا measure کر رہا تھا
اس حصے کا لنک: Chinchilla، اور اوپر والا sweep کیا measure کر رہا تھامارچ 2022 میں، Hoffmann et al. نے 70 million سے 16 billion parameters تک 400 سے زیادہ models train کیے اور تین independent routes سے opposite conclusion تک پہنچے۔2 ان کی Table 2 میں exponent کو 0.50، 0.49 اور 0.46 report کرتی ہے، Kaplan کے 0.73 کے مقابل۔ سادہ لفظوں میں: model size اور training data کو equal proportion میں grow کرنا چاہیے۔
ان کا second approach وہی ہے جو اس chapter کے شروع میں sweep نے scale کے millionth حصے پر reproduce کیا: budget fix کریں، اسی budget پر بہت سے sizes train کریں، final loss کو model size کے against plot کریں۔
| parameters | |||
|---|---|---|---|
| 98,624 | 5.3531 (171) | 4.8638 (542) | — |
| 150,320 | 5.4636 (74) | — | — |
| 194,208 | 5.5041 (44) | 4.8730 (140) | 4.4040 (442) |
| 295,808 | 5.5550 (19) | 4.9029 (60) | 4.3383 (190) |
| 665,280 | 5.7849 (3.8) | 5.1254 (12) | 4.4192 (38) |
| 1,280,768 | 5.8174 (1.0) | 5.1751 (3.2) | 4.5003 (10) |
| 3,101,568 | — | 5.4686 (0.5) | 4.7768 (1.7) |
| 5,315,072 | — | 5.5894 (0.2) | 4.8514 (0.6) |
| 15,053,568 | — | — | 5.3534 (0.1) |
Held-out loss nats per token میں، brackets میں tokens per parameter، ہر budget پر best model bold؛ dash وہ point ہے جو run نہیں کیا گیا، کیونکہ budget corpus سے زیادہ text demand کرتا تھا یا size اس sweep کے باہر تھا۔
Column کے نیچے پڑھیں: loss گرتا ہے، bottom out ہوتا ہے اور پھر چڑھتا ہے۔ کوئی model اپنے budget کے لیے اتنا ہی آسانی سے بہت بڑا ہو سکتا ہے جتنا بہت چھوٹا — پر 295,808 کی جگہ 665,280 parameters چننے کی penalty 0.08 nats ہے، جو اوپر fitted envelope پر وہ loss ہے جو correctly sized model 18 % کم compute کے ساتھ پہنچتا ہے۔ غلط shape چننا budget کا پانچواں حصہ ضائع کر دیتا ہے۔ یہ Chinchilla کی Figure 3 ہے، four hundred models کے بجائے ایک GPU پر ایک afternoon میں۔
اب across پڑھیں۔ پر best model swept میں سب سے چھوٹا ہے؛ پر یہ 295,808 parameters ہے، دونوں طرف bracketed۔ Budget بڑھنے پر optimum دائیں move کرتا ہے، یہی correction کا پورا content ہے۔ Paper کا third approach fit کریں — ہر run پر surface — اور کے subject to minimise کریں:
L(N, D) = 24.7 / N^0.195 + 46.8 / D^0.169 (E fits to ~0; see below)
implied N_opt ∝ C^0.464
compare Chinchilla 0.46-0.50 · Besiroglu 0.513 · Kaplan 0.730.46، ایک laptop سے، Kaplan کے 0.73 کے مقابل۔ Three-budget fit سے تین digits تک agreement luck ہے؛ پہلے digit تک agreement نہیں۔ Exponent travel کرتا ہے — constant نہیں، کیونکہ ان optima پر token-to-parameter ratio 170 سے 540 ہے، 20 نہیں۔ تین reasons، سب instructive۔ zero پر fit ہوتا ہے کیونکہ 4 nats سے اوپر loss پر run entropy floor کے قریب کہیں نہیں جو Chinchilla کے fit کو dominate کرتا ہے۔ Batch size اور learning rate کو per point tune کرنے کے بجائے fixed رکھا گیا، جو ان runs کو handicaps کرتا ہے جنہیں fewest steps ملتے ہیں — اور وہ large models ہیں: FLOPs پر 1.28-million-parameter model کو total 159 optimiser steps ملتے ہیں، Kaplan کے term کے مطابق کسی بھی model کو درکار چند ہزار سے بہت کم۔ Scaling law ایک regime کے اندر fit ہوتا ہے، اور یہ one Chinchilla سے چھ orders of magnitude نیچے بیٹھتا ہے۔
اسی لیے paper کا abstract: "current large language models are significantly undertrained"۔ Chinchilla demonstration ہے — 1.4 trillion tokens پر 70 billion parameters، Gopher کے 300 billion پر 280 billion جیسا same total compute، مگر 57 میں سے 51 MMLU tasks پر بہتر، 60 % کے مقابل 67.5 %۔2 چار گنا چھوٹا، ساڑھے چار گنا زیادہ text، same money، better model۔
اس famous ratio پر دو caveats۔ "Twenty tokens per parameter" paper میں کوئی sentence نہیں؛ paper صرف کہتا ہے کہ "for every doubling of model size the number of training tokens should also be doubled"؛ 20 Table 3 اور Chinchilla کے اپنے 1.4 T پر 70 B سے inference ہے۔ اور اس کی precision published سے worse ہے: Besiroglu et al. نے Figure 4 کی digitisation سے refit کیا، پایا کہ original parameters "fit the reconstructed data poorly" اور intervals "implausibly tight given the number of data points" تھے، اور honest range "between 4 and 40" tokens per parameter رکھی۔8
Chinchilla کے method کی ایک detail باب 1 کے learning-rate schedules کے promise کو پورا کرتی ہے۔ Cosine schedule کو token budget کے ساتھ match کرنا پڑتا ہے۔ جو model 10 million tokens دیکھے گا اسے اپنی learning rate 10 million tokens پر zero تک decay کرنی چاہیے؛ اسے 100 million کے لیے sized schedule دیں، early stop کریں، اور آپ descent کے بیچ کا loss پڑھ رہے ہیں ایسی rate پر جو بہت high ہے۔ Chinchilla ہر model کو four cycle lengths پر train کرتا ہے تاکہ بالکل اس کو control کرے؛ اوپر والا sweep اسی وجہ سے schedule کو budget سے set کرتا ہے۔
Scaling laws کیا promise نہیں کرتیں
اس حصے کا لنک: Scaling laws کیا promise نہیں کرتیںیہ field کا سب سے useful empirical result ہیں اور routinely oversold ہیں۔ چار limits۔
یہ loss predict کرتی ہیں، capability نہیں۔ left-hand side held-out text پر cross-entropy ہے۔ ان papers میں کوئی چیز یہ claim کرنے کا license نہیں دیتی کہ model correct SQL لکھے گا، harmful request refuse کرے گا یا tool use کرے گا۔ یہ Chapter 5 کا lesson دوبارہ ہے: loss کی prediction اس behaviour کی prediction نہیں جس کی آپ payment کر رہے ہیں۔
یہ fitted ہیں، derived نہیں۔ کوئی theory produce نہیں کرتی۔ Constants tokenizer کے ساتھ move کرتے ہیں — اسی لیے دو tokenizers کے across perplexity comparison meaningless ہے، جیسا Chapter 8 نے explain کیا — اور data mixture، architecture اور optimiser کے ساتھ بھی۔ ہر published law اس setup کی law ہے جس نے اسے produce کیا، اسی لیے Meta نے Llama 3 سے پہلے اپنی law refit کی۔5
یہ ہر step کے لیے fresh token assume کرتی ہیں، جو خاموشی سے infinite corpus assume کرتا ہے۔ Muennighoff et al. نے measure کیا کہ جب یہ ختم ہو جائے تو کیا ہوتا ہے: repeated data کے چار epochs تک تقریباً کچھ cost نہیں — 44 billion unique tokens پر 8.7-billion-parameter model نے انہیں چار بار دیکھ کر 178 billion unique tokens والے same model سے "only 0.5 % higher validation loss" پر finish کیا — جبکہ تقریباً sixteen epochs کے بعد additional compute کچھ نہیں خریدتا۔9
اور اب کوئی compute-optimal train نہیں کرتا۔ Chinchilla training کی cost minimise کرتا ہے؛ deployed model پھر ہر generated token کے لیے تقریباً FLOPs pay کرتا ہے، ہمیشہ۔ LLaMA 1 نے صاف کہا: "given a target level of performance, the preferred model is not the fastest to train but the fastest at inference"۔10 Sardana et al. نے کو instead minimise کر کے formalise کیا، اور پایا کہ جو بھی billion requests expect کرتا ہے اسے "smaller and longer than Chinchilla-optimal" train کرنا چاہیے۔11 Llama 3 کا §9.1 agree کرتا ہے: اس کے small models "far beyond the point of compute optimal training, effectively trading training compute for inference efficiency" train ہوتے ہیں۔5 Ratio obsolete نہیں؛ یہ ایک ایسے question کا جواب دیتا ہے جو اب پوچھا نہیں جا رہا۔
Emergent abilities، اور یہ بحث کہ آیا وہ real ہیں
اس حصے کا لنک: Emergent abilities، اور یہ بحث کہ آیا وہ real ہیںLoss smoothly گرتا ہے۔ Benchmark scores کبھی کبھی نہیں۔ Wei et al. نے ایسے cases collect کیے جہاں task training compute کے orders of magnitude تک chance پر بیٹھا رہتا ہے اور پھر jump کرتا ہے — GPT-3 میں تقریباً FLOPs پر تین-digit arithmetic ظاہر ہونا، MMLU کا اور کے درمیان guessing سے اوپر جانا — اور pattern کو name دیا: "an ability is emergent if it is not present in smaller models but is present in larger models"۔12 اگر یہ real property ہے تو cheap experiments سے extrapolate کرنا unsafe ہے، کیونکہ وہ capability جسے آپ خرید رہے ہیں شاید کسی ایسے scale پر موجود ہی نہ ہو جسے آپ afford کر کے test کر سکیں۔
Schaeffer, Miranda and Koyejo نے argue کیا کہ اس کا بڑا حصہ measurement کا artefact ہے، اور mechanism arithmetic ہے۔13 Per-token loss smoothly گرتا ہے، اس لیے ایک token کے right ہونے کی probability، ، gradually improve ہوتی ہے۔ Model کو -token answer پر exact string match سے score کریں تو آپ اس probability کو power تک raise کرتے ہیں — ایک smooth curve کو large power تک raise کریں تو وہ cliff جیسی لگتی ہے۔ Same outputs پر metric بدل کر ایسا metric لگائیں جو تمام tokens demand کرنے کے بجائے tokens count کرے، تو "the family's performance smoothly, continuously and predictably improves with increasing scale"۔
ان کا audit یاد رکھنے والا number ہے — "of the 39 preferred metrics in BIG-Bench, at most 5 display emergence"، جبکہ دو discontinuous metrics claimed cases کے 92 % سے زیادہ کی ذمہ دار ہیں — اور ان کی caution بھی: "nothing in this paper should be interpreted as claiming that large language models cannot display emergent abilities"۔ Chart میں jump metric کے بارے میں evidence ہے جب تک دوسری صورت ثابت نہ ہو۔ Chapter 29 وہ جگہ ہے جہاں یہ آپ کا مسئلہ بنتا ہے، کیونکہ hard-cutoff metric چننا ایک ایسا decision ہے جو آپ محسوس کیے بغیر کریں گے۔
Data کہاں سے آتا ہے
اس حصے کا لنک: Data کہاں سے آتا ہےCorpus pretraining run کا وہ حصہ ہے جس کے ساتھ کوئی equation attached نہیں، اور جہاں اکثر consequential decisions رہتے ہیں۔ Raw material web crawl ہے: Common Crawl کے August 2026 archive میں "2.14 billion web pages or 360 TiB of uncompressed content" ہیں، اس کا ایک month، download کرنے کے لیے free۔14 تقریباً کچھ بھی اپنی original حالت میں usable نہیں۔ T5 paper کہتا ہے crawl "largely comprises gibberish or boiler-plate text like menus, error messages, or duplicate text"، اور اس نے جو C4 pipeline introduce کی وہ blunt heuristics کی list ہے — صرف terminal punctuation پر ختم ہونے والی lines رکھیں، تین sentences سے کم pages drop کریں، curly brace یا public obscenity list کے word والی ہر page drop کریں — monthly text کے twenty terabytes کو تقریباً 750 GB میں بدلتے ہوئے۔15
Blunt صحیح لفظ ہے۔ Dodge et al. نے audit کیا کہ یہ filters کیا remove کرتے ہیں اور پایا کہ obscenity blocklist African-American English میں 42 % documents اور Hispanic-aligned English میں 32 % delete کرتی ہے، White-aligned English کے 6.2 % کے مقابل، جس سے corpus 97.8 % آخری category کا رہ جاتا ہے۔16 Dialect کے بارے میں no opinion رکھنے والی rule کی ایک opinion نکل آئی۔
پھر deduplication، جو housekeeping نہیں: Lee et al. نے C4 میں 61-word sentence کو 61,036 بار repeated پایا، اور دکھایا کہ deduplicating models کے "emit memorized text" کرنے کی rate دس گنا کم کرتی ہے، generated tokens کے 1.9 % سے 0.19 % تک۔17 مگر more always better نہیں — FineWeb کی team نے 96 crawls کے across globally deduplicate کیا، 4 trillion tokens ملے اور کوئی measurable gain نہیں، پھر ہر crawl کو separately deduplicate کیا، 20 trillion ملے، اور best existing corpus کو match کیا۔18
پھر contamination۔ Llama 3 نے اپنی contamination measure کر کے publish کی: AGIEval کا 98 %، BIG-Bench Hard کا 95 % اور HellaSwag کا 85 % training set سے 8-grams کے ذریعے overlapping، اور MMLU کے لیے overlap اتنا high کہ "it is impossible to get a good performance gain estimate"۔5 GPT-3 کا §4 ایک filtering bug report کرتا ہے جس نے benchmarks کو data میں چھوڑ دیا اور واپسی کا راستہ نہیں تھا: "because of cost considerations it was infeasible to retrain the model"۔7
Provenance unresolved part ہے۔ The Pile نے Books3 نام کا 100.96 GiB component ship کیا — corpus کا 12 % اور، paper کی اپنی consent table کے مطابق، private torrent tracker کی books؛19 اسے August 2023 میں copyright complaint کے بعد offline کر دیا گیا۔ September 2026 تک legal position unsettled ہے، اور trend کے طور پر cite کی گئی تین US rulings ایک دوسرے سے disagree کرتی ہیں۔ Alsup نے lawfully acquired books پر training کو "exceedingly transformative" پایا جبکہ pirated copies سے بنی library کو نہیں، اور Anthropic نے اس half کو $1.5 billion پر settle کیا، 482,460 works کے لیے، تقریباً $3,000 each، 20 July 2026 کو approved۔20 Chhabria نے Meta کو summary judgment دیا مگر لکھا کہ اس کی ruling "does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful"، صرف یہ کہ "these plaintiffs made the wrong arguments"۔21 Bibas نے Ross Intelligence کے خلاف ruling دیتے ہوئے note کیا کہ "only non-generative AI is before me today"۔22 کسی US appellate court نے اس question پر ruling نہیں دی۔
لوگ وہ parts کرتے ہیں جو loss نہیں کر سکتا۔ TIME نے January 2023 میں report کیا کہ OpenAI کے لیے Sama firm کے through toxic text label کرنے والے workers "between around $1.32 and $2 per hour" گھر لے جاتے تھے، child sexual abuse، torture اور self-harm بیان کرنے والے passages پڑھتے ہوئے، جبکہ OpenAI نے اس work کے لیے Sama کو $12.50 an hour pay کیا؛ Sama pay range اور quota دونوں dispute کرتا ہے۔23 یہ pretraining کے گرد filtering ہے، pretraining خود نہیں — مگر same invoice پر ہے، اور وہیں ایک انسان بیٹھتا ہے۔
Electricity real ہے اور عموماً غلط quote ہوتی ہے۔ سب سے careful published figure BLOOM کی ہے: run کے لیے 1,082,990 GPU-hours، 433 MWh اور 24.7 tonnes CO₂ equivalent، manufacturing اور idle nodes count کرنے پر 50.5؛24 Patterson et al. GPT-3 کو 1,287 MWh اور 552 tonnes پر رکھتے ہیں۔25 دو cautions۔ BLOOM کا advantage efficiency نہیں بلکہ 57 g CO₂ per kWh والا French nuclear grid ہے — اس نے OPT-175B سے زیادہ energy use کی۔ اور field کی سب سے quoted emissions figure، Strubell et al. کی neural architecture search کے لیے 626,155 lb، بعد میں 88 times too high نکلی، کیونکہ اس نے assume کیا تھا کہ search full model size پر run ہوئی جبکہ وہ proxy پر ہوئی تھی۔26 LBNL کی framing defensible ہے: US data centres نے 2024 میں 192 TWh use کیے، national electricity کا 4.7 % — یہ number ایک industry سے attached ہے، کسی ایک run سے نہیں۔27
Through-line وہ ہے جسے Bender et al. نے documentation debt کہا: "putting ourselves in a situation where the datasets are both undocumented and too large to document post hoc"۔28 اوپر ہر fact اس لیے موجود ہے کہ کسی نے دیکھا۔ ان corpora کے لیے جن کے پیچھے most people use models ہیں، کوئی نہیں دیکھ سکتا۔
اصل لاگت کیا ہے
اس حصے کا لنک: اصل لاگت کیا ہےاب وہ arithmetic جو ہر کوئی چاہتا ہے، چار cited inputs سے، تاکہ جب وہ stale ہوں تو واضح ہو کہ کس کو replace کرنا ہے۔
Peak throughput
اس حصے کا لنک: Peak throughputNVIDIA کا H100 page BF16 tensor-core throughput کے 1,979 teraFLOPS list کرتا ہے ایک footnote کے تحت جو "with sparsity" پڑھتا ہے۔29 کوئی pretraining run structured sparsity use نہیں کرتا، اس لیے dense figure اس کا half ہے: 989.5 TFLOP/s۔
Utilisation
اس حصے کا لنک: UtilisationLlama 3 کی Table 4 38–43 % BF16 model FLOPs utilisation report کرتی ہے۔ 40 % لیں: per GPU useful arithmetic کے 395.8 TFLOP/s۔5
Price
اس حصے کا لنک: PriceLambda کی on-demand price ایک 8×H100 SXM node کے لیے، accessed 2026-09-06: $3.99 per GPU-hour، یعنی node کے لیے $31.92 an hour۔30
Shape
اس حصے کا لنک: ShapeChinchilla کا ratio، ، دیتا ہے اور اس لیے ۔
| budget | H100-hours | FLOPs | compute-optimal params | tokens | on one 8×H100 node | GPUs to finish in 90 days |
|---|---|---|---|---|---|---|
| $100 | 25 | 3.6e19 | 546 M | 10.9 B | 3.1 h | 1 |
| $1,000 | 251 | 3.6e20 | 1.73 B | 34.5 B | 31.3 h | 1 |
| $10,000 | 2,506 | 3.6e21 | 5.46 B | 109 B | 13 days | 2 |
| $100,000 | 25,063 | 3.6e22 | 17.3 B | 345 B | 131 days | 12 |
| $1,000,000 | 250,627 | 3.6e23 | 54.6 B | 1.09 T | 4 years | 116 |
| $10,000,000 | 2,506,266 | 3.6e24 | 173 B | 3.45 T | 36 years | 1,160 |
| $100,000,000 | 25,062,657 | 3.6e25 | 546 B | 10.9 T | 358 years | 11,603 |
آخری دو columns کو ساتھ پڑھیں۔ $10,000 پر آپ کو ایک rented node پر fortnight میں 5-billion-parameter model ملتا ہے۔ $100,000,000 پر arithmetic 546 billion parameters کہتی ہے — اور تین ماہ کے لیے twelve thousand H100s wired together، جو ایسی چیز نہیں جسے credit card سے rent کیا جائے۔ تقریباً $100,000 کے بعد binding constraint money نہیں رہتا بلکہ cluster بن جاتا ہے۔
ایسی table پر trust کرنے سے پہلے اسے ان runs کے خلاف test کریں جن کی real cost published ہے — llm.c GPT-2 124M کو 8×A100 node پر "~90 minutes" میں "for about $20" reproduce کرتا ہے، اور GPT-2 1.6B کو 8×H100 node پر 24 hours میں $672 میں۔31
$672, against what $672 actually bought (llm.c GPT-2 1.6B, one 8xH100 node, 24 h)
this table predicts: 168 H100-hours N = 1.41 B params D = 28.3 B tokens
what was actually run: 192 H100-hours N = 1.558 B params D = 33.6 B tokens
Llama 3 405B, against Meta's own published GPU-hours
from the paper's 3.8e25 FLOPs at 40 % MFU: 26.67 M H100-hours
published in Meta's Llama 3.1 model card: 30.84 M H100-hours ratio 0.86دونوں about 15 % کے اندر ہیں، جو تقریباً اتنی accuracy ہے جتنی اس kind کے estimate deserve کرتے ہیں، اور اس accuracy سے کافی بہتر ہے جس کے ساتھ اسے عموماً quote کیا جاتا ہے۔
Headline comparison، دونوں definitions table پر
اس حصے کا لنک: Headline comparison، دونوں definitions table پراس subject میں سب سے repeated figure یہ ہے کہ 2019 میں تقریباً $43,000 cost کرنے والا GPT-2-class model آج چند tens of dollars میں reproduce ہو سکتا ہے۔ Modern half اچھی طرح documented ہے؛ historical half نہیں۔
Today. Karpathy کے nanochat README: "you can train your own GPT-2 capability LLM ... for only $48 (~2 hours of 8XH100 GPU node) ... On a spot instance, the total cost can be closer to ~$15."32 یہاں "GPT-2 capability" precise اور published ہے — GPT-2 کے CORE score 0.256525 کو beat کرنا — ایک leaderboard پر جس کی best entry 14 March 2026 تک 1.65 hours ہے۔ $48، $3 per GPU-hour assume کرتا ہے، Lambda کی list $3.99 سے کم؛ list پر یہ $64 کے قریب ہے۔
In 2019. کوئی primary source نہیں: OpenAI نے duration یا cost کبھی publish نہیں کی۔ Chain The Register، February 2019 سے چلتی ہے، جس نے "256 Google TPU3 cores" report کیا، price یا duration کے بغیر؛ پھر Synced، June 2019، جس نے note کیا کہ hardware Google Cloud پر $256 an hour cost کرتا تھا اور explicitly کہا کہ "OpenAI didn't specify the training duration"۔ $43,008، $256 an hour کو assumed 168 hours سے multiply کرنا ہے جس کا source کبھی نہیں ملا۔
تو honest headline یہ ہے: GPT-2 کے published benchmark score کو match کرنے والا model آج rented hardware پر $100 سے کہیں کم میں train ہو سکتا ہے، ایک 2019 cost کے مقابل جو کبھی publish نہیں ہوئی اور جس کا famous estimate duration کے بارے میں ایک unsourced guess پر ٹکا ہے۔ Collapse real ہے اور modern half credit card رکھنے والے ہر شخص کے لیے reproducible ہے؛ ratio ایسے number پر arithmetic ہے جو exist نہیں کرتا۔ Published training costs کی state عموماً یہی ہے۔ GPT-3 paper میں dollar amount بالکل نہیں، صرف Table D.1 میں FLOPs ہیں؛7 Llama 3 paper میں بھی کوئی نہیں۔5 آپ نے جو بھی training cost پڑھی ہے وہ FLOP count، hardware assumption اور price assumption سے estimate ہے — ہمیشہ پوچھنے کے قابل کہ کس کا۔
Base model کیا جانتا ہے، اور کب جاننا بند کیا
اس حصے کا لنک: Base model کیا جانتا ہے، اور کب جاننا بند کیاجو نکلتا ہے اس نے ایک fixed moment پر assembled fixed corpus دیکھا ہے، اور دو properties follow کرتی ہیں۔
پہلی knowledge cutoff ہے۔ Collection date کے بعد model کچھ نہیں جانتا — "uncertain ہے" نہیں، کچھ نہیں — اور یہ اسے ماننے کے بجائے fluently confabulate کرے گا، کیونکہ ایسا کہنا کبھی وہ behaviour نہیں تھا جس پر یہ train ہوا ہو۔ Llama 3.1 کا model card December 2023 دیتا ہے؛33 ہر model کا ایک ہوتا ہے، اور یہ training data کی property ہے، deployment کی نہیں۔ اس کے around work کرنا retrieval problem ہے، جو Chapter 19 ہے۔
دوسری یہ کہ base model answer دینے کے بجائے complete کرتا ہے۔ اسے "What is the capital of France?" دیں تو ایک plausible continuation ایک اور question ہے، کیونکہ corpus میں یہ string اکثر exercises کی list میں آتی ہے۔
آگے کیا آتا ہے
اس حصے کا لنک: آگے کیا آتا ہےText completer assistant نہیں۔ یہ instructions follow نہیں کرتا، کیونکہ corpus میں کچھ بھی اسے نہیں بتاتا تھا کہ request کو obey کرنا چاہیے، continue نہیں۔ اسے two participants والی conversation کا کوئی notion نہیں۔ یہ harmful prompt کی most probable continuation خوشی سے produce کرے گا، کیونکہ probable ہی وہ واحد چیز ہے جس کے لیے اسے optimise کیا گیا۔
اسے answer دینے والی چیز میں بدلنا ایک second stage لیتا ہے جس کی cost first کے ایک per cent کے fraction جتنی ہوتی ہے، اور جو تقریباً مکمل طور پر اسے مطلوب behaviour کی examples دکھانے اور پھر اس کے اپنے outputs کے pairs compare کرنے پر مشتمل ہے۔ یہی stage ہے جہاں سے instruction following، chat templates، refusals اور — یہ لوگوں کو حیران کرتا ہے — tool call کرنے کی ability آتی ہے۔ Chapter 11 وہ stage ہے: supervised fine-tuning، RLHF، DPO اور GRPO، اور یہ question کہ "aligned" کا مطلب کیا ہے اور decide کون کرتا ہے۔
Sources and method
اس حصے کا لنک: Sources and methodاس chapter کے ساتھ ساتھ پڑھنے کے لائق: Karpathy کا build-nanogpt اور اس کی accompanying video، جو full GPT-2 reproduction کو end to end ایسی pace پر walk کرتے ہیں جو یہ chapter نہیں کر سکتا؛ اور Stanford CS324، Large Language Models، جس کے data اور environmental impact پر lectures اوپر والے section سے زیادہ depth میں جاتے ہیں، اس material میں جسے یہ course ایک بار treat کر کے delegate کرتا ہے۔
حوالہ جات
اس حصے کا لنک: حوالہ جات-
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D. Scaling Laws for Neural Language Models. arXiv:2001.08361 (2020). تین power laws §1.2 میں Eqs. (1.1)–(1.3) ہیں اور full constants Appendix A, Table 5 میں؛ derivation §2.1 ہے؛ compute-allocation exponents Table 6 ہیں۔ Note کریں کہ دو compute laws ہیں، fixed batch size پر اور optimal batch size پر ؛ paper کہتا ہے latter "should be used to make predictions"۔ ↩ ↩2
-
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E. et al. Training Compute-Optimal Large Language Models. arXiv:2203.15556 (2022). Exponents Table 2 میں، projected budgets Table 3 میں، Gopher comparison §4 میں، parameter-counting convention Appendix F میں۔ Table 3 کے نیچے prose 175 B اور 280 B rows کے لیے Table 3 itself سے disagree کرتی ہے؛ quote کرنے والا version table ہے۔ ↩ ↩2 ↩3 ↩4
-
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D. et al. Mixed Precision Training. arXiv:1710.03740 (2017), ICLR 2018. FP32 master weights §3.1 میں، loss scaling §3.2 میں۔ ↩ ↩2
-
Rajbhandari, S., Rajbhandari, S., Ruwase, O. and He, Y. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 (2019), SC20. accounting §3.1 ہے؛ activations کے residual-state figures §3.2 ہیں۔ ↩
-
Grattafiori, A. et al. (Llama Team, AI @ Meta). The Llama 3 Herd of Models. arXiv:2407.21783 (2024). Compute budget اور token count §1 میں، refitted scaling law §3.2.1 میں، parallelism configuration اور MFU Table 4 میں، contamination analysis §5.1.4 میں، over-training statement §9.1 میں۔ Paper میں dollar figures یا emissions table نہیں۔ ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
Stanford CS336, Language Modeling from Scratch. Lecture 2 resource accounting cover کرتا ہے، lectures 5–8 GPUs، kernels اور parallelism، lectures 9 اور 11 scaling، lectures 13–14 data۔ یہ وہ course ہے جسے یہ chapter اپنی engineering delegate کرتا ہے، اور یہ public ہے۔ ↩
-
Brown, T. B. et al. Language Models are Few-Shot Learners. arXiv:2005.14165 (2020). Compute Appendix D, Table D.1 میں — جس میں ایک column literal طور پر "flops per param per token" headed ہے، جس کی ہر GPT-3 row کے لیے value 6 ہے۔ Contamination analysis §4 میں۔ ↩ ↩2 ↩3
-
Besiroglu, T., Erdil, E., Barnett, M. and You, J. Chinchilla Scaling: A replication attempt. arXiv:2404.10102 (2024). Chinchilla کا data اس کی Figure 4 digitise کر کے reconstruct کرتا ہے، refit کرتا ہے، اور corrected exponents اور بہت wider intervals report کرتا ہے۔ ↩
-
Muennighoff, N., Rush, A. M., Barak, B., Le Scao, T., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T. and Raffel, C. Scaling Data-Constrained Language Models. arXiv:2305.16264 (2023), NeurIPS 2023. Four-epoch result §6 ہے؛ sixteen-epoch half-life fitted ہے۔ ↩
-
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T. et al. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 (2023). §1 Chinchilla-optimal training کے خلاف inference-cost argument بیان کرتا ہے۔ ↩
-
Sardana, N., Portes, J., Doubov, S. and Frankle, J. Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws. arXiv:2401.00448 (2023), ICML 2024. ان کے §5 میں counterweight بھی ہے: extreme token ratios پر trained models improve کرتے رہتے ہیں، مگر "more slowly than scaling laws predict"۔ ↩
-
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S. et al. Emergent Abilities of Large Language Models. arXiv:2206.07682 (2022), TMLR. Definition §2 میں، examples اور compute thresholds §3–4 اور Table 1 میں۔ ↩
-
Schaeffer, R., Miranda, B. and Koyejo, S. Are Emergent Abilities of Large Language Models a Mirage? arXiv:2304.15004 (2023), NeurIPS 2023 outstanding paper. Metric argument §2 ہے، BIG-Bench meta-analysis §4، constructed vision example §5۔ ↩
-
Common Crawl, August 2026 Crawl Archive Now Available (CC-MAIN-2026-34), published 24 August 2026, accessed 2026-09-06. اس کا اپنا front page "over 300 billion pages spanning 15 years"، "totalling more than 10 petabytes" claim کرتا ہے — یہ پورے archive کی figure ہے، یہاں priced monthly crawl کی نہیں۔ ↩
-
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W. and Liu, P. J. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 (2019), JMLR 21(140). C4 filters §2.2 ہیں۔ Paper sizes bytes میں دیتا ہے، tokens میں نہیں؛ اس سے widely attributed 156-billion-token figure نیچے Dodge et al. سے ہے۔ ↩
-
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M. and Gardner, M. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758 (2021), EMNLP 2021. Dialect removal rates §5.3 ہیں؛ C4 میں benchmark contamination §4.2 ہے۔ ↩
-
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C. and Carlini, N. Deduplicating Training Data Makes Language Models Better. arXiv:2107.06499 (2021), ACL 2022. 61,036 repeats footnote 1 ہیں؛ memorisation figures §6.2، Table 4 میں ہیں، اور 50-token exact-match criterion کے تحت generated tokens کی percentages ہیں۔ ↩
-
Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L. and Wolf, T. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557 (2024), NeurIPS 2024 Datasets and Benchmarks. Deduplication result §3.4 ہے۔ Released dataset تب سے paper کے 15 trillion tokens سے آگے بڑھ چکا ہے۔ ↩
-
Gao, L., Biderman, S., Black, S. et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv:2101.00027 (2020). Books3 §2.3 اور Table 1 ہے؛ consent table Table 5 ہے۔ Corpus 825.18 GiB ہے، اس لیے title بھی round-down ہے۔ ↩
-
Bartz v. Anthropic, No. 4:24-cv-05417 (N.D. Cal.). Fair-use order 23 June 2025 (Dkt. 231); class certification 17 July 2025; final approval and judgment 20 July 2026 (Dkt. 680). Settlement صرف past inputs release کرتا ہے، outputs نہیں اور future conduct نہیں۔ ↩
-
Kadrey v. Meta, No. 3:23-cv-03417-VC (N.D. Cal.), summary judgment 25 June 2025 (Dkt. 598). Note کریں کہ torrenting پر distribution claim decide نہیں ہوا اور live رہتا ہے۔ ↩
-
Thomson Reuters v. ROSS Intelligence, No. 1:20-cv-00613-SB (D. Del.), revised opinion 11 February 2025 (Dkt. 770), Bibas J. Third Circuit کو interlocutory appeal پر (No. 25-2153), argued 11 June 2026، writing کے وقت undecided۔ ↩
-
Perrigo, B. Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. TIME, 18 January 2023. $2 senior reviewers کے لیے ceiling ہے جنہوں نے ہر target meet کیا؛ junior labellers، اکثریت، $1.32 گھر لے گئے۔ Sama کا rebuttal، same article میں quoted، $1.46–$3.74 اور lower quota دیتا ہے۔ ↩
-
Luccioni, A. S., Viguier, S. and Ligozat, A.-L. Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. arXiv:2211.02001 (2022), JMLR 24(253). Tables 1 اور 3۔ ↩
-
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M. and Dean, J. Carbon Emissions and Large Neural Network Training. arXiv:2104.10350 (2021). GPT-3 کی figures Table 4 ہیں؛ NAS estimate کی correction §4.1 ہے۔ ↩
-
Strubell, E., Ganesh, A. and McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. arXiv:1906.02243 (2019), ACL 2019. بالکل اسی وجہ سے پڑھنے کے قابل کہ اس کے سب سے quoted number کے ساتھ کیا ہوا: paper careful ہے، اپنی extrapolation state کرتا ہے، اور پھر بھی اس one line پر دو orders of magnitude غلط تھا جسے سب نے repeat کیا۔ ↩
-
Smith, S. J., Hubbard, A., Newkirk, A., Ganeshalingam, M., Holecek, B., Sartor, D., Mills, M. and Shehabi, A. United States Data Center Energy Usage Report: 2025 Update. LBNL-2001758 (18 June 2026). یہ historical series کے لیے widely cited 2024 report کو downward revise کرتا ہے؛ اگر آپ 2023 کے لیے 176 TWh figure quote کر رہے ہیں تو superseded edition quote کر رہے ہیں۔ ↩
-
Bender, E. M., Gebru, T., McMillan-Major, A. and Shmitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? FAccT '21, pp. 610–623. DOI 10.1145/3442188.3445922. "Documentation debt" §4.4 ہے۔ Note کریں کہ paper کی اپنی carbon figures Strubell et al. سے cited ہیں اور اوپر والی correction inherit کرتی ہیں — جو اس کے argument کی illustration ہے، refutation نہیں۔ ↩
-
NVIDIA. NVIDIA H100 Tensor Core GPU product page,
nvidia.com/en-us/data-center/h100/(accessed 2026-09-06). اس page پر FP64 کے علاوہ ہر tensor-core row footnote "with sparsity" رکھتی ہے؛ یہاں استعمال ہونے والی dense BF16 figure published 1,979 TFLOPS کا half ہے۔ ↩ -
Lambda. GPU Cloud pricing,
lambda.ai/pricing(accessed 2026-09-06). On-demand، per GPU per hour، tax سے پہلے۔ اس section میں prices اس course کی کسی بھی چیز سے تیزی سے stale ہوں گی؛ ان کے around arithmetic نہیں۔ ↩ -
Karpathy, A.
karpathy/llm.c, discussion #481, Reproducing GPT-2 (124M) in llm.c in 90 minutes for $20 (28 May 2024), اور discussion #677, Let's reproduce GPT-2 (1.6B): one 8XH100 node, 24 hours, $672, in llm.c (11 July 2024). ↩ -
Karpathy, A.
karpathy/nanochat, README and "time to GPT-2" leaderboard (accessed 2026-09-06). $48 figure اور "GPT-2 capability" کی CORE-score definition دونوں README میں ہیں؛ repository کا اپناspeedrun.shکہتا ہے "approximately 1.5 hours"، اس لیے two-hour figure کو rounded سمجھیں۔ ↩ -
Meta. Llama 3.1 model card,
models/llama3_1/MODEL_CARD.mdinmeta-llama/llama-models(accessed 2026-09-06). 405 B model کے 30.84 M H100-hours، 39.3 M total، 11,390 tCO2eq location-based figure، اور December 2023 data cutoff کا source۔ ↩