सामग्री पर जाएँ
10/30अध्याय 10 / 30

LLM की pretraining: डेटा, compute, scaling laws और लागत

एक laptop GPU पर 20 models से scaling law मापा गया और 6ND compute अनुमान को वास्तविक FLOP counter से जाँचा गया।

इस पेज पर

Chapter 9 एक ऐसे transformer block पर खत्म हुआ था जो train होता है। ऐसे कुछ blocks को stack करें, Chapter 8 के next-token loss को output पर लगाएँ, और invent करने के लिए कुछ नहीं बचता। जो भी बचता है, वह खरीदारी है.

यह सुनने से बड़ा बदलाव है। अब तक हर chapter ने पूछा था क्या यह सीखता है? — एक yes-or-no सवाल जिसे laptop दस मिनट में निपटा देता है। यह chapter ऐसा सवाल पूछता है जिसमें पैसा शामिल है: arithmetic की fixed मात्रा देखते हुए, सबसे अच्छा model कौन-सा है जिसे मैं खरीद सकता हूँ? जवाब एक formula है, और 2018 में यह किसी के लिए भी obvious नहीं था।

यहाँ वही सवाल measurement से जवाब दिया गया है, एक laptop GPU पर। 98,624 से 15 million parameters तक के बीस models को Wikipedia के 174 million tokens पर scratch से train किया गया — 2,048-token BPE vocabulary, उसी तरह train की गई जैसे Chapter 7 में होती है, Chapter 9 का transformer। हर run को तीन compute budgets में से ठीक एक मिला और उससे एक operation भी ज़्यादा नहीं, इसलिए बड़ा model अनिवार्य रूप से कम text पढ़ता है। हर budget पर पहुँचा सबसे अच्छा held-out loss:

TEXT
budget C (FLOPs)   best loss   reached by a model of
       1.00e13       5.3531           98,624 params
       3.16e13       4.8638           98,624 params
       1.00e14       4.3383          295,808 params

fitted:  L = (Cc / C)^0.0913     over one decade of compute

दस गुना arithmetic loss को 19 % घटा देता है, और तीनों points log-log में सीधी line पर बैठते हैं। पहले नौ chapters में कुछ भी इसकी भविष्यवाणी नहीं करता। इसके पीछे कोई theorem नहीं है — यह एक empirical regularity है, जो इस laptop और datacentre के बीच के दस orders of magnitude पर अलग exponent के साथ बनी रहती है, और यही वह अकेला observation है जिसने एक industry को GPUs पर एक छोटे देश की GDP जितना खर्च करने के लिए राज़ी किया।

Objective में कुछ नहीं बदलता। Model अब भी next token predict करता है, loss अब भी Chapter 4 का cross-entropy है जिसे Chapter 8 के factorisation पर लगाया गया है, optimiser अब भी Chapter 6 का AdamW है। Pretraining कोई नया algorithm नहीं है; यह वही algorithm है जिसे इतने बड़े corpus पर चलाया गया है कि run का budget बनाना पड़ता है। दो बातें इसे संभव बनाती हैं: labels मुफ्त हैं, क्योंकि position tt का target t+1t+1 पर token है और वह text में पहले से मौजूद है; और Chapter 6 के आखिरी section ने objection हटा दिया था, क्योंकि classical rules से कहीं अधिक parameters वाला model टूटता नहीं, बेहतर होता है। जो निकलता है वह base model है — कुछ ऐसा जो answer देने के बजाय text जारी रखता है।

इनमें से किसी भी चीज़ का budget बनने से पहले उसे count करना पड़ता है, और field इसे एक formula से count करती है:

C6NDC \approx 6ND

जहाँ NN parameter count है, DD training tokens हैं और CC total floating-point operations हैं। Kaplan et al. इसे दो steps में derive करते हैं।1 Forward: प्रति parameter प्रति token 2 FLOPs, क्योंकि matrix multiply में हर parameter प्रति token एक बार use होता है, एक multiply और एक add में। Backward: forward का दोगुना, क्योंकि Chapter 5 का backward pass हर layer पर दो gradients compute करता है — layer के inputs के respect में, ताकि signal आगे यात्रा करता रहे, और उसके weights के respect में — हर एक forward जितने size का matrix multiply है, इसलिए 4N4N

यही पूरी derivation है, और इसे मानने के बजाय जाँचना चाहिए। PyTorch एक real FLOP counter, torch.utils.flop_counter.FlopCounterMode, ship करता है, जो model द्वारा dispatch किए गए हर operation को intercept करता है और actual work को total करता है। इसे चार orders of magnitude पर run करें, सबसे बड़े को meta device पर, जो shapes allocate करता है और memory नहीं:

flops.pyPYTHON
from torch.utils.flop_counter import FlopCounterMode

counter = FlopCounterMode(display=False)
with counter:                       
    loss = model(x, targets)[1]     
    loss.backward()                 
measured = counter.get_total_flops()
print(measured / (6 * n_params * n_tokens))
configurationNN without embeddingsNN totalmeasured, fwd+bwd÷ 6ND6ND (total NN)÷ 6ND6ND (no emb.)fwd+bwd ÷ fwd
dd 128, 4 layers, TT 256788,7367,254,4004.60e101.0319.4853.000
dd 512, 8 layers, TT 25625,183,23251,045,8883.26e111.0382.1043.000
dd 768, 12 layers, TT 102484,973,056124,356,8641.75e121.1451.6763.000
dd 1600, 48 layers, TT 10241,474,870,4001,556,920,0002.10e131.1001.1613.000
dd 4096, 32 layers, TT 20486,442,983,4246,582,444,0328.74e131.0801.1043.000
dd 8192, 80 layers, TT 819264,427,147,26465,544,929,2803.75e151.1631.1833.000

Forward+backward over forward ratio 3.000 है, हर scale पर exactly: कोई approximation नहीं जो संयोग से अच्छी निकली हो, बल्कि ऊपर की arithmetic identity एक ऐसे counter द्वारा round number के रूप में लौटाई गई है जिसे derivation के बारे में कुछ नहीं पता।

Measured total फिर 6ND6ND से 3 % और 17 % ऊपर बैठता है, जब NN embedding matrices को count करता है — और यह clause मायने रखता है, क्योंकि दो founding papers NN को अलग-अलग count करते हैं। Kaplan “all vocabulary and positional embeddings” को exclude करता है क्योंकि ऐसा करने से “significantly cleaner scaling laws” बनते हैं (§1.3); Chinchilla का Appendix F कहता है “we also count embeddings matrices in the total parameter count”.2 Wide vocabulary और narrow hidden dimension के लिए दोनों में नौ गुना फर्क है, जैसा पहली row दिखाती है।

Residual gap वह है जिसे 6ND6ND जानबूझकर छोड़ देता है: attention scores। Kaplan की Eq. (2.2) forward cost को 2N+2nlayernctxdmodel2N + 2\,n_{\text{layer}} n_{\text{ctx}} d_{\text{model}} लिखती है और second term छोड़ देती है क्योंकि dmodelnctx/12d_{\text{model}} \gg n_{\text{ctx}}/12 — 2020 में safe, अब कम safe, और यही वजह है कि T/dT/d बढ़ने पर ratio ऊपर drift करता है — इसी कारण यहाँ दो rows में TT 1,024 है और ratio गिरता है, 1.145 से 1.100 तक, जब dd 768 से 1,600 हो जाता है। यह वही O(T2)O(T^2) cost है जिसे Chapter 9 ने introduce किया और Chapter 16 price में बदलता है।

Compute तय करता है कि run कितना लंबा चलेगा; memory तय करती है कि वह शुरू भी हो सकता है या नहीं। Plain fp32 AdamW के साथ train करें और हर parameter चार numbers ढोता है: weight, उसका gradient, और Adam का running mean mm और variance vv — वे दो averages जिन्हें Chapter 6 में हाथ से बनाया गया था। चार bytes वाले चार numbers यानी 16 bytes per parameter, किसी भी activation से पहले। 8 GB laptop GPU पर measured, step के उस point पर resident allocation लेते हुए जहाँ कोई graph alive नहीं है:

modelvocabularybatchNN16N16N predictedresident measuredpeak in a stepthe difference
dd 512, 8 layers50,257851,045,888779 MB801 MB2,500 MB1,699 MB
dd 512, 8 layers4,096827,411,456418 MB426 MB1,043 MB617 MB
dd 256, 6 layers4,09685,839,36089 MB89 MB382 MB293 MB
dd 256, 6 layers4,096325,839,36089 MB89 MB1,259 MB1,170 MB
dd 256, 6 layers4,0961285,839,36089 MB89 MB4,771 MB4,681 MB

Prediction और measurement 3 % के भीतर agree करते हैं। Surprise आखिरी column है: activations model को dwarf कर देते हैं। वही 5.8-million-parameter model जिसे persistent state के लिए 89 MB चाहिए, batch 128 पर 4,681 MB activations चाहता है — model का बावन गुना — और इसका बड़ा हिस्सा transformer बिल्कुल नहीं है। यह logits हैं, प्रति token vocabulary size का एक vector, प्रति entry चार bytes: आखिरी row में 512 MB, पहली में 393 MB। Vocabulary size Chapter 7 में चुना गया था, और वह अब भी तय कर रहा है कि card पर क्या fit होगा।

कौन-सा term dominate करता है यह run के shape पर निर्भर करता है, इसलिए Micikevicius et al. कहते हैं memory “is dominated by activations”3 जबकि ZeRO कहता है कि 1.5-billion-parameter model को सिर्फ model states के लिए “at least 24 GB” चाहिए।4 ZeRO वही 16 bytes दूसरे रास्ते से पहुँचता है — fp16 weights के लिए 2Ψ2\Psi, fp16 gradients के लिए 2Ψ2\Psi, fp32 master weights और Adam के दो moments में से हर एक के लिए 4Ψ4\Psi — जो 70 billion parameters के लिए 1.12 terabytes है, किसी भी activation से पहले चौदह 80 GB GPUs के बराबर।

Frontier scale पर इनमें से कुछ भी एक device पर fit नहीं होता, इसलिए run को एक साथ चार तरीकों से split किया जाता है। Data parallelism हर GPU पर model की copy रखता है और gradients को average करता है — default, और वही जिसे ZeRO optimiser state की redundant copies रखने से इनकार करके बेहतर बनाता है। Tensor parallelism individual matrices को devices में split करता है। Pipeline parallelism हर device को layers का contiguous group देता है। Context parallelism sequence को ही split करता है, जिसकी जरूरत तभी पड़ती है जब TT इतना लंबा हो जाए कि attention term dominate करे। Llama 3 की Table 4 चारों को एक साथ list करती है: tensor 8, context up to 16, pipeline 16, data up to 128, कुल 16,384 H100 GPUs पर।5 यह course इसके बारे में बस इतना ही कहेगा; distributed training engineering अपने आप में एक semester है, और Stanford का CS336 वही semester है, lectures 5 से 8, code के साथ।6 Delegation के बाद जो बचता है वह एक number है, model FLOPs utilisation — GPU की peak arithmetic का वह fraction जो real run achieve करता है — जो साफ-सुथरे 6ND6ND को wall-clock time और इसलिए money में बदलता है।

January 2020 में Kaplan et al. ने transformers की grid train की और पाया कि test loss तीनों resources में से हर एक के साथ छह से अधिक orders of magnitude पर power law follow करता है।1 उनका §1.2 तीन fitted laws देता है:

L(N)=(NcN)αN,αN0.076,Nc8.8×1013L(N) = \left(\frac{N_c}{N}\right)^{\alpha_N}, \qquad \alpha_N \approx 0.076, \qquad N_c \approx 8.8 \times 10^{13}

साथ में data के लिए αD0.095\alpha_D \approx 0.095 और optimally allocated compute के लिए αCmin0.050\alpha_C^{\min} \approx 0.050। Constants universal नहीं हैं, और paper खुद ऐसा कहता है: “the precise numerical values of NcN_c, CcminC_c^{\min} and DcD_c depend on the vocabulary size and tokenization and hence do not have a fundamental meaning.”

Exponents tiny हैं: दस गुना parameters remaining loss से 100.0761.1910^{0.076} \approx 1.19 factor खरीदते हैं। यह कुछ भी नहीं जैसा लगता है, और यही यहाँ सबसे important fact है — returns खराब हैं और वे कभी रुकते नहीं। Small exponent वाला power law वादा करता है कि अगला order of magnitude मदद करेगा, पिछले से कम, हमेशा। Compute खरीदना gamble रहना बंद कर देता है और published exchange rate वाली purchase बन जाता है, और यही ठीक वह argument था जिसने capital unlock किया।

फिर prescription आया, और यहाँ paper इस तरह गलत था जिसने industry को बहुत पैसा खर्च करवाया। Kaplan की Table 6 NoptC0.73N_{\text{opt}} \propto C^{0.73} और DoptC0.27D_{\text{opt}} \propto C^{0.27} देती है: दस गुना compute का मतलब 5.4 गुना बड़ा model जिसे सिर्फ 1.9 गुना ज्यादा text खिलाया जाए। Abstract explicit है — “optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.” Field ने exactly यही किया: GPT-3 175 billion parameters on 300 billion tokens,7 Gopher 280 billion on 300 billion, Megatron-Turing NLG 530 billion on 270 billion।2 Across the board, प्रति parameter आधे token से दो tokens।

March 2022 में Hoffmann et al. ने 70 million से 16 billion parameters तक के 400 से अधिक models train किए और तीन independent routes से उल्टा conclusion पाया।2 उनकी Table 2 NoptCaN_{\text{opt}} \propto C^{a} में exponent aa को 0.50, 0.49 और 0.46 report करती है, Kaplan के 0.73 के against। Plain terms में: model size और training data समान proportion में बढ़ने चाहिए।

उनका दूसरा approach वही है जिसे इस chapter के ऊपर वाले sweep ने scale के millionth हिस्से पर reproduce किया: budget fix करें, कई sizes को exactly उसी budget पर train करें, final loss को model size के against plot करें।

parametersC=1013C = 10^{13}C=3.16×1013C = 3.16 \times 10^{13}C=1014C = 10^{14}
98,6245.3531 (171)4.8638 (542)
150,3205.4636 (74)
194,2085.5041 (44)4.8730 (140)4.4040 (442)
295,8085.5550 (19)4.9029 (60)4.3383 (190)
665,2805.7849 (3.8)5.1254 (12)4.4192 (38)
1,280,7685.8174 (1.0)5.1751 (3.2)4.5003 (10)
3,101,5685.4686 (0.5)4.7768 (1.7)
5,315,0725.5894 (0.2)4.8514 (0.6)
15,053,5685.3534 (0.1)

Held-out loss nats per token में, brackets में tokens per parameter, हर budget पर best model bold; dash ऐसा point है जिसे run नहीं किया गया, क्योंकि budget ने corpus से अधिक text माँगा या size वहाँ sweep किए गए sizes से बाहर था।

Column में नीचे पढ़ें: loss गिरता है, bottom out होता है और फिर चढ़ता है। एक model अपने budget के लिए उतना ही आसानी से बहुत बड़ा हो सकता है जितना बहुत छोटा101410^{14} पर 295,808 की जगह 665,280 parameters चुनने की penalty 0.08 nats है, जो ऊपर fitted envelope पर वह loss है जिसे सही size का model 18 % कम compute से reach करता है। Wrong shape चुनना budget का पाँचवाँ हिस्सा फेंक देता है। यह Chinchilla की Figure 3 है, चार सौ models के बजाय एक GPU पर एक afternoon में।

अब across पढ़ें। 101310^{13} पर best model swept में सबसे छोटा है; 101410^{14} पर यह 295,808 parameters है, दोनों sides से bracketed। Budget बढ़ने पर optimum दाईं ओर move करता है, और correction की पूरी बात यही है। Paper का third approach fit करें — हर run पर surface L(N,D)=E+A/Nα+B/DβL(N,D) = E + A/N^{\alpha} + B/D^{\beta} — और C=6NDC = 6ND के subject में minimise करें:

TEXT
L(N, D) = 24.7 / N^0.195 + 46.8 / D^0.169       (E fits to ~0; see below)
implied   N_opt ∝ C^0.464
  compare   Chinchilla 0.46-0.50 · Besiroglu 0.513 · Kaplan 0.73

0.46, laptop से, Kaplan के 0.73 के against। Three-budget fit से तीन digits तक agreement luck है; first digit तक agreement नहीं। Exponent यात्रा करता है — constant नहीं, क्योंकि इन optima पर token-to-parameter ratio 170 से 540 है, 20 नहीं। तीन वजहें, सभी instructive। EE zero fit होता है क्योंकि 4 nats से ऊपर loss पर run उस entropy floor के पास कहीं नहीं है जो Chinchilla के fit को dominate करता है। Batch size और learning rate को per point tune करने के बजाय fixed रखा गया, जो उन runs को handicap करता है जिन्हें सबसे कम steps मिलते हैं — और वे large models हैं: 101310^{13} FLOPs पर 1.28-million-parameter model को कुल 159 optimiser steps मिलते हैं, उन few thousand से बहुत नीचे जिन्हें Kaplan का SminS_{\min} term किसी भी model के लिए जरूरी बताता है। Scaling law एक regime के अंदर fit होता है, और यह Chinchilla से छह orders of magnitude नीचे बैठता है।

इसीलिए paper का abstract: “current large language models are significantly undertrained”. Chinchilla demonstration है — 70 billion parameters on 1.4 trillion tokens, Gopher के 280 billion on 300 billion जितना same total compute, 57 में से 51 MMLU tasks पर उसे beat करता हुआ, 60 % के against 67.5 %।2 चार गुना छोटा, साढ़े चार गुना ज्यादा text, वही पैसा, बेहतर model।

उस famous ratio पर दो caveats। “Twenty tokens per parameter” paper में sentence नहीं है, जो केवल इतना कहता है कि “for every doubling of model size the number of training tokens should also be doubled”; 20 Table 3 और Chinchilla के अपने 70 B on 1.4 T से inference है। और इसकी precision published से खराब है: Besiroglu et al. ने Figure 4 की digitisation से refit किया, पाया कि original parameters “fit the reconstructed data poorly” और intervals “implausibly tight given the number of data points” थे, और honest range को “between 4 and 40” tokens per parameter रखा।8

Chinchilla की method की एक detail Chapter 1 के learning-rate schedules वाले वादे को पूरा करती है। Cosine schedule को token budget से match करना पड़ता है। जो model 10 million tokens देखेगा उसे अपना learning rate 10 million tokens पर zero तक decay करना चाहिए; उसे 100 million के लिए sized schedule दें, जल्दी रोक दें, और आप descent के बीच का loss पढ़ रहे हैं, ऐसे rate पर जो बहुत ज्यादा है। Chinchilla exactly इसे control करने के लिए हर model को चार cycle lengths पर train करता है; ऊपर का sweep उसी कारण अपना schedule budget से set करता है।

वे field का सबसे useful empirical result हैं और routinely oversold हैं। चार limits।

वे loss predict करते हैं, capability नहीं। Left-hand side held-out text पर cross-entropy है। इन papers में कुछ भी यह claim करने की license नहीं देता कि model correct SQL लिखेगा, harmful request refuse करेगा या tool use करेगा। यह Chapter 5 का lesson फिर है: loss की prediction उस behaviour की prediction नहीं है जिसके लिए आप pay कर रहे हैं।

वे fitted हैं, derived नहीं। कोई theory αN=0.076\alpha_N = 0.076 produce नहीं करती। Constants tokenizer के साथ move करते हैं — इसी कारण दो tokenizers के बीच perplexity comparison meaningless है, जैसा Chapter 8 ने समझाया — और data mixture, architecture और optimiser के साथ भी। हर published law उस setup का law है जिसने उसे produce किया, इसलिए Meta ने Llama 3 से पहले अपना खुद का refit किया।5

वे हर step के लिए fresh token assume करते हैं, जो चुपचाप infinite corpus assume करता है। Muennighoff et al. ने measure किया कि जब यह खत्म हो जाता है तो क्या होता है: repeated data के चार epochs तक लगभग कुछ cost नहीं होता — 44 billion unique tokens पर 8.7-billion-parameter model, जिन्हें चार बार देखा गया, 178 billion unique tokens वाले same model से “only 0.5 % higher validation loss” पर finish हुआ — जबकि करीब sixteen epochs के बाद additional compute कुछ नहीं खरीदता।9

और अब कोई compute-optimal train नहीं करता। Chinchilla training की cost minimise करता है; deployed model फिर generated token प्रति roughly 2N2N FLOPs देता है, हमेशा। LLaMA 1 ने साफ कहा: “given a target level of performance, the preferred model is not the fastest to train but the fastest at inference”.10 Sardana et al. ने इसके बजाय 6NDtrain+2NDinference6ND_{\text{train}} + 2ND_{\text{inference}} minimise करके formalise किया, और पाया कि billion requests expect करने वाले किसी भी व्यक्ति को “smaller and longer than Chinchilla-optimal” train करना चाहिए।11 Llama 3 का §9.1 agree करता है: उसके small models “far beyond the point of compute optimal training, effectively trading training compute for inference efficiency” train होते हैं।5 Ratio obsolete नहीं है; वह ऐसे सवाल का जवाब देता है जो अब पूछा नहीं जा रहा।

Loss smoothly गिरता है। Benchmark scores कभी-कभी नहीं। Wei et al. ने ऐसे cases collect किए जहाँ कोई task training compute के orders of magnitude तक chance पर बैठा रहता है और फिर jump करता है — GPT-3 में लगभग 2×10222 \times 10^{22} FLOPs पर three-digit arithmetic दिखना, MMLU का 33 और 5×10235 \times 10^{23} के बीच guessing से ऊपर उठना — और pattern को नाम दिया: “an ability is emergent if it is not present in smaller models but is present in larger models”.12 अगर यह real property है, तो cheap experiments से extrapolate करना unsafe है, क्योंकि आप जो capability खरीद रहे हैं वह शायद ऐसे किसी scale पर exist ही न करे जिसे test करना आप afford कर सकें।

Schaeffer, Miranda और Koyejo ने argue किया कि इसका अधिकांश measurement का artefact है, और mechanism arithmetic है।13 Per-token loss smoothly गिरता है, इसलिए एक token के सही होने की probability, exp(L)\exp(-\mathcal{L}), धीरे-धीरे improve करती है। Model को exact string match से score करें, LL-token answer पर, और आप उस probability को LL power तक raise कर देते हैं — large power पर उठी smooth curve cliff जैसी दिखती है। उसी outputs पर ऐसा metric लगाएँ जो सभी tokens मांगने के बजाय tokens count करता हो, और “the family's performance smoothly, continuously and predictably improves with increasing scale”.

उनका audit याद रखने वाला number है — “of the 39 preferred metrics in BIG-Bench, at most 5 display emergence”, और claimed cases के 92 % से अधिक के लिए दो discontinuous metrics जिम्मेदार हैं — और उनकी caution भी: “nothing in this paper should be interpreted as claiming that large language models cannot display emergent abilities”. Chart में jump metric के बारे में evidence है, जब तक कुछ और साबित न हो। Chapter 29 वह जगह है जहाँ यह आपकी problem बनता है, क्योंकि hard-cutoff metric चुनना ऐसा decision है जिसे आप notice किए बिना करेंगे।

Corpus pretraining run का वह हिस्सा है जिसके साथ कोई equation attached नहीं है, और जहाँ अधिकतर consequential decisions रहते हैं। Raw material web crawl है: Common Crawl के August 2026 archive में “2.14 billion web pages or 360 TiB of uncompressed content” है, इसका एक महीना, download के लिए free।14 जैसा है वैसा लगभग कुछ भी usable नहीं है। T5 paper कहता है कि crawl “largely comprises gibberish or boiler-plate text like menus, error messages, or duplicate text” है, और उसने जो C4 pipeline introduce की वह blunt heuristics की list है — केवल terminal punctuation पर खत्म होने वाली lines रखें, तीन से कम sentences वाले pages drop करें, curly brace या public obscenities list के किसी word वाला कोई भी page drop करें — monthly text के बीस terabytes को लगभग 750 GB में बदलते हुए।15

Blunt सही word है। Dodge et al. ने audit किया कि वे filters क्या हटाते हैं और पाया कि obscenity blocklist African-American English में 42 % documents और Hispanic-aligned English में 32 % documents delete करती है, White-aligned English के 6.2 % के against, जिससे corpus का 97.8 % आखिरी category रह जाता है।16 Dialect के बारे में कोई opinion न रखने वाले rule की एक opinion निकली।

फिर deduplication, जो housekeeping नहीं है: Lee et al. ने C4 में 61-word sentence को 61,036 बार repeated पाया, और दिखाया कि deduplicating models के “emit memorized text” करने की rate को दस गुना काट देता है, generated tokens के 1.9 % से 0.19 % तक।17 हालांकि more बेहतर नहीं है — FineWeb की team ने 96 crawls में globally deduplicate किया, 4 trillion tokens मिले और कोई measurable gain नहीं, फिर हर crawl को अलग deduplicate किया, 20 trillion मिले, और best existing corpus match किया।18

फिर contamination। Llama 3 ने अपना measure किया और publish किया: AGIEval का 98 %, BIG-Bench Hard का 95 % और HellaSwag का 85 % training set से 8-grams द्वारा overlap करता है, और MMLU के लिए overlap इतना high कि “it is impossible to get a good performance gain estimate”.5 GPT-3 के §4 में filtering bug report है जिसने benchmarks को data में छोड़ दिया और वापसी का कोई रास्ता नहीं था: “because of cost considerations it was infeasible to retrain the model”.7

Provenance unresolved हिस्सा है। The Pile ने Books3 नाम का 100.96 GiB component ship किया — corpus का 12 % और, paper की अपनी consent table के अनुसार, private torrent tracker से books;19 copyright complaint के बाद August 2023 में उसे offline कर दिया गया। September 2026 तक legal position unsettled है, और trend के रूप में cite किए गए तीन US rulings एक-दूसरे से disagree करते हैं। Alsup ने lawfully acquired books पर training को “exceedingly transformative” पाया जबकि pirated copies से बनी library को नहीं, और Anthropic ने उस half को $1.5 billion में settle किया, 482,460 works cover करते हुए, roughly $3,000 each, 20 July 2026 को approved।20 Chhabria ने Meta को summary judgment दिया जबकि लिखा कि उनका ruling “does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful”, केवल इतना कि “these plaintiffs made the wrong arguments”.21 Bibas ने Ross Intelligence के खिलाफ ruling देते हुए note किया कि “only non-generative AI is before me today”.22 किसी US appellate court ने इस सवाल पर ruling नहीं दी है।

लोग वे हिस्से करते हैं जो loss नहीं कर सकता। TIME ने January 2023 में report किया कि OpenAI के लिए Sama firm के through toxic text label करने वाले workers “between around $1.32 and $2 per hour” घर ले जाते थे, child sexual abuse, torture और self-harm describe करने वाले passages पढ़ते हुए, जबकि OpenAI ने Sama को इस काम के लिए $12.50 an hour pay किया; Sama pay range और quota दोनों पर dispute करता है।23 यह pretraining के आसपास की filtering है, खुद pretraining नहीं — लेकिन यह same invoice पर है, और यही वह जगह है जहाँ एक person बैठता है।

Electricity real है और अक्सर misquoted होती है। सबसे careful published figure BLOOM का है: run के लिए 1,082,990 GPU-hours, 433 MWh और 24.7 tonnes CO₂ equivalent, manufacturing और idle nodes count करने पर 50.5;24 Patterson et al. GPT-3 को 1,287 MWh और 552 tonnes पर रखते हैं।25 दो cautions। BLOOM का advantage efficiency नहीं बल्कि 57 g CO₂ per kWh वाला French nuclear grid है — उसने OPT-175B से ज्यादा energy use की। और field की सबसे-quoted emissions figure, Strubell et al. की neural architecture search के लिए 626,155 lb, बाद में 88 times too high निकली, क्योंकि उसने assume किया था कि search full model size पर चली जबकि वह proxy पर चली थी।26 LBNL का framing defensible है: US data centres ने 2024 में 192 TWh use किया, national electricity का 4.7 % — एक number जो industry से attached है, किसी एक run से नहीं।27

Through-line वह है जिसे Bender et al. ने documentation debt कहा: “putting ourselves in a situation where the datasets are both undocumented and too large to document post hoc”.28 ऊपर का हर fact इसलिए exist करता है क्योंकि किसी ने देखा। जिन models को ज्यादातर लोग use करते हैं उनके पीछे के corpora के लिए, कोई नहीं देख सकता।

अब वह arithmetic जो हर कोई चाहता है, चार cited inputs से, ताकि जब वे stale हों तो obvious रहे कि किसे replace करना है।

NVIDIA का H100 page BF16 tensor-core throughput के 1,979 teraFLOPS list करता है, एक footnote के तहत जिसमें “with sparsity” लिखा है।29 कोई pretraining run structured sparsity use नहीं करता, इसलिए dense figure उसका आधा है: 989.5 TFLOP/s

Llama 3 की Table 4 38–43 % BF16 model FLOPs utilisation report करती है। 40 % लें: प्रति GPU 395.8 TFLOP/s useful arithmetic।5

Lambda का 8×H100 SXM node के लिए on-demand price, accessed 2026-09-06: $3.99 per GPU-hour, यानी node के लिए $31.92 an hour।30

Chinchilla का ratio, D=20ND = 20N, C=6ND=120N2C = 6ND = 120N^2 देता है और इसलिए N=C/120N = \sqrt{C/120}

budgetH100-hoursFLOPscompute-optimal paramstokenson one 8×H100 nodeGPUs to finish in 90 days
$100253.6e19546 M10.9 B3.1 h1
$1,0002513.6e201.73 B34.5 B31.3 h1
$10,0002,5063.6e215.46 B109 B13 days2
$100,00025,0633.6e2217.3 B345 B131 days12
$1,000,000250,6273.6e2354.6 B1.09 T4 years116
$10,000,0002,506,2663.6e24173 B3.45 T36 years1,160
$100,000,00025,062,6573.6e25546 B10.9 T358 years11,603

आखिरी दो columns साथ पढ़ें। $10,000 पर आपको एक rented node पर पखवाड़े में 5-billion-parameter model मिलता है। $100,000,000 पर arithmetic 546 billion parameters कहता है — और तीन months के लिए wired together बारह हजार H100s, जो ऐसा कुछ नहीं है जिसे credit card से rent किया जाए। लगभग $100,000 के बाद binding constraint money रहना बंद कर देता है और cluster बन जाता है।

ऐसी table पर भरोसा करने से पहले, उसे उन runs से test करें जिनकी real cost published है — llm.c GPT-2 124M को 8×A100 node पर “~90 minutes” में “about $20” में reproduce करता है, और GPT-2 1.6B को 8×H100 node पर 24 hours में $672 में।31

TEXT
$672, against what $672 actually bought (llm.c GPT-2 1.6B, one 8xH100 node, 24 h)
  this table predicts:      168 H100-hours   N = 1.41 B params   D = 28.3 B tokens
  what was actually run:    192 H100-hours   N = 1.558 B params  D = 33.6 B tokens

Llama 3 405B, against Meta's own published GPU-hours
  from the paper's 3.8e25 FLOPs at 40 % MFU:   26.67 M H100-hours
  published in Meta's Llama 3.1 model card:    30.84 M H100-hours    ratio 0.86

दोनों लगभग 15 % के भीतर, जो इस तरह के estimate को मिलनी चाहिए accuracy के करीब है और उस accuracy से काफी बेहतर है जिसके साथ इसे आमतौर पर quote किया जाता है।

इस subject में सबसे repeated figure यह है कि 2019 में करीब $43,000 cost करने वाला GPT-2-class model आज कुछ दर्जन dollars में reproduce किया जा सकता है। Modern half अच्छी तरह documented है; historical half नहीं।

आज। Karpathy के nanochat README: “you can train your own GPT-2 capability LLM ... for only $48 (~2 hours of 8XH100 GPU node) ... On a spot instance, the total cost can be closer to ~$15.”32 यहाँ “GPT-2 capability” precise और published है — GPT-2 के CORE score 0.256525 को beat करना — ऐसे leaderboard पर जिसकी 14 March 2026 तक best entry 1.65 hours है। $48 $3 per GPU-hour assume करता है, Lambda की list $3.99 से नीचे; list पर यह $64 के करीब है।

2019 में। कोई primary source नहीं है: OpenAI ने कभी duration या cost publish नहीं की। Chain The Register, February 2019 से चलती है, जिसमें “256 Google TPU3 cores” report है, price और duration नहीं; फिर Synced, June 2019, note करता है कि Google Cloud पर hardware cost $256 an hour थी और explicitly कहता है कि “OpenAI didn't specify the training duration”. $43,008 यानी $256 an hour times assumed 168 hours, जिसका source कभी किसी ने नहीं दिया।

इसलिए honest headline है: GPT-2 के published benchmark score को match करने वाला model आज rented hardware पर $100 से काफी कम में train किया जा सकता है, ऐसे 2019 cost के against जो कभी publish नहीं हुई और जिसका famous estimate duration के बारे में unsourced guess पर टिका है। Collapse real है और modern half credit card वाले किसी भी व्यक्ति द्वारा reproducible है; ratio ऐसे number पर arithmetic है जो exist नहीं करता। Published training costs की generally यही state है। GPT-3 paper में कोई dollar amount नहीं है, केवल Table D.1 में 3.14×10233.14 \times 10^{23} FLOPs;7 Llama 3 paper में भी कोई नहीं।5 आपने जो भी training cost पढ़ी है वह FLOP count, hardware assumption और price assumption से estimate है — हमेशा पूछने लायक कि किसका।

जो निकलता है उसने fixed moment पर assemble किया गया fixed corpus देखा है, और इससे दो properties follow करती हैं।

पहली है knowledge cutoff। Collection date के बाद model कुछ नहीं जानता — “uncertain” नहीं, कुछ नहीं — और वह ऐसा कहने के बजाय fluently confabulate करेगा, क्योंकि ऐसा कहना कोई behaviour नहीं था जिस पर उसे train किया गया था। Llama 3.1 का model card December 2023 देता है;33 हर model का एक होता है, और यह training data की property है, deployment की नहीं। इसके around काम करना retrieval problem है, जो Chapter 19 है।

दूसरी यह है कि base model answers के बजाय complete करता है। उसे “What is the capital of France?” दें और plausible continuation एक और question है, क्योंकि corpus में वह string अक्सर exercises की list में आती है।

Text completer assistant नहीं है। यह instructions follow नहीं करता, क्योंकि corpus में कुछ भी नहीं बताता कि request को obey किया जाना चाहिए, continue नहीं। इसे दो participants वाली conversation की कोई notion नहीं है। यह harmful prompt का सबसे probable continuation खुशी से produce करेगा, क्योंकि probable ही वह अकेली चीज़ थी जिसके लिए इसे optimise किया गया था।

इसे answer देने वाली चीज़ में बदलने में दूसरी stage लगती है, जिसकी cost पहली की fraction of a per cent है, और जो लगभग पूरी तरह इसे आपके चाहें behaviour के examples दिखाने और फिर इसके अपने outputs के pairs compare करने से बनी है। वही stage है जहाँ से instruction following, chat templates, refusals और — यह लोगों को surprise करता है — tool call करने की ability आती है। Chapter 11 वही stage है: supervised fine-tuning, RLHF, DPO और GRPO, और यह सवाल कि “aligned” का मतलब क्या है और decide कौन करता है।


इस chapter के साथ पढ़ने लायक और भी: Karpathy का build-nanogpt और उसका accompanying video, जो full GPT-2 reproduction को end to end ऐसे pace पर walk कराते हैं जो यह chapter नहीं कर सकता; और Stanford CS324, Large Language Models, जिसकी data और environmental impact वाली lectures ऊपर के section से ज्यादा deep जाती हैं, उस material में जिसे यह course एक बार treat करके delegate करता है।

  1. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D. Scaling Laws for Neural Language Models. arXiv:2001.08361 (2020). तीन power laws §1.2 में Eqs. (1.1)–(1.3) हैं और full constants Appendix A, Table 5 में हैं; 6N6N derivation §2.1 है; compute-allocation exponents Table 6 हैं। Note करें कि दो compute laws हैं, fixed batch size पर αC=0.057\alpha_C = 0.057 और optimal batch size पर αCmin=0.050\alpha_C^{\min} = 0.050; paper कहता है कि बाद वाले को “should be used to make predictions”. 2

  2. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E. et al. Training Compute-Optimal Large Language Models. arXiv:2203.15556 (2022). Exponents Table 2 में, projected budgets Table 3 में, Gopher comparison §4 में, parameter-counting convention Appendix F में। Table 3 के नीचे की prose 175 B और 280 B rows के लिए Table 3 से disagree करती है; quote करने वाला version table है। 2 3 4

  3. Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D. et al. Mixed Precision Training. arXiv:1710.03740 (2017), ICLR 2018. FP32 master weights §3.1 में, loss scaling §3.2 में। 2

  4. Rajbhandari, S., Rajbhandari, S., Ruwase, O. and He, Y. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 (2019), SC20. 16Ψ16\Psi accounting §3.1 है; activations के residual-state figures §3.2 हैं।

  5. Grattafiori, A. et al. (Llama Team, AI @ Meta). The Llama 3 Herd of Models. arXiv:2407.21783 (2024). Compute budget और token count §1 में, refitted scaling law §3.2.1 में, parallelism configuration और MFU Table 4 में, contamination analysis §5.1.4 में, over-training statement §9.1 में। Paper में कोई dollar figures और कोई emissions table नहीं है। 2 3 4 5 6

  6. Stanford CS336, Language Modeling from Scratch. Lecture 2 resource accounting cover करता है, lectures 5–8 GPUs, kernels और parallelism, lectures 9 और 11 scaling, lectures 13–14 data। यही वह course है जिसे यह chapter अपनी engineering delegate करता है, और यह public है।

  7. Brown, T. B. et al. Language Models are Few-Shot Learners. arXiv:2005.14165 (2020). Compute Appendix D, Table D.1 में — जिसमें एक column literally headed “flops per param per token” है, जिसका हर GPT-3 row के लिए value 6 है। Contamination analysis §4 में। 2 3

  8. Besiroglu, T., Erdil, E., Barnett, M. and You, J. Chinchilla Scaling: A replication attempt. arXiv:2404.10102 (2024). Chinchilla के data को उसकी Figure 4 digitise करके reconstruct करता है, refit करता है, और corrected exponents तथा काफी wider intervals report करता है।

  9. Muennighoff, N., Rush, A. M., Barak, B., Le Scao, T., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T. and Raffel, C. Scaling Data-Constrained Language Models. arXiv:2305.16264 (2023), NeurIPS 2023. Four-epoch result §6 है; sixteen-epoch half-life fitted RD15R_D^* \approx 15 है।

  10. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T. et al. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 (2023). §1 Chinchilla-optimal training के खिलाफ inference-cost argument state करता है।

  11. Sardana, N., Portes, J., Doubov, S. and Frankle, J. Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws. arXiv:2401.00448 (2023), ICML 2024. उनके §5 में counterweight भी है: extreme token ratios पर train किए गए models improve करते रहते हैं, लेकिन “more slowly than scaling laws predict”.

  12. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S. et al. Emergent Abilities of Large Language Models. arXiv:2206.07682 (2022), TMLR. Definition §2 में, examples और compute thresholds §3–4 तथा Table 1 में।

  13. Schaeffer, R., Miranda, B. and Koyejo, S. Are Emergent Abilities of Large Language Models a Mirage? arXiv:2304.15004 (2023), NeurIPS 2023 outstanding paper. Metric argument §2 है, BIG-Bench meta-analysis §4, constructed vision example §5।

  14. Common Crawl, August 2026 Crawl Archive Now Available (CC-MAIN-2026-34), published 24 August 2026, accessed 2026-09-06. इसका अपना front page “over 300 billion pages spanning 15 years”, “totalling more than 10 petabytes” claim करता है — यह whole archive का figure है, यहाँ priced monthly crawl का नहीं।

  15. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W. and Liu, P. J. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 (2019), JMLR 21(140). C4 filters §2.2 हैं। Paper sizes bytes में देता है, tokens में नहीं; उससे widely attributed 156-billion-token figure नीचे Dodge et al. से है।

  16. Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M. and Gardner, M. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758 (2021), EMNLP 2021. Dialect removal rates §5.3 हैं; C4 में benchmark contamination §4.2 है।

  17. Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C. and Carlini, N. Deduplicating Training Data Makes Language Models Better. arXiv:2107.06499 (2021), ACL 2022. 61,036 repeats footnote 1 हैं; memorisation figures §6.2, Table 4 में हैं, और 50-token exact-match criterion के तहत generated tokens के percentages हैं।

  18. Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L. and Wolf, T. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557 (2024), NeurIPS 2024 Datasets and Benchmarks. Deduplication result §3.4 है। Released dataset तब से paper के 15 trillion tokens से आगे बढ़ चुका है।

  19. Gao, L., Biderman, S., Black, S. et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv:2101.00027 (2020). Books3 §2.3 और Table 1 है; consent table Table 5 है। Corpus 825.18 GiB है, इसलिए title भी round-down है।

  20. Bartz v. Anthropic, No. 4:24-cv-05417 (N.D. Cal.). Fair-use order 23 June 2025 (Dkt. 231); class certification 17 July 2025; final approval and judgment 20 July 2026 (Dkt. 680). Settlement केवल past inputs release करता है, outputs नहीं और future conduct नहीं।

  21. Kadrey v. Meta, No. 3:23-cv-03417-VC (N.D. Cal.), summary judgment 25 June 2025 (Dkt. 598). Note करें कि torrenting पर distribution claim decide नहीं हुआ और live रहता है।

  22. Thomson Reuters v. ROSS Intelligence, No. 1:20-cv-00613-SB (D. Del.), revised opinion 11 February 2025 (Dkt. 770), Bibas J. Third Circuit (No. 25-2153) में interlocutory appeal पर, argued 11 June 2026, लिखते समय undecided।

  23. Perrigo, B. Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. TIME, 18 January 2023. $2 senior reviewers के लिए ceiling है जिन्होंने हर target meet किया; junior labellers, majority, $1.32 घर ले गए। Sama का rebuttal, उसी article में quoted, $1.46–$3.74 और lower quota देता है।

  24. Luccioni, A. S., Viguier, S. and Ligozat, A.-L. Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. arXiv:2211.02001 (2022), JMLR 24(253). Tables 1 और 3।

  25. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M. and Dean, J. Carbon Emissions and Large Neural Network Training. arXiv:2104.10350 (2021). GPT-3 के figures Table 4 हैं; NAS estimate की correction §4.1 है।

  26. Strubell, E., Ganesh, A. and McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. arXiv:1906.02243 (2019), ACL 2019. Precisely पढ़ने लायक है क्योंकि उसके सबसे-quoted number के साथ क्या हुआ: paper careful है, अपनी extrapolation state करता है, और फिर भी उस one line पर two orders of magnitude गलत था जिसे सबने repeat किया।

  27. Smith, S. J., Hubbard, A., Newkirk, A., Ganeshalingam, M., Holecek, B., Sartor, D., Mills, M. and Shehabi, A. United States Data Center Energy Usage Report: 2025 Update. LBNL-2001758 (18 June 2026). यह historical series के लिए widely cited 2024 report को downward revise करता है; अगर आप 2023 के लिए 176 TWh figure quote कर रहे हैं, तो आप superseded edition quote कर रहे हैं।

  28. Bender, E. M., Gebru, T., McMillan-Major, A. and Shmitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? FAccT '21, pp. 610–623. DOI 10.1145/3442188.3445922. “Documentation debt” §4.4 है। Note करें कि paper के अपने carbon figures Strubell et al. से cited हैं और ऊपर की correction inherit करते हैं — जो इसके argument का illustration है, refutation नहीं।

  29. NVIDIA. NVIDIA H100 Tensor Core GPU product page, nvidia.com/en-us/data-center/h100/ (accessed 2026-09-06). उस page पर FP64 छोड़कर हर tensor-core row में footnote “with sparsity” है; यहाँ use की गई dense BF16 figure published 1,979 TFLOPS की आधी है।

  30. Lambda. GPU Cloud pricing, lambda.ai/pricing (accessed 2026-09-06). On-demand, per GPU per hour, tax से पहले। इस section में prices course की किसी भी चीज़ से तेज stale होंगे; उनके आसपास की arithmetic नहीं।

  31. Karpathy, A. karpathy/llm.c, discussion #481, Reproducing GPT-2 (124M) in llm.c in 90 minutes for $20 (28 May 2024), और discussion #677, Let's reproduce GPT-2 (1.6B): one 8XH100 node, 24 hours, $672, in llm.c (11 July 2024)।

  32. Karpathy, A. karpathy/nanochat, README और “time to GPT-2” leaderboard (accessed 2026-09-06). $48 figure और “GPT-2 capability” की CORE-score definition दोनों README में हैं; repository का अपना speedrun.sh “approximately 1.5 hours” कहता है, इसलिए two-hour figure को rounded मानें।

  33. Meta. Llama 3.1 model card, models/llama3_1/MODEL_CARD.md in meta-llama/llama-models (accessed 2026-09-06). 405 B model के 30.84 M H100-hours, total 39.3 M, 11,390 tCO2eq location-based figure, और December 2023 data cutoff का source।


निर्माता

David Vicente Campos

NeuraLIA Labs के संस्थापक और MyRealFood के सह-संस्थापक

मैं लेओन विश्वविद्यालय से कंप्यूटर इंजीनियर हूँ। मैंने MyRealFood की सह-स्थापना की, जहाँ CTO के रूप में मैंने वह ऐप बनाया जिसे लाखों लोग बेहतर खान-पान के लिए इस्तेमाल कर चुके हैं, और मैंने NeuraLIA Labs की स्थापना की, जहाँ मैं AI प्रोडक्ट्स बनाता हूँ। यहाँ मैं उन बातों के बारे में लिखता हूँ जो इस सफ़र में मुझे समझनी पड़ीं, उस तरह जिस तरह काश किसी ने मुझे समझाई होतीं।

लेखक के बारे में और जानें

NeuraLIA Labs द्वारा प्रकाशित।

नए पोस्ट अपने इनबॉक्स में पाएं

AI समाचार, गाइड और प्रोडक्ट अपडेट — जब भी हम कुछ उपयोगी प्रकाशित करें, एक छोटा ईमेल।

कोर्स सूची

Abstract software decision engine with branching paths, probability nodes, and glowing gates.
jev13 मिनट पढ़ें

Jev AI मॉडल गद्य के लिए नहीं, निर्णयों के लिए बना है

TypeSafe AI का Jev ध्यान खींच रहा है क्योंकि यह सॉफ्टवेयर इंटेलिजेंस को संभावना की समस्या मानता है: सही शाखा चुनें, भरोसे का स्तर जोड़ें, और जब कोड को निर्णय चाहिए तो LLM से टेक्स्ट लिखवाने पर खर्च न करें।

Abstract agent runtime sorting documents, memory blocks and pointer nodes inside a bounded context frame.
context-engineering14 मिनट पढ़ें

लॉन्ग-होराइजन AI एजेंट्स के लिए कॉन्टेक्स्ट इंजीनियरिंग

लंबे समय तक चलने वाले एजेंट सिर्फ इसलिए असफल नहीं होते कि विंडो छोटी है। वे तब असफल होते हैं जब फ़ाइलें, टूल आउटपुट और पुराना इतिहास उस काम को ही पीछे धकेल देते हैं जिसे एजेंट को पूरा करना था।

मॉडल चुनने का काम LIA पर छोड़ने के लिए तैयार हैं?

हर AI मॉडल एक ही जगह — आज ही मुफ़्त शुरू करें।