GPT-3: The Scale Hypothesis Proven
175 billion parameters. OpenAI proves that simply making models bigger produces emergent intelligence.
Explore this event on the interactive timeline →In June 2020, OpenAI released GPT-3 with 175 billion parameters—100x larger than GPT-2. For the first time, a language model could write essays, code, poetry, and legal documents with minimal prompting. GPT-3 demonstrated "emergent abilities"—capabilities that appeared spontaneously at scale, without being explicitly trained. This proved the "scaling hypothesis": intelligence emerges from sufficient data and compute.
Key Numbers
- Parameters
- 175 Billion
- Training Cost
- ~$4.6 Million
- Training Data
- 300B tokens
Verified Facts
- GPT-3 had 175 billion parameters, trained on 300 billion tokens of internet text. Training cost was estimated at $4.6 million.
- It introduced "few-shot learning": give the model a few examples in the prompt, and it can perform tasks it was never explicitly trained for.
- GPT-3 could write functional code, translate languages, answer trivia, and even generate legal contracts—all from the same model with zero task-specific training.
- The API (June 2020 beta) was a key commercial milestone, though OpenAI took in only about $3.5M in total revenue that year; the $100M+ run rate came later, in the 2023 ChatGPT era.