OpenAI’s 10th anniversary, its new model and the race for superintelligence

by
0 comments
OpenAI's 10th anniversary, its new model and the race for superintelligence

OpenAI marked its tenth anniversary in December 2025 with big news: the release of GPT-5.2, a model positioned to master knowledge work, and a bold prediction from CEO Sam Altman that superintelligence is now, in his view, practically inevitable within the next decade. The new model, which followed an internal “Code Red” directive to accelerate development amid intensifying competition with Google’s Gemini 3, introduces significant improvements in long-context understanding, agentic tool use and the ability to execute complex, real-world tasks from start to finish.

The release also signals a deeper shift: OpenAI is moving away from abstract intelligence scores toward metrics that track how well AI performs actual paid work. That shift, and what OpenAI’s first decade means for the businesses now built on its models, was the subject of an extended discussion between SmarterX and Marketing AI Institute founder Paul Roetzer and his co-host on episode 186 of The Artificial Intelligence Show.

Intelligence benchmarks have stopped being informative

For years, AI progress was measured by benchmark suites that function essentially as IQ tests for machines. According to Roetzer, those metrics have hit a ceiling. “Tests of intelligence are basically saturated,” he noted — frontier models already perform at or beyond top human level on most standardized tests, so further gains on those scales are increasingly hard for anyone to feel in practice.

GDPval: measuring models against paid work

With GPT-5.2, OpenAI is leaning heavily on a newer benchmark called GDPval, which evaluates models on 1,320 well-specified tasks drawn from 44 occupations across the sectors that contribute most to US GDP — producing real-world deliverables such as legal briefs, engineering drawings and nursing care plans rather than multiple-choice answers.

The headline result is eye-opening: OpenAI reports that GPT-5.2 Thinking beat or tied top industry professionals on roughly 71% of GDPval comparisons, nearly double the win rate of its predecessor GPT-5.1, with the Pro variant reaching about 74%. Roetzer’s argument is that this is the measurement that matters — tracking models against actual work is how economic disruption becomes visible before it arrives, and by that measure it is already arriving.

Faster and cheaper than human experts — with caveats

The implications go beyond output quality. OpenAI reports that GPT-5.2 Thinking completed GDPval tasks at more than 11 times the speed and less than 1% of the cost of the human experts it was compared against. The company frames this as unlocking economic value — emphasizing help with spreadsheets, presentations and coding — and notably does not frame it in terms of replacing jobs. The subtext about job displacement is nonetheless hard to ignore, and it is the reason GDPval-style results are being watched as closely by economists as by developers.

A changed mission

The model landed during a wave of retrospectives on the company’s origins. Founded as a nonprofit in December 2015, OpenAI’s original mission was to advance digital intelligence for the benefit of humanity, unconstrained by the need to generate profit; its founding statement promised broad sharing of benefits and research. The contrast with the present is stark — Roetzer argues that essentially none of the founding paragraph still describes the organization. Today OpenAI is a major commercial force in a global race toward artificial general intelligence, competing directly with Google, Anthropic and others, and Altman used the anniversary to predict that superintelligence will be built within the coming decade.

Whatever one makes of that prediction, the trajectory — from idealistic research lab to one of the most consequential commercial entities in technology — is among the defining business stories of the era.

Limitations and what to watch

Several cautions apply to the headline numbers. GDPval is OpenAI’s own benchmark: the task selection, grading methodology and win-rate framing all come from the company whose models are being scored, and independent replication is limited so far. Win rates on well-specified tasks also overstate readiness for real jobs, which involve ambiguity, accountability and context that benchmark tasks strip away — OpenAI itself positions the results as evidence for AI-plus-human-oversight rather than substitution. And speed-and-cost comparisons measure inference, not the review time humans must still spend checking outputs. For small businesses, the practical takeaway is unchanged from earlier model cycles: capabilities are rising and prices are falling, but returns still come from disciplined deployment, not from the benchmark charts.

Related Articles