The State of Frontier AI in 2025
The AI industry moves fast enough that writing about it risks immediate obsolescence, but here’s where things stand. OpenAI released GPT-4 in March 2023, followed by GPT-4 Turbo (November 2023), GPT-4o (May 2024, with native multimodality), and the o1/o3 reasoning models (late 2024/early 2025). The jump from GPT-3.5 to GPT-4 was dramatic — roughly a 30-percentage-point improvement on most benchmarks. The jump from GPT-4 to GPT-4o was more about multimodality, speed, and cost than raw intelligence improvement.
GPT-5 has been the subject of intense speculation. OpenAI CEO Sam Altman has made statements ranging from “GPT-5 will be a significant leap” to the more measured “diminishing returns may apply.” The model was reportedly in training through late 2024, with a potential release window in 2025. But delays are common in frontier AI development — GPT-4 was completed in mid-2022 and didn’t ship until March 2023, representing about 8 months of safety testing and alignment work.
The Scaling Debate
The fundamental question: does scaling transformer models with more compute, more data, and more parameters continue to yield proportional improvements? The evidence from 2023-2024 is mixed.
On one hand, scaling has been remarkably consistent. The “scaling laws” papers from OpenAI (2020) and DeepMind (2022) showed a reliable relationship between compute, data, parameters, and performance. GPT-4 is estimated to have roughly 1.76 trillion parameters (a mixture-of-experts model), and the jump from GPT-3’s 175 billion was roughly linear in performance gains per log-unit of compute.
On the other hand, there are signs of diminishing marginal returns. Anthropic’s Claude 3.5 Sonnet (mid-2024) matched or exceeded GPT-4 on many benchmarks despite likely having fewer parameters. Google’s Gemini Ultra 1.0 scored slightly above GPT-4 on MMLU (90.0% vs 86.4%) but the gap wasn’t transformational. These data points suggest that architecture innovations, training data quality, and post-training techniques (RLHF, constitutional AI, etc.) are becoming as important as raw parameter count.
The open-source community complicates this picture further. Meta’s LLaMA 3 (released 2024 with 405B parameters) achieved performance competitive with GPT-4 on many benchmarks while being freely available. Chinese lab DeepSeek’s V3 model (December 2024) was trained for a reported $5.6 million in compute — a fraction of what GPT-4 reportedly cost — and delivered comparable performance. The efficiency gap between leading labs may be narrowing.
Multimodality Is the New Table Stakes
GPT-4o marked OpenAI’s shift to natively multimodal models — the same model processes text, images, and audio without separate encoder/decoder pipelines. Google’s Gemini was designed as natively multimodal from the start. Anthropic’s Claude 3 added image understanding. Every frontier model in 2025 is multimodal.
What’s next is video understanding and generation. OpenAI’s Sora (announced February 2024, released late 2024/early 2025) generates minute-long videos from text prompts. Google’s Veo 2 competes directly. Runway’s Gen-3 Alpha focuses on creative video tools. The integration of video understanding into general-purpose models — so an AI can watch a video and reason about what’s happening in it — is the next frontier and may be a core capability of GPT-5.
The Competitive Landscape
OpenAI has real competition for the first time since GPT-3.5 made them a household name:
- Anthropic (Claude): Founded by former OpenAI researchers, their models prioritize safety and reasoning. Claude 3.5 Sonnet is genuinely excellent at coding and long-context analysis. Their 200K token context window set a standard others had to match.
- Google DeepMind (Gemini): Google’s integration of DeepMind and Google Brain under one roof is bearing fruit. Gemini Ultra is competitive with GPT-4 on most measures, and Google’s distribution advantage (Android, Search, Workspace) gives them reach OpenAI can’t match.
- Meta (LLaMA): Meta’s open-weight strategy has made LLaMA the foundation of the open-source AI ecosystem. Thousands of fine-tuned variants exist. The LLaMA 3 405B model proved that open models can compete with closed ones.
- xAI (Grok): Elon Musk’s AI venture released Grok-2 in late 2024, approaching GPT-4 level performance with a distinctive “rebellious” personality. Their advantage: real-time X/Twitter data integration and massive compute resources (100,000+ GPUs in their Memphis cluster).
- DeepSeek, Qwen, Mistral: Chinese and European labs are proving that frontier AI isn’t a two-country game. DeepSeek’s efficiency claims in particular have challenged assumptions about the cost of training frontier models.
What GPT-5 Might Actually Deliver
Based on the trajectory and what’s publicly known, here’s a realistic set of expectations for GPT-5: genuinely better reasoning — not just pattern matching, but multi-step logical inference that GPT-4 still struggles with. Native video understanding. Longer context windows (possibly 1M+ tokens). Better factuality and reduced hallucination rates. And perhaps most importantly, better agentic capabilities — the ability to use tools, browse the web, and execute multi-step tasks autonomously.
Whether that constitutes “AGI” depends entirely on your definition of that famously slippery term. It almost certainly won’t be. But it might be the first AI system that feels less like a very clever autocomplete and more like a genuine reasoning engine.
