More

    AI Giants Unleash 4 Frontier Models in 3 Weeks as the Race Enters Overdrive


    Key Takeaways

    Grok 4.5 from xAI, Claude Opus 5 from Anthropic, the GPT-5.6 family from OpenAI, and Kimi K3 from Moonshot AI each target the same problem: getting an AI system to carry a multi-hour task, from research to coding to structured reporting, without losing track of the plan.

    Grok 4.5 Trains on Real Developer Sessions

    xAI released Grok 4.5 on July 8. The model runs on a 1.5 trillion parameter base and was trained in part on real usage data from Cursor, the coding platform SpaceXAI acquired earlier this year. Pricing sits at $2 per million input tokens and $6 per million output tokens, with a 500,000 token context window.

    Elon discussing the release of Grok 4.5 on July 8, 2026.

    Elon Musk described Grok 4.5 as an Opus-class model that runs faster and at lower cost. Independent trackers place it fourth on the Artificial Analysis Intelligence Index, ahead of every open weight model, at pricing more than 60% below Claude Opus 4.8 and GPT-5.5. On Terminal-Bench 2.1, a test that scores how well a model completes command-line engineering tasks, Grok 4.5 scored 83.3%.

    The bigger shift is token efficiency. xAI says the model needs roughly a fifth of the output tokens Opus 4.8 required for comparable tasks, which lowers the cost of running long agent sessions.

    Claude Opus 5 Holds the Line on Price

    Anthropic released Claude Opus 5 on July 24, positioning it as a model that reaches close to the performance of Claude Fable 5, the company’s most capable release, at half of Fable’s $10 input and $50 output pricing. Opus 5 itself carries the same $5 input and $25 output pricing as its predecessor, Opus 4.8.

    Claude X post screenshot.
    Claude Opus 5 announcement via X.

    The model ships with a 1 million token context window, 128,000 max output tokens, and an adjustable reasoning effort setting that ranges from low to a new xhigh mode. Anthropic says Opus 5 sets new marks on Frontier-Bench and GDPval-AA, two coding and knowledge work evaluations, though it trails the restricted Claude Mythos 5 model on cybersecurity tasks. Opus 5 is now the default model on Claude Max and the strongest option on Claude Pro.

    OpenAI Splits GPT-5.6 Into Three Tiers

    OpenAI moved GPT-5.6 to general availability on July 9 after a two-week preview limited to roughly 20 organizations vetted by the U.S. government, following an executive order tied to frontier model safety review. The family ships as three tiers: Sol, the flagship, priced at $5 input and $30 output per million tokens; Terra, a mid-tier model at $2.50 and $15 that OpenAI says matches GPT-5.5 at half the cost; and Luna, a fast, low-cost tier at $1 and $6.

    OpenAI X post screenshot.
    ChatGPT 5.6 Sol announcement via X.

    OpenAI reports Sol leads the Artificial Analysis Coding Agent Index and hits 88.8% on Terminal-Bench 2.1, rising to 91.9% when the model runs four sub-agents in parallel under its new ultra mode. All three tiers carry OpenAI’s highest internal risk rating for cyber and biological misuse potential, which triggered added review during the government-gated preview.

    Kimi K3 Pushes Open Weight Models Toward the Frontier

    Moonshot AI released Kimi K3 on July 16, a 2.8 trillion parameter mixture of experts model with 896 total experts and 16 active per task. The model carries a 1 million token context window and native multimodal input, with API pricing at $3 input and $15 output per million tokens. Moonshot has committed to publishing full open weights by July 27.

    Kimi Moonshot X post screenshot.
    Kimi K3 announcement via X.

    K3 is the largest open-weight model released to date, roughly 75% bigger than the previous largest widely used open model. Independent trackers place K3 fourth among current frontier systems, behind Claude Fable 5 and GPT-5.6 Sol but ahead of Claude Opus 4.8.

    Why the Gap With Elder Models Matters

    Context windows below 200,000 tokens once forced developers to break large codebases or research packets into fragments. Every model in this group now runs at 500,000 tokens or beyond, with three of the four at 1 million, letting a single session hold a full repository or a stack of primary source documents.

    Agent reliability has moved as well. Systems from the 2023 and 2024 period often lost their plan after a handful of tool calls. The models released this month are built to sustain dozens of coordinated steps and recover when a tool call returns bad data instead of stalling out.

    For developers, the practical effect is a lower cost per finished task rather than a higher ceiling on any single benchmark. Effort controls in Opus 5, tiered pricing in GPT-5.6, and the efficiency claims behind Grok 4.5 point toward the same goal: letting teams choose how much compute a task deserves instead of paying flagship prices for every request.

    For enterprises and policymakers, the government-gated rollout of GPT-5.6 signals that oversight is now built into release schedules for the largest models, not added after the fact. Anthropic’s brief, government-directed suspension of Fable and Mythos access in June was an earlier version of the same pattern.



    Source link

    Latest stories

    You might also like...