The strangest thing about the current generation of AI models is that they have become dramatically more capable without necessarily becoming dramatically better writers.
That sounds contradictory until you spend much time with them. In 2026, frontier models can solve difficult mathematics problems, debug complex software, operate computers, search the web, reason across huge collections of documents and carry out multi-stage tasks that would have looked implausible only a couple of years ago. The improvement in coding and reasoning, in particular, has been astonishing.
Ask the same models to produce a genuinely memorable 1,500-word magazine article, though, and the progress is harder to see. The writing is usually competent, often very good and occasionally excellent, but the familiar habits remain: over-explaining, generic transitions, unnecessary summaries and a tendency to settle into a polished but unmistakably synthetic register.
That makes the current AI writing race more interesting than the old question of which chatbot produces the prettiest paragraph.
OpenAI now has GPT-5.6 Sol. Anthropic has Claude Opus 5. Google has Gemini 3.1 Pro Preview alongside its rapidly expanding Gemini 3.x family. On August 12, xAI added Grok 4.6 to the field.
All four are extremely capable systems, but the old shorthand — Claude for prose, ChatGPT for everything else, Gemini for Google users — no longer describes the market particularly well. Most of all, the idea that Claude should automatically be crowned the best writer now deserves much more scrutiny than it used to.
Writing quality is proving to be one of the more subjective and strangely nonlinear areas of AI progress, and that may tell us something important about how machine intelligence is developing.
The four AI writing empires
Here is the Gilded Age version of the market.
| Platform | Current frontier contender | What it increasingly optimizes for | Writing advantage | Biggest question |
|---|---|---|---|---|
| Claude | Claude Opus 5 | Deep reasoning, agents, coding, knowledge work | Editing, interpretation, nuanced rewriting | Has Claude lost some of its old prose magic? |
| ChatGPT | GPT-5.6 Sol | General knowledge work and end-to-end workflows | Research + structure + editing + production | Is competence replacing personality? |
| Gemini | Gemini 3.1 Pro Preview | Multimodal reasoning and Google integration | Source-heavy, grounded writing | Can information advantage become distinctive prose? |
| Grok | Grok 4.6 | Real-time knowledge, coding and agents | Live commentary and internet-native research | Is immediacy enough to make it a serious writer? |
The specifications themselves are reaching the point of absurdity. GPT-5.6 Sol operates with million-token-class context. Claude Opus 5 offers a one-million-token context window and up to 128,000 output tokens. Gemini 3.1 Pro accepts more than one million input tokens, while Grok 4.6 arrives with a still enormous 500,000-token window.
Those numbers matter, particularly for research-heavy work, but they are becoming less useful as a way of distinguishing the models. Once every major system can ingest books, reports, transcripts and large folders of documents, the more important question is what the model actually does with all that material.
That is where the companies begin to look quite different. Anthropic, OpenAI, Google and xAI are not simply building competing text generators anymore. They are building different versions of a general-purpose knowledge worker, and writing is increasingly just one component of a much larger system.
The spiky theory of AI
One useful way to understand the current situation is through the idea of the jagged frontier, a term popularized by Wharton professor Ethan Mollick and his research collaborators.
The basic observation is that AI capability does not progress in a smooth, predictable line. A model can perform one task at what looks like an expert level, then fail badly at another task that seems much simpler. The frontier is uneven.
The image above traces an evolving debate over what the path to AGI might actually look like. Its intellectual roots go back to the “jagged technological frontier” described by researchers including Ethan Mollick and Fabrizio Dell’Acqua in 2023: AI systems can be astonishingly capable at one task while failing at another that seems equally simple. In late 2025, Tomas Pueyo visualized this idea as an expanding, irregular “blob” of AI capability, arguing that as models improve, the jagged edges will gradually fill in until AI becomes broadly competent across the full range of human tasks — the mainstream AGI thesis shown in the top row. Adam Hunt later adapted that graphic to illustrate a more provocative alternative: perhaps the frontier never smooths out. Instead, certain capabilities such as coding, mathematics or scientific reasoning could shoot far beyond human ability while other areas — common sense, social judgment or aspects of language — improve much more slowly or even stagnate. The result would not necessarily be a neat, human-like general intelligence, but an increasingly strange and uneven “spiky” intelligence: superhuman in a handful of domains while remaining surprisingly mediocre in others.
That idea feels even more relevant today. The strongest models are developing enormous spikes of capability in areas such as coding, mathematics, scientific reasoning and agentic computer use. Progress in those areas is fast, visible and relatively easy to measure. Writing is a different proposition.
The models have undoubtedly improved at factual accuracy, instruction following, long-context comprehension, editing and structural reasoning. They are better at handling a complicated brief, keeping track of a large document and following an editorial style guide than earlier generations were.
What is much harder to argue is that the best AI prose today feels proportionately more distinctive, stylish or human than the best AI prose from a year or two ago.
There is a fairly obvious reason for that. Code has tests. Mathematics has answers. Software can be executed, proofs can be checked and agents can either complete a task successfully or fail to complete it. Writing has no equivalent.
There is no unit test for a strong opening paragraph, and there is no universally accepted score for whether one sentence has better rhythm than another. More awkwardly, many of the qualities we value in good writing are difficult to reconcile with the optimization pressures placed on an assistant model. Great prose can be ambiguous, abrasive, funny, incomplete, emotionally strange or deliberately indirect. A model trained to be clear, safe, helpful and comprehensive can easily become the kind of writer who explains everything twice.
That may be one reason AI can improve dramatically at calculus without becoming proportionately funnier, or become a much better programmer without acquiring noticeably better taste.
The benchmark problem: AI can measure code better than prose
The model launches themselves offer some evidence for this. Look at what the AI labs choose to benchmark.
Anthropic’s Claude Opus 5 launch is full of Frontier-Bench, CursorBench, ARC-AGI, AutomationBench, OSWorld, scientific evaluations and tests of professional knowledge work. Anthropic has plenty of numbers showing that Opus is becoming a better software engineer, researcher and agent. What you will not find is an equally definitive public benchmark establishing that Opus 5 is a better novelist or essayist than Opus 4.8.
Google’s Gemini 3.1 Pro announcement follows a similar pattern. Google discusses creative work, but the hard claims are dominated by reasoning, coding and difficult problem-solving.
OpenAI has been more willing to talk explicitly about writing. When it launched GPT-5, for example, the company described it as its most capable writing collaborator yet and showed examples of poems and other prose. Even there, however, the quantitative evidence focused primarily on mathematics, coding, multimodal reasoning, health and knowledge work rather than presenting a single definitive measure of literary quality.
xAI is an interesting partial exception. When it launched Grok 4.1, the company highlighted performance on the independent Creative Writing v3 benchmark, which tests models across 32 writing prompts and compares outputs against detailed rubrics.
That is useful, but it immediately exposes another problem with writing benchmarks: Creative Writing v3 is itself judged by an LLM. In other words, one machine is being asked to decide whether another machine has taste.
Researchers are trying to improve on this. WritingBench covers more than 1,200 writing tasks across six broad domains and 100 subdomains, including persuasive, creative, informative and technical writing. EQ-Bench’s Longform Writing benchmark makes models plan, revise and sustain a story over multiple long-form turns. LitBench goes further by using thousands of human preference judgments, in part because evaluating literary quality with another general-purpose model remains unreliable.
The writing benchmarks worth knowing
| Benchmark | What it tries to measure | Why it matters | The catch |
| Creative Writing v3 | Creative writing across 32 prompts | xAI publicly used it to demonstrate Grok 4.1 | LLM-judged |
| Longform Writing | Planning, revision and sustained fiction | Tests coherence beyond a clever paragraph | Still uses an AI judge |
| WritingBench | 1,200+ tasks across 100 writing subdomains | Much broader than fiction alone | Automated evaluation remains imperfect |
| LitBench | Alignment with human creative-writing preferences | Built around thousands of human-labelled comparisons | More about reliable evaluation than a simple consumer leaderboard |
None of this makes writing benchmarks useless. They are getting better and they give us far more information than intuition alone.
They simply cannot settle the question in the way a coding or mathematics benchmark can, because writing eventually runs into taste.
Cormac McCarthy is not great because he would score perfectly against a corporate style guide. A board memo and a punk-rock manifesto are not improved by applying exactly the same rubric to both. Much of what makes writing distinctive comes from knowing when to break the expected pattern rather than follow it.
That is an unusually awkward challenge for systems built around predicting what should most plausibly come next.
Claude: has the former writing champion lost its mojo?
For years, Claude was the default recommendation among a certain kind of serious AI writer, and not without reason.
Earlier Claude models often felt less mechanical than the competition. They could handle rhythm and implication well, preserve the character of source material during a rewrite and, at their best, resist the urge to explain every point to death.
That reputation has carried forward, but the response to recent models has become much more mixed.
Anthropic’s development priorities are also revealing. Claude Opus 5’s headline improvements revolve heavily around software engineering, autonomous work, tool use, knowledge work, computer operation and scientific research. Those are valuable improvements, but they do not necessarily tell us whether Opus is now a more enjoyable or distinctive prose stylist.
Among heavy Claude users, there has been a noticeable amount of disagreement about that question. Some say newer Opus models have become more verbose, argumentative or mannered, with a recognizable form of “Claude-speak.” Others still regard Claude as comfortably ahead for dialogue, editing and creative work.
Anecdotes from Reddit and X obviously do not constitute a scientific evaluation, and every major model release produces a predictable wave of people insisting that the previous version was better. Even so, the disagreement matters because it marks a change in perception.
A couple of years ago, “Claude writes best” was close to conventional wisdom among many power users. In 2026, it is much more obviously an opinion.
Claude’s writing verdict
| Category | Verdict |
| Natural prose | Still potentially excellent, but inconsistent |
| Editing existing writing | Excellent |
| Voice imitation | Very strong with examples and instructions |
| Fiction | Strong contender, no longer automatic winner |
| Long-form reasoning | Excellent |
| Research workflow | Strong and improving |
| Model mannerisms | A growing complaint among some users |
| Overall | The former champion now has to defend the title |
Claude still belongs in any serious writer’s toolkit. It remains especially good when it has strong source material, a detailed editorial brief and examples of the voice it is supposed to preserve.
What I would no longer do is assume, before comparing outputs, that Claude must be the best writer simply because Claude used to be the best writer.
ChatGPT: the writer that is becoming a newsroom
OpenAI’s strategy is less about owning the crown for individual sentences and more about owning the entire process around them.
GPT-5.6 Sol is positioned as a frontier model for complex professional work, with OpenAI emphasizing coding, research, knowledge work, design and long-running workflows. The August update also specifically pushed the model toward more focused answers and less unnecessary elaboration.
That sounds like a minor behavioral adjustment until you consider how much AI writing suffers from excessive explanation. Many models know the facts, understand the argument and can organize the material, but still feel compelled to introduce every point, restate it and then offer a tidy conclusion.
For writers, learning when not to say something is a meaningful capability improvement.
ChatGPT’s larger advantage, however, is the environment OpenAI has built around the model. Projects can hold instructions, source material and previous work. Research can extend onto the web. Documents can be developed over multiple sessions. Connected tools increasingly allow the model to move from source gathering to analysis, drafting, editing and production without forcing the user to rebuild the context at every stage.
A 1,500-word article may require finding regulatory filings, checking earlier coverage, pulling figures out of PDFs, comparing several sources, validating quotes, deciding what the story actually is, structuring the argument and then rewriting the piece several times. The prose itself is only one layer. OpenAI appears increasingly determined to absorb the rest of that workflow.
ChatGPT’s writing verdict
| Category | Verdict |
| Natural prose | Very good |
| Editing | Excellent |
| Research-heavy journalism | Excellent |
| Argument construction | Excellent |
| Voice consistency | Strong with sufficient examples |
| Workflow/tool integration | Probably the strongest overall proposition |
| Main weakness | Can still drift toward polished generic competence |
| Overall | Best one-tool choice for many professional writers |
That does not necessarily make ChatGPT the model most likely to produce the best single paragraph in a blind test.
It may make it the system in which the largest share of the actual writing job gets done.
Gemini: Google owns the library
Google’s strategic advantage is easy to understand. The company already owns much of the infrastructure through which professional information flows, and AI writing is increasingly inseparable from information retrieval.
Gemini 3.1 Pro supports more than one million input tokens and is designed for sophisticated reasoning across large collections of text, images, video, audio and PDFs. Google’s own positioning focuses heavily on reasoning, synthesis and complex problem-solving rather than literary style.
That fits the wider pattern. The frontier labs are investing enormous resources in reasoning and agency because those capabilities can be measured, productized and turned into useful software.
Gemini’s strength for writers comes from everything surrounding the model.
Google has Docs, Drive and Gmail, while NotebookLM has become one of the more compelling mainstream tools for source-grounded research. That gives Gemini a natural advantage whenever writing begins not with a blank page but with a messy pile of source material.
A novelist may care relatively little about that infrastructure. An analyst or journalist working from dozens of reports, email threads, transcripts and spreadsheets probably cares a great deal.
Gemini does not need to convince every writer that it is Shakespeare. It can be enormously useful simply by being unusually good at finding, organizing and connecting the right information before the writing begins.
Gemini’s writing verdict
| Category | Verdict |
| Raw prose | Capable, but not its clearest differentiator |
| Research synthesis | Excellent |
| Large source collections | Excellent |
| Multimodal research | Excellent |
| Google Workspace users | Potentially unmatched integration |
| Distinctive voice | Still something I would edit aggressively |
| Overall | The researcher’s AI writer |
For research-led work, that may be enough to make Gemini the best first stop even if another model handles the final prose pass.
Grok: the correspondent standing in the street
Grok is the hardest model in this comparison to judge because its latest version is barely out of the gate.
Grok 4.6 arrived on August 12, 2026, so any sweeping declaration about its long-term writing quality should be treated cautiously. xAI’s current positioning again emphasizes coding, software engineering, tool use and agentic work rather than presenting 4.6 primarily as a creative-writing model.
Where Grok genuinely differs from the others is its relationship with X.
For better and for worse, X remains one of the places where breaking news, market narratives, political arguments, memes, rumors and expert commentary emerge in real time. That makes Grok unusually useful for writers covering technology, crypto, markets, politics and internet culture.
The obvious danger is that immediacy and reliability are not the same thing.
Social media can tell you what people are talking about long before conventional sources catch up, but it can also amplify claims that are incomplete, misleading or simply false. The most useful role for Grok is therefore not to act as the final authority but to identify the live narrative, trace where claims are coming from and show you what deserves verification.
Used that way, it feels less like a copy editor and more like a reporter who happens to be standing in the middle of the crowd.
xAI has also shown more interest in writing benchmarks than some people realize. Its Grok 4.1 release specifically promoted performance on Creative Writing v3, even though the 4.6 launch has shifted the emphasis back toward coding and agentic work.
Again, the development path looks spiky rather than linear.
Grok’s writing verdict
| Category | Verdict |
| Raw prose | 4.6 is too new for a definitive call |
| Real-time awareness | Major strength |
| X-native research | Unique advantage |
| Commentary and zeitgeist | Potentially excellent |
| Creative writing pedigree | More interesting than many assume |
| Deep source-grounded writing | Requires verification discipline |
| Overall | The wildcard — and the live news desk |
And now the words themselves are getting watermarked
Just as the argument about writing quality becomes more complicated, another issue is moving rapidly into the foreground: provenance.
Anthropic has confirmed that supported new Claude models will embed an invisible, machine-readable watermark into generated text. The policy is linked to the European Union AI Act’s transparency requirements, and Anthropic says supported models launched in the EU from August 2, 2026 include the marking from launch. The company is also applying the system more broadly rather than restricting it only to European users.
According to Anthropic, the watermark is embedded into the generated text itself and can survive ordinary copying and pasting, although extensive rewriting, paraphrasing or translation may weaken or remove it.
The important detail is that the mark does not prove Claude wrote the underlying material.
Anthropic explicitly says Claude-generated markings can appear when Claude has proofread, translated, summarized or otherwise processed text that originally came from a human. (support.claude.com)
That distinction is going to matter enormously.
Consider a journalist who writes an entire article and asks Claude to remove typos and tighten a few sentences. Or a novelist who uses Claude to suggest cuts. Or a student who writes an essay unaided but asks an AI to check the grammar.
Those documents may be AI-assisted without being meaningfully AI-authored.
The technical signal may therefore tell us something useful about provenance while telling us much less about authorship.
Claude isn’t actually first
Anthropic is not starting from an empty field.
Google DeepMind already uses SynthID to embed invisible signals into AI-generated content, including text. In the Gemini app and web experience, Google says the system alters token probabilities in a way that leaves a statistical signature while preserving the meaning and apparent quality of the output. (deepmind.google)
OpenAI is moving in the same general regulatory direction, but its implementation is currently different. The company supports the EU Code of Practice on Transparency of AI-Generated Content and already uses provenance systems for supported generated images and audio.
As of August 13, 2026, however, OpenAI’s public provenance documentation does not describe an equivalent watermark for ordinary ChatGPT text. (openai.com)
xAI’s public documentation is more limited again. Grok-generated images and videos carry visible or machine-readable provenance information, but the company’s current consumer documentation does not describe an equivalent system for normal text responses. (docs.x.ai)
The AI text watermark situation
| Platform | Text watermarking status | What writers should know |
| Claude | Rolling out on supported new models | Invisible text mark can indicate Claude processed the material, not necessarily authored it |
| Gemini | Yes, via SynthID in Gemini app/web | Google already embeds an imperceptible statistical signal into generated text |
| ChatGPT | No equivalent public text implementation documented yet | OpenAI supports EU transparency rules and already marks supported images/audio |
| Grok | No public text watermark documented | xAI currently documents watermarks for generated images and video |
This table is almost certain to change, but the direction of travel is clear.
The watermark wars could change which AI writers choose
There is a strong case for provenance technology. The internet is filling rapidly with synthetic content, and platforms, publishers and researchers need better ways to understand where material came from.
Text, however, creates a particularly messy problem because writing is rarely a single act of generation.
A human can write 95% of a document and ask an AI to revise the other 5%. An AI can generate the first draft and a human can substantially rewrite it. A person can dictate an original argument and ask a model only to organize it. A newsroom can pass the same document through several editors, some human and some artificial.
At what point does the finished document become “AI-generated”?
That question sounds philosophical until institutions begin attaching consequences to the answer.
For journalists, authors, academics and students, the difference between AI-generated and AI-assisted is potentially huge. Yet watermarking technology may detect machine involvement without being able to explain what that involvement actually consisted of.
It could also create an unexpected competitive dynamic between the platforms. If one AI tool reliably marks everything it edits while another does not, writers who dislike having their workflow encoded into the finished text may begin taking that into account when choosing a model.
That does not make watermarking inherently bad. Provenance is likely to become increasingly important as synthetic media spreads.
It does mean that provenance policy could become a product feature in its own right, alongside context windows, prices and model quality.
There is also a limit to what these systems can tell us. A detectable Claude watermark cannot establish whether an argument was original, whether a factual claim is true or whether the ideas came from the human or the machine. Likewise, the absence of a watermark is not proof of human authorship.
Anthropic itself warns against making those inferences.
That caveat may become very important once schools, publishers, employers, search engines and social platforms begin building policies around these signals.
So which AI should a writer actually use?
Here is where I currently land.
| If your main job is… | Start with… | Why |
| Research-heavy article | ChatGPT | Best combination of research, reasoning, drafting and editing |
| Huge source dossier | Gemini | Google + NotebookLM is a formidable research stack |
| Creative rewrite | Claude | Still capable of excellent editorial transformation |
| Fiction | Claude, but test against ChatGPT and Grok | Claude’s crown is no longer uncontested |
| Breaking tech/crypto commentary | Grok + ChatGPT | Grok finds the live narrative; ChatGPT disciplines it |
| Editing your own prose | Claude or ChatGPT | Both can be stronger editors than autonomous authors |
| Corporate documents | ChatGPT | Broad end-to-end production workflow |
| Google Workspace-heavy work | Gemini | The integration advantage is obvious |
| Finding what the internet thinks right now | Grok | X is its distinctive weapon |
| One subscription for a professional writer | ChatGPT | Best overall utility rather than necessarily best sentence-level prose |
The table is useful, but it also hides an increasingly obvious reality: professional writers do not necessarily need to choose one model and remain loyal to it.
The better approach may be to use them as specialized tools and, where the work matters, make them check one another.
The four-model newsroom
| Stage | Best tool | What it does |
| 1. Source gathering | Gemini / ChatGPT | Find primary documents and assemble research |
| 2. Live narrative check | Grok | Surface reactions, debates and emerging claims |
| 3. Thesis development | ChatGPT | Turn information into an argument |
| 4. First draft | ChatGPT / Claude | Produce long-form prose from the research |
| 5. Voice pass | Claude / ChatGPT | Rewrite mechanical or generic passages |
| 6. Adversarial edit | A different model | Attack assumptions and identify weak sections |
| 7. Fact check | ChatGPT / Gemini | Re-open sources and verify claims |
| 8. Final human edit | You | Remove everything that sounds like a machine wrote it |
That last stage is still doing a lot of work. It may continue doing so for quite a while.
The benchmark nobody has solved: taste
AI companies understandably like benchmarks. They give researchers a way to measure progress and allow model developers to make specific claims rather than simply insisting that the new model “feels smarter.”
Writing does not fit neatly into that framework. The new generation of creative-writing benchmarks is useful precisely because researchers are trying to measure things such as style, coherence, originality, emotional impact and long-form consistency rather than relying entirely on anecdote. LitBench’s use of human preference data is especially interesting because it acknowledges the limits of asking one language model to judge another.
Even so, no benchmark can completely remove the subjective part of the equation. Taste depends heavily on context and audience. The prose that works brilliantly in a literary novel may be completely wrong for a financial report. An aggressive first-person column and a corporate board memo should not sound remotely alike, even though a generalized writing evaluator might reward many of the same qualities in both.
There is also a deeper technical tension here. Language models are exceptionally sophisticated systems for modelling what language is likely to come next. Great writers, by contrast, often become memorable precisely because they know when not to choose the expected phrase.
That does not mean AI will never become an exceptional writer. There is already enough evidence to rule out that kind of complacency.
It does suggest that the path from greater reasoning ability to greater style is not automatic. The ability to solve harder problems can improve rapidly while literary taste, humour and originality move at a much slower pace.
That may turn out to be one of the defining characteristics of the current generation of AI.
The Gilded Age verdict
The old hierarchy is breaking down. Claude still has a serious claim as one of the best models for editing, stylistic transformation and creative collaboration, but its status as the automatic prose champion is much less secure than it once appeared. Opus 5 can be significantly more capable in many measurable ways while being less appealing to some writers. There is no contradiction in that.
ChatGPT increasingly looks like the strongest overall writing system rather than necessarily the strongest pure prose model. Research, files, memory, editing, tool use and production are becoming part of the same environment, which matters enormously for professional work.
Gemini has the most obvious structural advantage when writing begins with large quantities of information. Google’s ability to connect AI with documents, email, search and NotebookLM means Gemini does not need to win every stylistic comparison to become indispensable for certain kinds of research-heavy work.
Grok remains the wildcard. Its proximity to X and the real-time internet gives it a distinctive role for news, markets, technology and cultural commentary, while xAI’s previous performance on dedicated creative-writing benchmarks suggests it should not be dismissed as simply the internet’s snarkiest chatbot.
None of the four has solved taste, and the arrival of text watermarking now adds another complication. Writers are going to have to think not only about what a model produces, but also about what traces it leaves behind and what institutions may eventually infer from those traces.
For years, it was easy to assume that beautiful prose would emerge naturally once models became sufficiently intelligent.
The reality looks more complicated. AI can become much better at mathematics without becoming funnier. It can make enormous advances in software engineering while developing conversational habits that some users find more irritating. It can improve across dozens of benchmarks without producing a correspondingly dramatic leap in prose style.
That does not suggest AI development is slowing down. If anything, it suggests we are finally getting a clearer look at the shape of machine intelligence.
Progress is uneven, and the skills humans tend to bundle together under the word “intelligence” do not necessarily improve at the same rate. Writing may turn out to be one of the best places to see that difference.

