Hero title card headed 'Claude Opus 5.5 and the watermelon test', subtitled 'The most readable Claude yet — and it still buries the point', with three badges reading Flesch-Kincaid 6.95, Reading Ease 68.4, and September 22, 2026, and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Claude Opus 5.5 and the Watermelon Test: Anthropic Fixed the Prose, but the Point Is Still Buried

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

One of the samples that travelled furthest out of launch week was a short story about a watermelon. A few hundred words, no benchmark attached, no chart — just a model asked to write something small and getting it mostly right. That is a strange artifact to sit near the centre of a flagship release, but it points at the thing the vendor chose to lead with when it shipped Claude Opus 5.5 on September 22, 2026: not another point on Terminal-Bench, but prose. The company's framing is that this model writes less "Claudish" — its word for the formulaic, over-hedged, throat-clearing register that Claude Opus 5 drew complaints for, and that Claude Fable 5.1, GPT-6 Astra and GPT-5.6 Sol have each been measured against since. So the question worth answering is not whether the demo was charming. It is whether the prose measurably improved, by how much, and where it still fails.

What Anthropic says it fixed

Screenshot of Anthropic's Claude Platform release-notes page, showing the Claude Opus 5.5 entry dated September 22, 2026, with the model ID claude-opus-5-5, a 1M-token context window and 128K maximum output.

Anthropic's claim is specific enough to test. Opus 5.5, it says, puts the most important information first, uses less jargon, and follows a user's stated writing rules more closely than previous Opus models. The supporting material is vendor-reported and unreproduced: internal fact-checking tests that Anthropic says passed 16 of 18 attempts, against zero of 18 for both Claude Fable 5.1 and Claude Opus 5. Customer quotes in the launch material point the same direction — Ramp's John Ruelas describing output that "writes like a good colleague," Box reporting roughly 40% less verbosity on its own workloads.

Two things are worth separating here. The first is that "less Claudish" is a real, nameable failure mode, not a marketing invention. Practitioners have been complaining about condensed, badly structured Claude prose since the post-4.6 Opus generations, and the complaint is specific: too much preamble, too many restatements of the prompt, a habit of arriving at the point two paragraphs after the reader needed it. The second is that Anthropic's own numbers for having fixed it are exactly that — Anthropic's own numbers. The 16-of-18 figure has no independent replication, and "40% less verbosity" is a customer's internal measurement, not a published methodology.

The release also carries a non-writing change worth knowing: Opus 5.5 is the first Opus model shipping with cybersecurity, biology and frontier-LLM-development safeguards matching those on Claude Fable 5.1. When a request trips one of them, Anthropic says it transparently routes the request to another model rather than refusing outright.

The numbers that aren't Anthropic's

A single-column readability scoreboard for Claude Opus 5.5: Flesch-Kincaid grade 6.95 versus Opus 5 at 7.97; Reading Ease 68.4 versus 61.35; Fable 5.1 and GPT-6 Astra at 7.43 and 7.52 grade; GPT-5.6 Sol at 8.74 grade and 56.62 ease; ranked 4th of 5 tested as an editor; and an AA Intelligence Index of 58, first of 212 models, with a sourcing footer.

The most useful independent read on the writing claim comes from Every's Vibe Check, which ran the same prompts across five frontier models and scored the output with two standard readability instruments. On that prompt set, Claude Opus 5.5 came out as the most readable of the group — and the gap to the model it replaced is the story:

• Flesch-Kincaid grade level — Claude Opus 5.5 at 6.95, against Claude Opus 5 at 7.97

• Flesch Reading Ease — 68.4 for Opus 5.5, against 61.35 for Opus 5

• Claude Fable 5.1 — 7.43 grade level, 66.09 Reading Ease

• GPT-6 Astra — 7.52 grade level, 64.55 Reading Ease

• GPT-5.6 Sol — 8.74 grade level, 56.62 Reading Ease

Read the two instruments together, because they move in opposite directions and that is the point: a lower grade level and a higher ease score both mean simpler prose. Opus 5.5 lands about a full grade below Opus 5, roughly half a grade below both Claude Fable 5.1 and GPT-6 Astra, and nearly two grades below GPT-5.6 Sol. In plain terms, a seventh-grade reading level instead of an eighth-grade one.

That is a real improvement and it is also a narrow one. This is one outlet's prompt set, not a benchmark with a fixed task list and a leaderboard behind it. It tells you the direction of travel and roughly the size of the step. It does not tell you that Opus 5.5 writes well on your work, and Every's own qualitative notes make clear the reviewer did not think so either.

Where it still buries the point

The same review that produced the readability numbers is blunt about the failure that survived the fix. Asked to open an essay, Opus 5.5 took 37 to 39 sentences to cover what the reviewer covered in 21, and spent a full paragraph on a point the reviewer made in eight words. The reviewer's summary is that the model "still buries the point" — and the sharper version of the complaint is that it can correctly diagnose what a paragraph needs, then produce a revision that ignores its own diagnosis.

As an editor it fared worse than as a writer. Ranked against the other models on editing tasks, Opus 5.5 came in behind Claude Opus 5, GPT-5.6 Sol and GPT-6 Astra — the model is better at producing readable sentences than at recognising which sentence should have come first. The reviewer's verdict lands in a sentence worth quoting because it is the most useful thing in the whole review: "I'm happier with it as a collaborator than with the prose it produces."

The practical read follows directly from that. Draft with it, and draft at high effort or above — the lower effort settings are where the padding gets worst. Supply your own examples rather than asking it to invent them. Then move the point to the top yourself, or hand the draft to something that will. What Opus 5.5 gives you is a faster path to a clean first draft and a genuinely better collaborator for the passes after it. It does not give you an editor.

What the extra words cost

Screenshot of the Artificial Analysis model page for Claude Opus 5.5, showing an Intelligence Index of 58 ranked first of 212 models, a cost rank of 93 of 212 at $4.00 per million input and $20.00 per million output with a 95% cache discount and $5.98 per Index task, and a verbosity rank of 95 of 212 at 260 million output tokens against a median of 88 million.

There is a cost dimension to the verbosity problem that the writing reviews mostly skip, and it cuts against the launch narrative. Artificial Analysis, which runs its own evaluation rather than relying on vendor figures, places Claude Opus 5.5 first of 212 models on its Intelligence Index at a score of 58 — but 95th of 212 on verbosity, having generated 260 million output tokens across that Index run against a median of 88 million. It cost $5.98 per Intelligence Index task to evaluate, and $8,708.20 in total.

Set that beside Anthropic's own efficiency claim and the two do not obviously agree. The vendor-reported position is that Opus 5.5 uses fewer tokens and generates output more than 30% faster than Opus 5. Artificial Analysis's independent run found the model very verbose relative to its peers. Both can be true at once — different task mixes, different effort settings, different definitions of "fewer tokens" — but nobody should read the launch framing as an established fact about token economy. On price the picture is cleaner: $4 per million input and $20 per million output, down from $5/$25, with a 95% cache discount in the Artificial Analysis figures.

If you are paying per output token, verbose prose is not a style preference. It is the line item.

Running it against what you already have

One honesty note before the practical part: Claude Opus 5.5 is not one of our routes. We do not host it, and nothing below should be read as a claim that we do. You get it through Anthropic's own API and several third-party platforms.

What is on OrcaRouter today is the model it replaced and the tier above it. Claude Opus 5 is on OrcaRouter today, and so is Claude Fable 5.1, both at the provider's list price with 0% markup passed through — which means a vendor price change lands on our side the same day rather than whenever a reseller gets around to it. If you are trying to decide whether Opus 5.5's prose is worth the move, that is a useful pairing: you can run the comparison on one key without a second contract or a code change, because both models are one API for 200+ models away on the same endpoint.

The other reason to keep a second model in the loop is the failure mode this article is about. A model that buries the point is a model you want to be able to route around mid-task, and automatic failover exists precisely so a bad generation does not become a stalled pipeline. If you would rather not pick one model at all, the routing DSL composes several into a single call and model fusion runs a panel of models answering together — which is a reasonable shape for editing work, where the failure is judgement rather than capability.

Who should act, and who should wait

If your complaint about Claude Opus 5 was that you had to rewrite its openings, Opus 5.5 is a real fix for roughly one grade level of that problem, and the readability data says the improvement is measurable rather than vibes. If your complaint was that it takes three paragraphs to get to the point, nothing in the independent evidence says that changed. Those are different complaints and the launch coverage has been treating them as one.

For creative work specifically, the watermelon story is not evidence of anything except that people reach for small, warm, low-stakes prompts when they want to feel out a new model. That is a reasonable thing to do. It is also exactly the setting where a tendency to bury the point is least visible, because nobody is reading a watermelon story for the thesis. Put the model in front of a document where the reader needs the conclusion in the first two sentences and you will find out which of the two problems you actually had.

The thing to watch next is whether the next model in the family — Sonnet 5.5 and Haiku 5.5 are both expected in the coming weeks — inherits the readability gain, and whether anyone publishes a readability number for editing rather than drafting. The drafting number is the easy one. The editing number is where Opus 5.5 still ranks behind three of its competitors.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily