Anthropic’s latest model, Claude Opus 5.5, appears to be writing noticeably differently than its predecessor, according to new data from the AI benchmarking tool Arena. The changes touch on some of the most recognizable markers of AI-generated text, including a sharp drop in em dash usage.
Arena analyzed high-reasoning Text Arena responses from August and September 2026, comparing Opus 5 and Opus 5.5 across a dozen writing measures. It found that 10 of 12 metrics shifted toward what it characterizes as more natural, less obviously machine-generated prose.
Key stylistic shifts
- Em dash frequency fell from 15.2 per 1,000 words in Opus 5 to just 0.8 in Opus 5.5, a reduction of roughly 95%.
- Semicolon usage dropped from 6.10 to 1.64 per 1,000 words.
- Average sentence length shortened from 12.14 words to 10.03 words.
The em dash has become something of a shorthand indicator for detecting AI-written content, so a near-elimination of it in Opus 5.5 could make outputs harder to flag as machine-generated using simple heuristics. Simpler wording and shorter sentences point to a broader effort by Anthropic to make Claude’s writing read less like formulaic AI slop.
There is a tradeoff, however. Opus 5.5 is more verbose overall, with average response length increasing from 453 words to 481 words, making it the longest-writing Opus model Arena has measured so far.
Why it matters for security teams
While this isn’t a vulnerability disclosure, the shift has relevance for defenders who rely on stylistic fingerprints, such as em dash and semicolon frequency or sentence structure, to help identify AI-generated phishing emails, disinformation, or fraudulent content. As models like Claude Opus 5.5 move away from these patterns, detection approaches built around surface-level writing quirks may need to be reevaluated. The broader trend suggests that distinguishing human from AI-authored text based on stylistic tells alone will likely keep getting harder as vendors iterate on their models.
