It sounds like a contradiction: a more intelligent AI that writes worse. Yet that is exactly what happened with Claude, according to an explanation from an Anthropic engineer that has been circulating among AI researchers and developers. The core issue comes down to how training signals interact with creative quality, and the answer reveals a tension that sits at the heart of modern large language model development.

Smarter Models, Blander Output

As Anthropic pushed Claude's reasoning and factual accuracy forward, the feedback mechanisms used during training inadvertently pulled writing quality in the other direction. Reinforcement learning from human feedback, the dominant method for aligning AI behavior, tends to reward outputs that feel safe, clear, and broadly acceptable to a wide pool of evaluators. Creative writing, by contrast, rewards risk, specificity, and voice. Those two incentive structures can pull in opposite directions. Evaluators rating responses for helpfulness and accuracy are not necessarily selecting for the kind of prose that makes a reader want to keep reading.

Key Facts

  • Claude's benchmark scores on reasoning and knowledge tasks improved while some users reported flatter, more generic writing output.
  • Human feedback used in RLHF training tends to favor safe, agreeable responses over stylistically bold ones.
  • The problem is not unique to Claude; it reflects a structural tension in how frontier AI models are trained.
  • Anthropic has acknowledged the issue and signaled that writing quality is an active area of focus.
  • Developers and power users were among the first to notice the shift, particularly in longer-form creative tasks.

The engineer's explanation pointed to how averaged human preferences flatten outliers. A sentence that is genuinely surprising or stylistically distinctive might score lower in aggregate evaluations simply because it divides opinion, even if skilled writers would rate it highly. Over thousands of training iterations, those small penalties accumulate. The model learns to produce writing that offends no one and excites no one either. This is sometimes called "mode collapse" in practice, though the technical definition is narrower. The end result is prose that is competent but forgettable.

The problem is that what humans rate as 'good' in a quick evaluation is not always what makes writing actually worth reading. There's a gap between legibility and quality that training can easily fall into.Anthropic engineer, via The Decoder
Claude AI Handboek by Leon Tindemans
Get the Claude AI Handboek
458 pages on getting more out of Claude, by AI expert Leon Tindemans. A printed book, written in Dutch, shipped worldwide with track and trace.
View the book →

A Known Trade-Off With No Easy Fix

This tension is not new, and it is not specific to Claude. Anyone following the latest Claude AI news will have seen similar complaints surface around other major model updates industry-wide. But the Anthropic engineer's candid account is notable for how directly it names the mechanism. Rather than attributing the regression to a bug or an oversight, the explanation treats it as a structural feature of current training pipelines, one that requires deliberate effort to counteract rather than a simple patch.

Addressing the problem means either developing evaluation methods that better capture writing quality, or finding ways to weight the feedback of more discerning reviewers more heavily. Neither approach is straightforward. Training data curation, specialized writing evaluators, and targeted fine-tuning are all options that teams across the industry are exploring. Anthropic's work on Claude's model family suggests the company is aware of this gap and is working to close it, though no specific timeline or method has been announced publicly.

For users who rely on Claude for creative or editorial work, the explanation offers some reassurance that the regression is understood rather than mysterious. It also hints that future updates could restore some of the expressive range that earlier versions demonstrated. The challenge is doing that without sacrificing the reasoning gains that make the newer models more useful for technical and analytical tasks.

The episode is a useful reminder that benchmark scores measure specific capabilities. A model that scores higher on reasoning tasks is genuinely better at reasoning. That improvement does not automatically carry over to style, voice, or narrative craft. Those qualities require their own training signals, their own evaluation frameworks, and their own dedicated attention. For now, the gap between intelligence and eloquence in AI systems remains very much open.

Further reading: Learn more about Claude's model family, read our background on Anthropic, or browse the latest Claude AI news.