This was Opus week. On May 28 Anthropic shipped Claude Opus 4.8, and the number that matters is not the agentic-coding benchmark (up to 69.2%) - it is the honesty one: the new model is roughly four times less likely than 4.7 to let a flaw in its own code pass unremarked. That is the gap between an agent you supervise line by line and one you can leave running.

Anthropic held pricing flat at $5 / $25 per million tokens, added a fast mode that is 2.5x quicker and 3x cheaper than before, and - the same day - closed a $65B raise. The flagship got better and cheaper to run in the same breath. Which made it the perfect week to ask whether the flagship is still the only answer: we ran Opus 4.8 head to head against typed on fourteen medium-to-hard coding tasks, graded against real test suites. They tied, 14-14. The quality gap has narrowed to a billing decision.

Below: the Opus 4.8 details and its same-day Copilot GA, a busy Claude Code release train, and a broader week where OpenAI put Codex hands-on-keyboard inside Windows, MiniMax teased a 15.6x decoding speedup, and Mistral pivoted to a full-stack enterprise play.


From the Yaw blog

Anthropic this week

Claude Code this week

The broader week