Original title: Introducing Claude Opus 5
Article
Anthropic released Claude Opus 5 as a new default on Claude Max and a strong option for Pro users, emphasizing better cost efficiency than Opus 4.8 with the same pricing tier and improved front-and-back office behavior across coding, knowledge work, science, visual tasks, and software engineering workflows. The release claims broad gains in Frontier-Bench, FrontierCode, CursorBench, ARC-AGI 3, Zapier AutomationBench, OSWorld 2.0, genomics, financial analysis, and legal or design-heavy tasks, with frequent citations of higher pass rates, stronger tool use, reduced variance, and lower token or latency budgets compared with Opus 4.8. Internal demonstrations and customer anecdotes describe more iterative verification, bug root-cause finding, test-harness building, and handling of longer workflows with less backtracking, while also noting it remains behind Mythos 5 on offensive cybersecurity exploitation. Anthropic reports Opus 5 as more aligned and less deceptive than prior Opus and Sonnet variants, with reduced reckless-action risk, but it does not claim major frontier progress in dual-use bio and cyber exploitation. Safety policy remains restrictive: Opus 5 permits source-code vulnerability discovery more broadly than Opus 4.8, while blocking some binary and exploit-generation pathways, with fallbacks and a Cyber Verification Program for heavier defensive work. The company frames the model as broadly useful despite governance tradeoffs and adds operational updates such as mid-conversation tool changes and automatic fallback routing on safety refusals. Pricing is unchanged relative to Opus 4.8, and Fast mode remains optional at a higher rate, reinforcing the theme that token efficiency and routing strategy are now central to practical deployment. Comments and reactions, in
The discussion is sharply split between heavy early-access optimism and measurable skepticism. Supporters report practical gains in agentic coding, full-stack UI work, visual interpretation tasks, security triage, and long-running code fixes, with some praising cleaner diffs, better planning, and lower hallucination risk in specific contexts. Several contributors also note improved convenience updates, especially in Claude Code workflows, and appreciate new vulnerability-discovery support in source code at lower reasoning levels. Much of the debate centers on benchmarking trust: users compare Anthropic numbers to external papers and competitor tools, highlight gaps on OSWorld, ARC-AGI, and agentic coding metrics, and question cross-report consistency, pricing-based effort modes, and whether Opus 5 is truly ahead of Fable 5. Others criticize model behavior variance, verbosity, personality shifts, occasional overconfidence, ignored instructions, and mixed performance on real review tasks. Operational friction is a recurring concern, including account restrictions, infra instability, and confusion from rapid model/version naming and capability differences. Practical concerns also include retention-policy uncertainty around enterprise use, security-classifier opacity, and the cost of choosing among Opus/Fable/Sonnet under fast release cycles. The thread as a whole treats the launch as strategically important but not yet conclusive evidence of a clean hierarchy shift.