/ THE SHORT ANSWER
Run a controlled evaluation on a small set of representative workflows. Score output quality, completion rate, human correction, latency, policy refusals, integration effort and total cost per accepted result. OpenAI reports stronger results on several computer-use and agentic benchmarks, while Anthropic reports lower typical token-billed cost than its previous Fable model and a distinct enterprise safeguard model. The benchmark harnesses and operating policies differ, so neither provider's launch material proves that one model is best for every business.
- 01GPT-6 Astra and Claude Fable 5.1 were announced on 3 September and 1 September 2026 respectively.
- 02OpenAI positions Astra for computer use and professional workflows, while Anthropic positions Fable 5.1 for agentic, coding and enterprise work.
- 03Provider benchmark results are useful signals, but different harnesses, tools and policies limit direct comparison.
- 04Cyber safeguards can affect legitimate testing workflows, so refusal and escalation behaviour belongs in the evaluation.
- 05Total cost should be measured per accepted business result, not only per token.
/ dotSuper point of view
Model choice has become an operating-design decision. The winner is the system that completes valuable work with the least hidden review, policy and migration cost.
What changed
OpenAI released GPT-6 Astra on 3 September 2026 and described it as a model built for computer use, professional work, forms, CRM systems, research, documents, spreadsheets and website creation. The company says it is rolling Astra out across paid ChatGPT plans and through its API, Azure and Amazon Bedrock. OpenAI lists standard API prices of $10 per million input tokens and $50 per million output tokens.
Anthropic released Claude Fable 5.1 on 1 September. It says the model is generally available, while Claude Mythos 5.1 uses the same model with different safeguards through trusted programs. Anthropic estimates that Fable 5.1 reduces typical token-billed cost by 25 percent compared with Fable 5, with larger reductions possible on highly agentic work.
Both providers published benchmark tables. OpenAI reports Astra at 64.6 percent on Terminal-Bench Science and 57.9 percent on Terminal-Bench 4.0, compared with figures of 52.6 percent and 55.8 percent for Fable 5.1. Anthropic's own release reports those Fable results alongside other coding and automation tests. OpenAI also notes method modifications on some comparisons. These are provider-reported results, not one independently controlled head-to-head test.
- Astra expands the emphasis on computer use and end-to-end professional tasks.
- Fable 5.1 focuses on agentic efficiency and enterprise safeguard choices.
- Pricing, tool use and safety policy now shape the model experience as much as raw generation quality.
Why it matters to businesses
A benchmark advantage does not automatically become an operating advantage. A model can score well and still create extra work through inconsistent tool calls, hard-to-review output, slow responses or excessive escalation. Conversely, a model with a slightly lower public score may be the better choice if it fits a company's data boundaries, interface, review process and existing cloud environment.
Cyber policy is especially relevant for security teams. OpenAI says Astra reaches its Critical capability threshold for cybersecurity and applies layered safeguards. Its separate safety note reports stronger refusal of disallowed cyber requests, while acknowledging that legitimate defensive work can experience friction. Anthropic says Fable 5.1 reduces false positives compared with Fable 5, while continuing to restrict penetration testing and exploit-generation tasks outside trusted arrangements.
Procurement should therefore compare the complete service, not an abstract model. That includes data handling, regional availability, rate limits, tool reliability, support, audit evidence, model-change notices and an exit path. The business case rests on accepted output per rupee or dollar, plus the cost of the humans and controls needed to make that output safe.
| Dimension | What to measure | Why it matters |
|---|---|---|
| Work quality | Accepted outputs without correction | Shows real usefulness, not stylistic preference |
| Control | Refusals, escalations and policy exceptions | Reveals operational friction and risk |
| Cost | Model, tool, review and retry cost per result | Captures the actual unit economics |
| Reliability | Completion, latency and recovery from tool failure | Determines whether a workflow can run consistently |
| Portability | Prompts, tools and data that can move | Limits future switching cost |
What to do next
Choose three to five workflows where model performance has a measurable business consequence. Examples include preparing a sales brief, updating a CRM record, reconciling a spreadsheet or researching a supplier. Build a small evaluation set using real but appropriately protected inputs, including normal cases, ambiguous cases and failure cases.
Give each model the same tools, instructions and acceptance criteria where possible. Record output quality, completion time, retries, refusals, reviewer minutes and total API cost. Keep provider-specific optimisations as a second test so that a fair baseline is not confused with the best achievable implementation.
Decide by workflow rather than forcing one model across the company. A routed system can use different models for different tasks, provided the governance remains understandable. Store prompts, evaluation examples and decision criteria outside the provider interface so the organisation retains its operating knowledge.
- Measure cost per accepted result, including human review and retries.
- Test policy behaviour with legitimate edge cases before security teams depend on the model.
- Document a rollback and provider-switching path before production use.
- Re-run the evaluation when a model, price or safeguard policy changes.
What this page cannot conclude
- 01The performance figures in this briefing come from provider releases and are not an independent dotSuper benchmark.
- 02Benchmark harnesses, prompts, tools, token budgets and safety settings can differ, which limits direct comparison.
- 03Quoted prices and availability can vary by tier, region, platform and later provider changes.
- 04A workflow evaluation should use protected data and appropriate legal, security and procurement review.
Sources
- 01Introducing GPT-6 AstraOpenAI · accessed Sep 7, 2026
- 02Our path to AstraOpenAI · accessed Sep 7, 2026
- 03Claude Fable and Mythos 5.1Anthropic · accessed Sep 7, 2026
Our editorial standard · Found an error? Send a correction with its source.