When you choose an AI model for a bug fix or a documentation task, the score is only the starting point. You still need to know how much useful work it completes and how much checking it needs.
This guide looks at the published GPT-6 Astra and Claude Fable 5.1 results for developers and students. You’ll also find a repeatable trial protocol, a blank evaluation log, and a cost calculator you can use on your own tasks.
The September releases have renewed the AGI debate. Alongside the evaluations, we’ll look at what Sam Altman, Dario Amodei, and Greg Brockman hope this technology will make possible. Their expectations are useful context for deciding what to build next.
Can the model fix the bug, support an answer with the right source, or produce instructions a teammate can follow? Keep that outcome in view as you compare the scores.
What changed with Astra and Fable 5.1?
GPT-6 Astra
OpenAI’s September launch emphasizes software engineering, computer use, science, and multistep professional work, with a phased rollout.
Read OpenAI’s announcementClaude Fable 5.1
Anthropic emphasizes coding, knowledge work, and sustained problem-solving. Fable is generally available. Mythos 5.1 uses the same underlying model with different safeguards and restricted access.
Read Anthropic’s announcementWhat these models accomplish also depends on their tools, instructions, memory, permissions, and review process. Choosing the model is one part of building the system.
You can improve how the system finds documents, checks a change, or asks for help. Those decisions can improve the result alongside any gain from a more capable model.
What the benchmark setup changes
Benchmarks let us compare systems using repeatable scoring rules. Check the task version and test conditions before comparing two numbers. Results can change with the setup, as the ARC-AGI example below shows.
Compare the reported benchmark scores
Terminal-Bench 4.0
View all chart data
| Benchmark | GPT-6 Astra | Fable 5.1 |
|---|---|---|
| Terminal-Bench 4.0 | 57.9 | 55.8 |
| Humanity’s Last Exam · with tools | 57.2 | 65.0 |
| ARC-AGI-2 | 95.0 | 90.0 |
Use the table to shortlist candidates. Then compare them on your own tasks, including the time you spend checking and repairing their output.
Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1, up from 24.7% for Fable 5. Its reported standard error is ±3.5–4.5 percentage points per model. The launch also notes that safeguards and fallback models affect some evaluations. See the conditions.
Which model belongs in your workflow?
Run a small pilot that compares the complete setup, including tools and human review. The steps below are a suggested protocol. This journal hasn’t run an Astra or Fable trial.
- Define a pass before generating an answer.
For a bug fix, require the relevant tests to pass, the reported behavior to be corrected, and the reviewer to accept the change. For a document answer, require the claim to be supported by the correct source. Keep critical failure rules separate from style preferences.
- Record the setup.
Keep task inputs, available documents, tool permissions, time limits, and retry budgets consistent. Record model IDs, dates, prompts, reasoning settings, and harness versions. If you can’t match a condition, name that difference in the result.
- Keep unfamiliar cases aside.
Use some cases to develop the workflow and different cases to evaluate it. Include missing information, ambiguous requests, and mistakes the system should catch. A small pilot can expose failures without establishing how reliably the system handles every possible task.
- Count the work around the model.
Log rejected attempts, retries, review time, and corrections. A fast first answer may still be expensive to turn into accepted work. Use the calculator below to make the cost of that work explicit.
- Set limits on where you’ll use it.
Check that the setup meets your quality requirements and permission boundaries before comparing cost and turnaround. Revisit the decision when the model, task, or workflow changes. A lower bill doesn’t compensate for an unacceptable error.
What does an accepted result cost?
Enter the totals from your trial. This example assumes 20 attempts, 16 accepted results, $8 in API spend, and 120 minutes of review at $30 per hour. These are illustrative values, not measured Astra or Fable results or quoted model prices.
The trial total includes $60.00 of review and correction time. A lower cost is useful only when the result meets your quality requirement.
The CSV includes your current calculation, assumptions, and limitations. It is not proof of measured model performance.
Method: accepted tasks per attempt = unique accepted task results ÷ all completed attempts, including retries. Trial cost = API spend + (review minutes ÷ 60 × hourly rate). Cost per accepted result = trial cost ÷ accepted results. The starting example is $8 + (120 ÷ 60 × $30) = $68. Then $68 ÷ 16 = $4.25.
This estimate covers API spend and review time. It excludes setup, infrastructure, subscriptions, ongoing maintenance, and the consequences of undetected mistakes. Calculations run in your browser. No inputs go to a server or persist between visits. Exporting saves a file on your device.
Download the blank evaluation log (CSV)Record task IDs, settings, attempts, outcomes, costs, and review time. Use the same pass rule for every candidate and count each successful task once. Report success after retries separately from first-attempt success.
Why this page includes the data table
The benchmark explorer keeps all three datasets in a readable HTML table, with evaluation conditions beside the bars. The controls switch between tasks. They don’t average unlike benchmarks into an overall intelligence score. Open the table to compare the exact values without relying on the bars alone.
Are we entering the AGI era?
AGI usually describes broadly human-level cognitive capability across many tasks. The exact threshold is disputed. Google DeepMind’s From AGI to ASI uses a working definition around median human performance and distinguishes it from superintelligence that exceeds large, coordinated expert groups.
Astra under two evaluation setups
GPT-6 Astra · ARC-AGI-3 Semi-Private
ARC Prize calls the result meaningful progress in generalization while explicitly declining to treat benchmark saturation as proof of AGI. Its environments have bounded rules and goals. The scoring methodology combines completed levels with action efficiency relative to human participants. “99.9%” is not a percentage of human intelligence.
Independent reporting from ACS Information Age captures the same disagreement. Researchers Toby Walsh and Rebecca Johnson question whether benchmark results justify broad claims about general intelligence.
What Altman, Amodei, and Brockman hope AI will make possible
The people building AI have ambitions that extend well beyond the next benchmark. Their essays explain those ambitions in their own words. Read them as arguments about the future, separate from the evaluations above.
Sam Altman: more people able to build
In Reflections, published in January 2025, Altman described a shift in ambition from AGI toward superintelligence. He argued that more capable tools could accelerate scientific discovery and increase prosperity, and connected gradual deployment with learning from use and giving society time to adapt.
His February 2025 essay Three Observations argues that cheaper access to useful intelligence could make more work feasible. He imagines software agents as coworkers that still need human direction and supervision. Greater prosperity, he cautions, doesn’t automatically create greater equality.
A developer can start with a smaller question: is a useful project now within reach? That might be a local-language support tool or a better way to prepare technical guidance.
Dario Amodei: faster scientific research
In his October 2024 essay Machines of Loving Grace, Amodei imagines AI exceeding leading experts across many fields and accelerating scientific and health advances. His scenario assumes that powerful AI exists first. Experiments, data, and physical constraints still affect what happens next. Today’s benchmark scores don’t establish that this scenario has arrived.
Greg Brockman: put discoveries to use
In The OpenAI Mission, co-written with Ilya Sutskever in March 2019, Brockman imagined systems that could help turn scientific breakthroughs into products and services, including more accessible healthcare. They made safety and the distribution of benefits central to that ambition. It’s a historical mission statement about what they wanted to build.
More capable models
Evaluate relevant tasks.
Dependable workflows
Review complete work.
Wider access
Make access affordable and usable.
Useful outcomes
Measure benefit for people.
What could this mean for the IT industry?
A small team could try an improvement it previously postponed. Someone who knows a local industry could test a product idea. A student could attempt a more demanding project. These are practical reasons to be interested in better AI tools.
They don’t guarantee new jobs. Some tasks will shrink, and the transition won’t be even. Training, access, and how employers manage the change will affect who benefits.
If you’re deciding what to learn or build, consider these areas:
Build the whole workflow
Consider a support agent that gathers evidence, proposes a fix, and hands the change to a reviewer. Someone has to connect the systems, set permissions, write instructions, and handle failures. That’s engineering work around the model.
Test the awkward cases
Build tests, logs, security boundaries, and a rollback plan around the demo. Check the tasks your organization actually handles, including the ones that fail halfway through. Evaluation and secure integration are useful skills to develop here.
Bring expertise to the interface
Spend time with the people who’ll use the system. How does a clinic schedule appointments? What happens when a warehouse order goes wrong? How does a teacher prepare a lesson? Those details help you decide what the software needs to do.
Try a focused product idea
A small team can use AI to draft prototypes and compare designs before committing to a larger build. Look for recurring problems in local businesses, underserved languages, or a profession you know. Find someone who wants the problem solved, then test the prototype with them.
Junior developers still need to debug, model data, read unfamiliar code, and explain tradeoffs. Use an assistant during practice, then investigate and run the code it produces. Can you spot an incomplete answer and work out what’s missing?
Experienced engineers can support that learning by reviewing the reasoning behind a change, setting realistic exercises, and giving people responsibility gradually.
A 90-day project plan
Pick a manageable project and try this learning plan. The goal is to gather evidence about what works, with enough time to investigate what doesn’t.
- DAYS 01–30
Understand one workflow
Choose a recurring task. Record how it works today, where errors occur, and who checks the outcome. Build a baseline with a handful of representative examples. Learn the underlying tools well enough to spot an incorrect result.
- DAYS 31–60
Build a small assistant you can review
Connect only the data and tools it needs. Define the actions it can take and the points that need human review. Add tests, logs, and a straightforward way to undo a change. Compare model options on the same examples.
- DAYS 61–90
Test it and write up the results
Test with real users and unfamiliar cases. Report failures as clearly as successes. Write up the problem, your decisions, evidence of improvement, and remaining limits for your portfolio. Explain what the results support, even if the improvement is modest.
Beyond AGI: ASI and the singularity
Artificial superintelligence, or ASI, concerns capabilities beyond human-level general intelligence. The singularity imagines technological change accelerating beyond our ability to predict it, potentially through AI improving AI. DeepMind’s research report explores possible routes and uncertainties.
Four possible paths beyond AGI
Scale
Expand the resources available to capable systems.
New methods
Find more effective algorithms or architectures.
Self-improvement
Use AI to help improve future AI systems.
Agent collectives
Coordinate many systems on larger problems.
Altman’s June 2025 essay The Gentle Singularity imagines AI-assisted research helping improve later AI systems. He distinguishes that feedback from a system autonomously rewriting itself. Whether that progress can sustain itself remains something to demonstrate.
More accessible learning, better research tools, and software for overlooked communities are goals worth working toward. Bring educators, designers, security specialists, and domain experts into that work alongside researchers and software engineers. Their knowledge helps decide what to build, who can use it, and how to check that it works.
Common questions
Do GPT-6 Astra and Claude Fable 5.1 prove that AGI is here?
No single result settles it. These releases show strong capabilities, but researchers still disagree about the AGI threshold. Check which task was tested, under what conditions, and how reliably the system performed.
Is an ARC-AGI score a percentage of human intelligence?
No. ARC-AGI-3 combines task completion and action efficiency against a human baseline. The percentage describes that score. It isn’t an IQ score or a measure of every human ability.
How are Claude Fable 5.1 and Mythos 5.1 different?
Anthropic describes them as the same underlying model with different safeguards and access arrangements. Fable is generally available. Mythos is offered through restricted, trusted-access programs.
What should a student or junior developer learn first?
Start with programming fundamentals, debugging, data literacy, and communication. Use AI on a small project, then test its output and explain how it works. You need enough understanding to catch a plausible but incorrect answer.
When will ASI or the singularity arrive?
There is no verified date. Build adaptable skills and reliable systems as capabilities develop.
Choose one task.
Keep a record of what happens.
Use the evaluation log and calculator on a task you already understand. Record the accepted work, corrections, and time spent. Those results will tell you where to try the model next.
Sources & further reading
Benchmark source check: September 7, 2026. Leadership sources added and reviewed September 8, 2026. Their original publication dates appear below. Forecasts are attributed to their authors. The product framework, proposed trials, and learning plan are this journal’s analysis.
- OpenAI: GPT-6 AstraRelease announcement and comparative benchmark table.
- Anthropic: Claude Fable 5.1 and Mythos 5.1Capabilities, availability, and evaluation conditions.
- ARC Prize: GPT-6 Astra on ARC-AGI-3Independent evaluation and harness differences.
- ARC-AGI-3 scoring methodologyHow completion and efficiency contribute to the score.
- Google DeepMind: From AGI to ASIResearch overview. read the full paper.
- ACS Information Age: The AGI claim and expert disagreementIndependent reporting on the debate.
- Sam Altman: ReflectionsJanuary 2025. Primary account of his outlook and deployment approach.
- Sam Altman: Three ObservationsFebruary 2025. His economic and social expectations.
- Dario Amodei: Machines of Loving GraceOctober 2024. Conditional benefits scenario.
- Greg Brockman & Ilya Sutskever: The OpenAI MissionMarch 2019. Historical mission statement.
- Sam Altman: The Gentle SingularityJune 2025. His view of progress toward superintelligence.
Reporting context: The Guardian reports Brockman’s AGI-era framing of Astra’s launch. The StartupFortune discussion highlights disagreement over definitions. Both are secondary accounts. Only the public preview of Alex Heath’s Sources article was accessible. Its subscriber-only body is not used as evidence here.
Additional reading: The Verge · Bloomberg. These linked articles were not used to establish the factual claims or chart values above.
Claps, saves, topic follows, and comment previews are for this visit only. Comments are not published.
