Skip to content
The Journal
All stories
ARTIFICIAL INTELLIGENCE15 min read

GPT-6 Astra & Claude Fable 5.1:
benchmarks and the AGI debate.

Read the results, compare the leaders’ expectations, and test what these models could do in your own work.

Research analysis · Published benchmark dataUpdated Method & limitations

When you choose an AI model for a bug fix or a documentation task, the score is only the starting point. You still need to know how much useful work it completes and how much checking it needs.

This guide looks at the published GPT-6 Astra and Claude Fable 5.1 results for developers and students. You’ll also find a repeatable trial protocol, a blank evaluation log, and a cost calculator you can use on your own tasks.

The September releases have renewed the AGI debate. Alongside the evaluations, we’ll look at what Sam Altman, Dario Amodei, and Greg Brockman hope this technology will make possible. Their expectations are useful context for deciding what to build next.

Start with a task you can check

Can the model fix the bug, support an answer with the right source, or produce instructions a teammate can follow? Keep that outcome in view as you compare the scores.

01 / THE RELEASES

What changed with Astra and Fable 5.1?

GPT-6 Astra

OpenAI’s September launch emphasizes software engineering, computer use, science, and multistep professional work, with a phased rollout.

Read OpenAI’s announcement

Claude Fable 5.1

Anthropic emphasizes coding, knowledge work, and sustained problem-solving. Fable is generally available. Mythos 5.1 uses the same underlying model with different safeguards and restricted access.

Read Anthropic’s announcement

What these models accomplish also depends on their tools, instructions, memory, permissions, and review process. Choosing the model is one part of building the system.

You can improve how the system finds documents, checks a change, or asks for help. Those decisions can improve the result alongside any gain from a more capable model.

02 / THE EVIDENCE

What the benchmark setup changes

Benchmarks let us compare systems using repeatable scoring rules. Check the task version and test conditions before comparing two numbers. Results can change with the setup, as the ARC-AGI example below shows.

FIG. 01

Compare the reported benchmark scores

Terminal-Bench 4.0

GPT-6 Astra57.9%
Claude Fable 5.155.8%
Source: OpenAI’s launch table, September 2026. The chart uses the highest reported score at any reasoning effort. Research and API test conditions may differ from production use. Each test stands on its own. There is no combined score.
View all chart data
Reported benchmark scores (%)
BenchmarkGPT-6 AstraFable 5.1
Terminal-Bench 4.057.955.8
Humanity’s Last Exam · with tools57.265.0
ARC-AGI-295.090.0

Use the table to shortlist candidates. Then compare them on your own tasks, including the time you spend checking and repairing their output.

ANOTHER SIGNAL TO WATCH

Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1, up from 24.7% for Fable 5. Its reported standard error is ±3.5–4.5 percentage points per model. The launch also notes that safeguards and fallback models affect some evaluations. See the conditions.

03 / MAKE A DECISION WITH EVIDENCE

Which model belongs in your workflow?

Run a small pilot that compares the complete setup, including tools and human review. The steps below are a suggested protocol. This journal hasn’t run an Astra or Fable trial.

  1. Define a pass before generating an answer.

    For a bug fix, require the relevant tests to pass, the reported behavior to be corrected, and the reviewer to accept the change. For a document answer, require the claim to be supported by the correct source. Keep critical failure rules separate from style preferences.

  2. Record the setup.

    Keep task inputs, available documents, tool permissions, time limits, and retry budgets consistent. Record model IDs, dates, prompts, reasoning settings, and harness versions. If you can’t match a condition, name that difference in the result.

  3. Keep unfamiliar cases aside.

    Use some cases to develop the workflow and different cases to evaluate it. Include missing information, ambiguous requests, and mistakes the system should catch. A small pilot can expose failures without establishing how reliably the system handles every possible task.

  4. Count the work around the model.

    Log rejected attempts, retries, review time, and corrections. A fast first answer may still be expensive to turn into accepted work. Use the calculator below to make the cost of that work explicit.

  5. Set limits on where you’ll use it.

    Check that the setup meets your quality requirements and permission boundaries before comparing cost and turnaround. Revisit the decision when the model, task, or workflow changes. A lower bill doesn’t compensate for an unacceptable error.

What does an accepted result cost?

Enter the totals from your trial. This example assumes 20 attempts, 16 accepted results, $8 in API spend, and 120 minutes of review at $30 per hour. These are illustrative values, not measured Astra or Fable results or quoted model prices.

Count every completed attempt, including failures and retries.
Agree on a pass rule before the trial. Count each accepted result once.
Include charges for failed attempts and retries.
Include checking and correction for all attempts, even rejected results.
Your assumption for valuing review time, not a salary estimate.
Accepted tasks / attempts80%
API + review cost$68.00
Cost per accepted result$4.25

The trial total includes $60.00 of review and correction time. A lower cost is useful only when the result meets your quality requirement.

The CSV includes your current calculation, assumptions, and limitations. It is not proof of measured model performance.

Method: accepted tasks per attempt = unique accepted task results ÷ all completed attempts, including retries. Trial cost = API spend + (review minutes ÷ 60 × hourly rate). Cost per accepted result = trial cost ÷ accepted results. The starting example is $8 + (120 ÷ 60 × $30) = $68. Then $68 ÷ 16 = $4.25.

This estimate covers API spend and review time. It excludes setup, infrastructure, subscriptions, ongoing maintenance, and the consequences of undetected mistakes. Calculations run in your browser. No inputs go to a server or persist between visits. Exporting saves a file on your device.

Download the blank evaluation log (CSV)

Record task IDs, settings, attempts, outcomes, costs, and review time. Use the same pass rule for every candidate and count each successful task once. Report success after retries separately from first-attempt success.

Why this page includes the data table

The benchmark explorer keeps all three datasets in a readable HTML table, with evaluation conditions beside the bars. The controls switch between tasks. They don’t average unlike benchmarks into an overall intelligence score. Open the table to compare the exact values without relying on the bars alone.

04 / THE BIG QUESTION

Are we entering the AGI era?

AGI usually describes broadly human-level cognitive capability across many tasks. The exact threshold is disputed. Google DeepMind’s From AGI to ASI uses a working definition around median human performance and distinguishes it from superintelligence that exceeds large, coordinated expert groups.

FIG. 02

Astra under two evaluation setups

GPT-6 Astra · ARC-AGI-3 Semi-Private

Standard harnessMax reasoning effort
62.7%
Provider AdapterHigh reasoning effort
99.9%
ARC Prize’s independent evaluation. The adapter preserves private reasoning state and supports context compaction. These results use different harnesses and reasoning settings.

ARC Prize calls the result meaningful progress in generalization while explicitly declining to treat benchmark saturation as proof of AGI. Its environments have bounded rules and goals. The scoring methodology combines completed levels with action efficiency relative to human participants. “99.9%” is not a percentage of human intelligence.

Independent reporting from ACS Information Age captures the same disagreement. Researchers Toby Walsh and Rebecca Johnson question whether benchmark results justify broad claims about general intelligence.

05 / FROM CAPABILITY TO BENEFIT

What Altman, Amodei, and Brockman hope AI will make possible

The people building AI have ambitions that extend well beyond the next benchmark. Their essays explain those ambitions in their own words. Read them as arguments about the future, separate from the evaluations above.

Sam Altman: more people able to build

In Reflections, published in January 2025, Altman described a shift in ambition from AGI toward superintelligence. He argued that more capable tools could accelerate scientific discovery and increase prosperity, and connected gradual deployment with learning from use and giving society time to adapt.

His February 2025 essay Three Observations argues that cheaper access to useful intelligence could make more work feasible. He imagines software agents as coworkers that still need human direction and supervision. Greater prosperity, he cautions, doesn’t automatically create greater equality.

A developer can start with a smaller question: is a useful project now within reach? That might be a local-language support tool or a better way to prepare technical guidance.

Dario Amodei: faster scientific research

In his October 2024 essay Machines of Loving Grace, Amodei imagines AI exceeding leading experts across many fields and accelerating scientific and health advances. His scenario assumes that powerful AI exists first. Experiments, data, and physical constraints still affect what happens next. Today’s benchmark scores don’t establish that this scenario has arrived.

Greg Brockman: put discoveries to use

In The OpenAI Mission, co-written with Ilya Sutskever in March 2019, Brockman imagined systems that could help turn scientific breakthroughs into products and services, including more accessible healthcare. They made safety and the distribution of benefits central to that ambition. It’s a historical mission statement about what they wanted to build.

AN ORIGINAL FRAMEWORKFrom capability to everyday value
  1. More capable models

    Evaluate relevant tasks.

  2. Dependable workflows

    Review complete work.

  3. Wider access

    Make access affordable and usable.

  4. Useful outcomes

    Measure benefit for people.

Hashan’s framework for assessing progress: test the model, check the complete workflow, and consider who can use it. The diagram connects those decisions to useful outcomes. It does not predict when they will happen.

What this means to me as a developer

I’m drawn to Altman’s hope that more people can act on their ideas. My own Online Web Toolkit began as a Markdown-to-PDF CLI for documentation I needed in my engineering work. It grew into a shared website. I care about that step: turning something technically possible into something another person can use.

Consider an AI assistant for release documentation. Check whether a teammate could follow its setup steps and examples without the author beside them. Do its links point to the relevant changes? How much needs correcting? Count that review time alongside drafting time. This is a proposed evaluation, not a deployed assistant’s results.

The trial protocol gives you a way to record those checks before choosing a model.

06 / THE HUMAN OPPORTUNITY

What could this mean for the IT industry?

A small team could try an improvement it previously postponed. Someone who knows a local industry could test a product idea. A student could attempt a more demanding project. These are practical reasons to be interested in better AI tools.

They don’t guarantee new jobs. Some tasks will shrink, and the transition won’t be even. Training, access, and how employers manage the change will affect who benefits.

If you’re deciding what to learn or build, consider these areas:

01

Build the whole workflow

Consider a support agent that gathers evidence, proposes a fix, and hands the change to a reviewer. Someone has to connect the systems, set permissions, write instructions, and handle failures. That’s engineering work around the model.

02

Test the awkward cases

Build tests, logs, security boundaries, and a rollback plan around the demo. Check the tasks your organization actually handles, including the ones that fail halfway through. Evaluation and secure integration are useful skills to develop here.

03

Bring expertise to the interface

Spend time with the people who’ll use the system. How does a clinic schedule appointments? What happens when a warehouse order goes wrong? How does a teacher prepare a lesson? Those details help you decide what the software needs to do.

04

Try a focused product idea

A small team can use AI to draft prototypes and compare designs before committing to a larger build. Look for recurring problems in local businesses, underserved languages, or a profession you know. Find someone who wants the problem solved, then test the prototype with them.

Junior developers still need to debug, model data, read unfamiliar code, and explain tradeoffs. Use an assistant during practice, then investigate and run the code it produces. Can you spot an incomplete answer and work out what’s missing?

Experienced engineers can support that learning by reviewing the reasoning behind a change, setting realistic exercises, and giving people responsibility gradually.

07 / A PRACTICAL START

A 90-day project plan

Pick a manageable project and try this learning plan. The goal is to gather evidence about what works, with enough time to investigate what doesn’t.

  1. DAYS 01–30

    Understand one workflow

    Choose a recurring task. Record how it works today, where errors occur, and who checks the outcome. Build a baseline with a handful of representative examples. Learn the underlying tools well enough to spot an incorrect result.

  2. DAYS 31–60

    Build a small assistant you can review

    Connect only the data and tools it needs. Define the actions it can take and the points that need human review. Add tests, logs, and a straightforward way to undo a change. Compare model options on the same examples.

  3. DAYS 61–90

    Test it and write up the results

    Test with real users and unfamiliar cases. Report failures as clearly as successes. Write up the problem, your decisions, evidence of improvement, and remaining limits for your portfolio. Explain what the results support, even if the improvement is modest.

08 / LOOKING FURTHER AHEAD

Beyond AGI: ASI and the singularity

Artificial superintelligence, or ASI, concerns capabilities beyond human-level general intelligence. The singularity imagines technological change accelerating beyond our ability to predict it, potentially through AI improving AI. DeepMind’s research report explores possible routes and uncertainties.

FIG. 03

Four possible paths beyond AGI

01

Scale

Expand the resources available to capable systems.

02

New methods

Find more effective algorithms or architectures.

03

Self-improvement

Use AI to help improve future AI systems.

04

Agent collectives

Coordinate many systems on larger problems.

Conceptual pathways described by Google DeepMind. They may overlap and still face unresolved bottlenecks. This graphic is not a timeline or a probability estimate.

Altman’s June 2025 essay The Gentle Singularity imagines AI-assisted research helping improve later AI systems. He distinguishes that feedback from a system autonomously rewriting itself. Whether that progress can sustain itself remains something to demonstrate.

More accessible learning, better research tools, and software for overlooked communities are goals worth working toward. Bring educators, designers, security specialists, and domain experts into that work alongside researchers and software engineers. Their knowledge helps decide what to build, who can use it, and how to check that it works.

09 / A FEW GOOD QUESTIONS

Common questions

Do GPT-6 Astra and Claude Fable 5.1 prove that AGI is here?

No single result settles it. These releases show strong capabilities, but researchers still disagree about the AGI threshold. Check which task was tested, under what conditions, and how reliably the system performed.

Is an ARC-AGI score a percentage of human intelligence?

No. ARC-AGI-3 combines task completion and action efficiency against a human baseline. The percentage describes that score. It isn’t an IQ score or a measure of every human ability.

How are Claude Fable 5.1 and Mythos 5.1 different?

Anthropic describes them as the same underlying model with different safeguards and access arrangements. Fable is generally available. Mythos is offered through restricted, trusted-access programs.

What should a student or junior developer learn first?

Start with programming fundamentals, debugging, data literacy, and communication. Use AI on a small project, then test its output and explain how it works. You need enough understanding to catch a plausible but incorrect answer.

When will ASI or the singularity arrive?

There is no verified date. Build adaptable skills and reliable systems as capabilities develop.

Choose one task.
Keep a record of what happens.

Use the evaluation log and calculator on a task you already understand. Record the accepted work, corrections, and time spent. Those results will tell you where to try the model next.

Sources & further reading

Benchmark source check: September 7, 2026. Leadership sources added and reviewed September 8, 2026. Their original publication dates appear below. Forecasts are attributed to their authors. The product framework, proposed trials, and learning plan are this journal’s analysis.

  1. OpenAI: GPT-6 AstraRelease announcement and comparative benchmark table.
  2. Anthropic: Claude Fable 5.1 and Mythos 5.1Capabilities, availability, and evaluation conditions.
  3. ARC Prize: GPT-6 Astra on ARC-AGI-3Independent evaluation and harness differences.
  4. ARC-AGI-3 scoring methodologyHow completion and efficiency contribute to the score.
  5. Google DeepMind: From AGI to ASIResearch overview. read the full paper.
  6. ACS Information Age: The AGI claim and expert disagreementIndependent reporting on the debate.
  7. Sam Altman: ReflectionsJanuary 2025. Primary account of his outlook and deployment approach.
  8. Sam Altman: Three ObservationsFebruary 2025. His economic and social expectations.
  9. Dario Amodei: Machines of Loving GraceOctober 2024. Conditional benefits scenario.
  10. Greg Brockman & Ilya Sutskever: The OpenAI MissionMarch 2019. Historical mission statement.
  11. Sam Altman: The Gentle SingularityJune 2025. His view of progress toward superintelligence.

Reporting context: The Guardian reports Brockman’s AGI-era framing of Astra’s launch. The StartupFortune discussion highlights disagreement over definitions. Both are secondary accounts. Only the public preview of Alex Heath’s Sources article was accessible. Its subscriber-only body is not used as evidence here.

Additional reading: The Verge · Bloomberg. These linked articles were not used to establish the factual claims or chart values above.

Claps, saves, topic follows, and comment previews are for this visit only. Comments are not published.

ABOUT THE AUTHOR

Hashan Shalitha

Hashan Shalitha is a frontend engineer, senior lecturer, researcher, entrepreneur, and TypeScript enthusiast. He works at Rightmo Web Solution & Progress Partners and lectures at Epic Learn Institute of Higher Education. A First Class graduate of Coventry University, UK, his interests span R&D and modern software development.

Background, projects & editorial approach

Comments

GPT-6 Astra & Claude Fable 5.1: Benchmarks and the AGI debate

Preview only. Your comments are not published and disappear when you leave this page.

No comment previews yet

Add a thought above to see it here. Only you can see these previews.

Share this story

GPT-6 Astra & Claude Fable 5.1: Benchmarks and the AGI debate

You can also select and copy the link directly.

Audio options

GPT-6 Astra & Claude Fable 5.1: Benchmarks and the AGI debate

Read aloud is unavailable in this browser. Listen mode needs a browser with speech synthesis support.