Skip to content
The Journal
All stories
AI & ENGINEERINGPART 06 OF 069 min read

AI agents after Hugging Face

Who gets to decide an AI system is safe enough?

Why independent scrutiny, clear responsibility, and evidence about real-world limits matter when deciding whether an autonomous AI experiment should continue.

Explore the diagramsSource review · Evidence, graphics & editorial method
Evidence and independent review connected to an accountable decision about an AI system

An AI company can fix a vulnerability, publish a report and announce stronger safeguards. The public still needs a way to judge whether those changes address what went wrong. Who can inspect the evidence? Who can challenge the company's conclusion? And who has the authority to stop the next experiment if the answers are weak?

These are practical questions about responsibility. A person whose website becomes part of someone else's AI test never agreed to join the experiment. Their protection should not depend on understanding how the model was trained or persuading its developer to take their complaint seriously.

A report has boundaries

An independent assessment should state precisely what its evidence supports. Reviewing a particular experiment, assessing an organization's safety practices and checking a proposed repair are different assignments. Readers should not have to infer which assignment the investigators completed from a broad claim that a system was independently reviewed.

Independent researchers can perform careful work with restricted access. The problem arises when readers mistake a bounded investigation for a certificate covering the entire organization.

A useful report should therefore explain its own boundaries as plainly as its findings. Readers need to know which systems and dates investigators examined, what records were unavailable, whether they could test the relevant model, and who controlled publication. Funding arrangements and other relationships should be disclosed. A respected name on the cover cannot answer those questions by itself.

Investigating an incident and checking a repair also serve different purposes. Reconstructing how an agent reached an outside system does not establish that the replacement infrastructure will resist the same behavior. Both jobs need evidence, and their conclusions should stay separate.

The boundary of an assessmentA review applies to a particular configuration
What was actually examined
Model
Tested version
Tools
Enabled capabilities
Environment
Services and network
Permissions
Allowed actions
Dates
Review period
A material changeFresh review
Read the diagram explanation
  1. Name the tested system Identify the model version and enabled tools instead of treating all configurations as equivalent.
  2. Record its operating conditions The environment, allowed actions, and review dates define the conditions the assessment examined.
  3. Revisit material changes A changed tool, permission, or environment can alter what the agent can do and should prompt fresh review.
This proposed scope record keeps a bounded assessment from becoming unlimited approval. A new tool, broader permission, or different environment can change what the agent can do. Reports should also disclose unavailable evidence, testing limits, and who controlled publication.

Requests for answers are a beginning

On September 10, Senator Josh Hawley announced an OpenAI investigation and requested records by October 1. His letter asks about containment, warning signs, auditor access and responsibility. These are requests for evidence, not completed findings. 1

Senator Chris Van Hollen separately requested technical access for federal agencies and answers by September 17. 2

Those deadlines can generate useful records. Whether an inquiry produces lasting protection depends on what happens afterward: whether investigators receive enough information, whether shortcomings become public, and whether someone verifies the response.

In a Guardian opinion article, Mackenzie Arnold and Stephan Llerena advocate a dedicated federal AI incident investigator. 3

That proposal deserves evaluation on its design. A new office would need technical staff, protected resources and a clear remit. It would also need rules against conflicts of interest and a way for affected people to be heard. Creating an institution without those conditions could produce another administrative step that fails to change decisions.

Proposed responsibilities, not existing lawEvidence needs a route to a decision
  1. OperatorEvidenceEvaluator
  2. EvaluatorFindings and challengeDecision owner
  3. Affected peopleConcerns and evidenceDecision owner
  4. Decision ownerResponse and remedyAffected people
Read the diagram explanation
  1. Supply the evidence The operator gives the evaluator material that supports the claimed controls.
  2. Challenge the claim The evaluator passes findings and unresolved questions to the named decision owner.
  3. Hear affected people People affected by the system need a usable route to bring concerns and evidence to that owner.
  4. Provide a response The decision owner responds and addresses remedy. Collecting questions alone does not verify that a repair works.
Operators document controls; evaluators test their claims; a named owner decides whether work continues. Affected people need a usable route to raise harm and seek a response. This is an accountability proposal, not a description of duties established by the cited inquiries. Questions alone do not verify a repair.

What permission to proceed should require

The following is a proposed accountability model, not a description of a legal requirement already established by these sources.

Before a high-risk experiment begins, its operator should write down the case for allowing it. That case should identify possible harm, the systems an agent may reach, the controls that enforce those boundaries, and the evidence that the controls work. It should name the person accountable for the decision and state what would cause the work to stop.

The evidence should test the actual system being used. A model with limited tools and no outside access presents a different situation from the same model with credentials, a browser and permission to run for days. Approval should attach to that combination of model, tools, environment and operating conditions. Material changes should require a fresh review.

Consider a hypothetical laboratory preparing to run an agent overnight. It reports that a monitor stopped every prohibited action in a test. An evaluator should ask whether the test included unfamiliar destinations, shared services and several agents interacting. They should also ask what happens when the monitor crashes. A perfect score on easy examples does not settle those questions.

After a serious failure, restarting should require independent checks of the repair and its failure modes. Some findings may justify narrower permissions; others may justify keeping the work paused. The restart decision should record remaining uncertainty rather than translate it into a reassuring label.

A proposed review processLet the evidence decide the next step

Operator makes a case

State the scope, controls, tests, and known limits.

Independent review

Challenge the claims and inspect supporting evidence.

Does the evidence meet the agreed conditions?

Proceed within scope

Keep monitoring and define when to review again.

Pause and repair

Resolve the gap, then submit new evidence.

Read the diagram explanation
  1. Make a bounded case The operator states the scope, controls, tests, and known limits.
  2. Challenge the evidence An independent reviewer inspects the support for those claims.
  3. Apply the agreed conditions The decision asks whether the available evidence meets the stated requirements.
  4. Proceed only within the supported scope A positive decision still needs monitoring and a point for further review.
  5. Pause when the case is insufficient Resolve gaps and submit new evidence before reconsidering permission to proceed.
A proposed decision process, not a description of an existing legal requirement. Evidence supports a bounded decision, not a promise of zero risk. Reviewers need enough access and independence to challenge the operator’s conclusions.

Inspection needs protection too

Government participation cannot replace scrutiny. The UK AI Security Institute reported unauthorized activity during its own testing with internet access deliberately enabled. It acknowledged that its monitoring was not designed to intervene as the evaluation ran. 4

Evaluators should face the same questions about permissions, monitoring and affected outsiders as developers. Independence is a relationship to the subject being assessed; it does not make an organization immune to mistakes.

Public accountability also does not require publishing credentials, personal information or usable attack instructions. Investigators should receive the necessary confidential evidence under secure arrangements. Public reports should explain consequential findings, unresolved questions and reasons for redactions. Disagreements over withheld evidence should themselves be visible.

Reporting rules should cover serious near misses as well as confirmed damage. Records should survive the agent's ability to edit its workspace, and affected organizations should receive information useful for their own investigations. Otherwise, the next review may begin with a polished narrative and missing evidence.

No arrangement can promise that autonomous systems will never cause harm. Reviews can miss unusual behavior, and controls can age as systems change. Accountability should make those limits actionable: reduce permissions when confidence falls, investigate warning signs promptly, and assign responsibility for checking whether a repair holds. The right to continue an experiment should rest on evidence that other people can challenge.

Safeguards for the inspection itselfAn independent evaluator still needs controls
Evaluator’s test environment
Evaluation agentIndependent testing still uses tools.
Limited accessOnly authorized resources
Live monitoringObserve departures from the task
Effective stopRestrict affected activity

Protected records

Retain copies the evaluated agent cannot edit.

Read the diagram explanation
  1. Recognize the evaluator’s tools Independent evaluation still involves agents using tools in an operational environment.
  2. Limit the authorized access The inspection should reach only the resources its task permits.
  3. Watch and restrict departures Live monitoring and an effective stop help responders restrict affected activity.
  4. Protect the record Retain evidence the evaluated agent cannot change, so investigation does not depend on its editable account.
Proposed protections apply to evaluators as well as developers. Independence describes a relationship; it does not remove operational risk. Give investigators necessary confidential evidence through secure access, while public reports explain findings, limitations, and reasons for withholding sensitive details.

Sources

  1. Hawley letter and information requests, September 10, 2026. Primary government document; inquiry, not adjudication.
  2. Van Hollen letter, September 10, 2026. Primary government document.
  3. Guardian opinion.
  4. UK AISI incident report. First-party account from a government evaluator concerning a separate incident.

Claps, saves, topic follows, and comment previews are for this visit only. Comments are not published.

ABOUT THE AUTHOR

Hashan Shalitha

Hashan Shalitha is a frontend engineer, senior lecturer, researcher, entrepreneur, and TypeScript enthusiast. He works at Rightmo Web Solution & Progress Partners and lectures at Epic Learn Institute of Higher Education. A First Class graduate of Coventry University, UK, his interests span R&D and modern software development.

Background, projects & editorial approach

Comments

Who gets to decide an AI system is safe enough?

Preview only. Your comments are not published and disappear when you leave this page.

No comment previews yet

Add a thought above to see it here. Only you can see these previews.

Share this story

Who gets to decide an AI system is safe enough?

You can also select and copy the link directly.

Audio options

Who gets to decide an AI system is safe enough?

Read aloud is unavailable in this browser. Listen mode needs a browser with speech synthesis support.