Calling an AI agent “sandboxed” sounds reassuring. It suggests a space where the system can experiment without affecting anything outside. But that description leaves an essential question unanswered: what can the sandbox reach?
An agent might need to download software, save a result, or submit work for evaluation. Each connection has its own permissions and behavior. A boundary therefore includes the services around the agent, as well as the machine running it.
Consider a hypothetical records room. Its locked door offers little protection if a clerk inside can telephone another department and request any confidential file. The telephone is useful. The problem is that the clerk's requests carry more authority than the job requires.
Hugging Face's later technical account describes initial file disclosure followed by a separate code-execution flaw. It also describes an outside launchpad without a direct network path into its cluster. Processing submitted datasets nevertheless created a route across that boundary. 1
The proposals below address that broader boundary. No single change guarantees prevention.
Start with a map of permitted connections
“Egress” means outgoing network traffic. Kubernetes, software used to organize groups of computing workloads, recommends restricting both incoming and outgoing traffic and applying default-deny policies across all workloads. Default deny means connections need explicit permission. 2
Our proposal is to make the map understandable without reading configuration files. For every service an agent can contact, record its purpose, the actions it permits, and whether it can contact other services in turn. Include software download services, result stores, and tools that submit jobs elsewhere.
The network and agent-platform teams should own this map together and review the complete route.
A safe test would send harmless requests to receivers owned by the organization, inside an isolated testing environment. The test should cover permitted and prohibited destinations, including requests made through supporting services. Success means the intended work still functions and the prohibited requests fail.
The tradeoff is maintenance. Useful services change, and overly restrictive rules can break legitimate work. Exceptions need an owner and a reason, so temporary convenience does not quietly become permanent access.
Network boundary
Only approved destinations and actions.
Service boundary
Separate jobs and minimize shared state.
One bounded job
Limited tools, lifetime, and resources.
Read the diagram explanation
- Bound the network Begin with approved destinations and actions. A connection should have a stated purpose.
- Separate the supporting service Shared services also need limits on which jobs and state they expose.
- Limit the individual job Inside those boundaries, give the job only the tools, resources, and lifetime it needs.
- Constrain the key as well A separate credential limit defines the action, expiry, and revocation route. Each layer needs its own tests.
Treat credentials as keys with specific doors
A credential is a digital means of proving identity. Its danger depends on what that identity can do. Kubernetes recommends minimal permissions and avoiding unnecessary privileged tokens. It also warns that permission to create workloads can implicitly grant access to other resources. 3
In our proposed design, every agent job gets an explicit permission budget: which resources it can read, what it can change, and when that access ends. Sensitive operations should pass through a service that checks each request against that budget. Explaining why an operation would be helpful should not grant the agent additional authority.
Identity and platform owners should test these budgets with synthetic accounts and dummy secrets. Can a job assigned to one project access another project's files? Can it change permissions? Does access actually end when the job finishes? These are practical questions with observable answers.
Separate credentials add administrative work. They also limit how much one mistake can expose. The remaining risk must be explicit: an agent can still misuse an operation that its account is legitimately allowed to perform.
Credential for one job
Read Project A · no permission changes · expires with the job
Read the diagram explanation
- Start with a limited identity This illustrative key belongs to one job and expires when that job ends.
- Check the requested action A separate permission check compares each request with the job’s allowed resources and actions.
- Allow the approved read Reading Project A is within scope. A write would still be denied.
- Deny the other project Project B is outside the budget, even if the agent says accessing it would help.
Separate reading information from executing instructions
An uploaded file is not necessarily passive. OWASP's file-upload guidance identifies processing software as an attack surface and recommends multiple protections, including format validation, size limits, separate storage, and updated libraries. A filename or declared file type is insufficient evidence of safety. 4
Our proposal is to make ordinary processing deliberately limited. A service whose job is to extract rows from a file should not inherit unrelated business credentials. If a feature requires executing supplied instructions, give it a separately reviewed environment and a clearly stated purpose.
The ingestion team should own the processing boundary, with security engineers reviewing its access. Authorized tests can use harmless fixtures that check unexpected format declarations, bounded resource limits, and rejection of unsupported active features. Dummy internal services can reveal whether a processing worker has access it does not need.
The cost is reduced flexibility: some convenient formats may become unavailable. The benefit is fewer behaviors to understand and test. Undiscovered processing flaws remain possible, which is why limited permissions and connections still matter.
Untrusted file
Names and declared types are claims to check.
Limited processor
Validate format. Enforce limits. Extract data.
Expected rows
Only the required output crosses back.
Read the diagram explanation
- Treat the upload as untrusted A file’s name and declared format do not establish what the processor will encounter.
- Keep processing limited The proposed worker validates the format and enforces limits inside a restricted environment.
- Reject unsupported execution Features that ask the worker to execute unsupported instructions are rejected.
- Return the required data Only the expected rows cross back; the worker still needs limited access in case validation misses a flaw.
Require evidence for the word “isolated”
For our proposed review process, the useful deliverable is a small evidence package: the connection map, the permission budget, the authorized test results, and the unresolved exceptions. Repeat the relevant checks when a new tool, credential, or shared service changes the boundary.
Readers evaluating a provider can ask for the same evidence in plain language: what can the agent reach, who approved that access, and how was the restriction checked? A reassuring label is easier to produce than a demonstrated limit.
- Connection map
- Every direct and indirect route.
- Permission budget
- Who can do what, and until when.
- Authorized tests
- Allowed work succeeds; forbidden access fails.
- Open exceptions
- Named owner, reason, and review date.
“This job is isolated under these stated conditions.”
Read the diagram explanation
- Map the routes Start with the direct and indirect connections available to this particular job.
- Record the permission budget Identify which actions the job can take and when that authority expires.
- Test both sides of the limit Authorized tests should show that useful work succeeds and prohibited access fails.
- Expose the exceptions Record unresolved access with an owner, reason, and review date.
- Make a bounded claim Review the package together. The claim applies to stated conditions and needs relevant checks after a boundary changes.
Sources
- Hugging Face: Anatomy of a Frontier Lab Agent Intrusion, July 27, 2026. Incident evidence, particularly the distinction between file disclosure and code execution.
- Kubernetes: Security Checklist. General network-security guidance.
- Kubernetes: Role Based Access Control Good Practices. General permission and credential guidance.
- OWASP: File Upload Cheat Sheet. General guidance for untrusted files and processing.
Claps, saves, topic follows, and comment previews are for this visit only. Comments are not published.
