NSF SBIR Phase I · NSF 26-510 Project Pitch Draft v1 · 24 July 2026 Roadmap · Trade · Peacebuilding

Capability-based security for autonomous AI agent tool access

A least-authority trust substrate for AI agents that execute real actions — unforgeable, attenuable, fail-closed capability tokens gating every tool call, with a mechanically-proven never-brick safety property.

Draft for your review. This is a working draft grounded in the behavior of a deployed sovereign system. It is a federal submission — verify every factual claim before you file, and fill the bracketed You supply fields (company, PI, team, market evidence). NSF allows two pitches per company per 12 months, so spend the first deliberately.
Proven demonstrated by the deployed system Phase I risk the genuine research unknown You supply operator input required
01

The Technology Innovation

Autonomous AI agents increasingly take real actions — calling tools, hitting APIs, moving money, changing infrastructure — through interfaces like the Model Context Protocol and function-calling. The security model for those actions has not kept pace. In practice an agent is handed broad, long-lived credentials, or a human is asked to approve every individual call. The first is unsafe; the second does not scale to fleets of agents doing continuous work.

Our innovation is an object-capability (ocap) security substrate for agent tool access. Every tool invocation is gated by an unforgeable, cryptographically-signed capability token that names exactly which tools it grants, carries an expiry, and can be attenuated — re-delegated with strictly reduced authority — without a central broker in the loop. Verification is fail-closed: absent or invalid authority is a refusal, never a default-allow. Proven

Layered on the capability core is a mechanical “never-brick” safety property: operations that could cause irreversible harm must be read-only, reversible, or fail-safe by construction, and a gate refuses anything that cannot prove it — the guarantee is enforced mechanically, not asserted in a policy document. The whole runtime is deterministic by construction (integer-exact, no floating-point nondeterminism), and every safety-critical component ships with an adversarial gate carrying explicit negative controls. Proven

Why this is novel and not incremental: decades of capability-security theory never reached mainstream OS or application security, which settled on ambient-authority models (broad credentials, coarse role-based access). Autonomous agent fleets are the first surface where least-authority-by-construction and attenuate-only delegation are not just elegant but necessary — and we have a working implementation, not a paper. This is a different trust model matched to a new problem, not a hardening of OAuth or RBAC.

02

Technical Objectives & Challenges

≤ 500 words

Phase I resolves the genuine unknowns that stand between a working system and a rigorously-guaranteed one. Each objective has a measurable milestone.

  • O1Prove the attenuation algebra is monotone. Formally characterize delegation and machine-check that no composition of delegations can escalate authority under adversarial re-delegation. Challenge: closing every escalation path, not just the obvious ones. Phase I risk
    Milestone: a mechanized proof (or a counterexample that redirects the design).
  • O2Adversarial evaluation harness. A red-team suite that attempts token forgery, replay, confused-deputy, and cross-agent escalation across a simulated fleet, with negative controls (a test that passes when it should fail is the defect we hunt). Phase I risk
    Milestone: measured forgery-refusal rate; zero escalation paths found; every check paired with a negative control.
  • O3Generalize the never-brick gate. Extend the mechanical irreversible-action guarantee from its current scope to a broader property checker with proven soundness — no false “safe.” Phase I risk
    Milestone: a soundness argument plus a mutation harness that breaks the subject and confirms the gate catches it.
  • O4Interop without adoption. A shim that brings the security property to third-party agent frameworks (MCP servers, function-calling APIs) without requiring them to adopt our runtime — proven to preserve least-authority end to end. Phase I risk
    Milestone: interop demonstrated with N external frameworks within a stated overhead budget.
03

The Market Opportunity

Enterprises are deploying autonomous agents into production — operations, software, customer-facing systems — faster than the security model for agent actions has matured. Every organization fielding a fleet hits the same wall: how do you grant an agent exactly the authority a task needs, revocably and auditably, at fleet scale, without a human in every loop?

Today’s answers — broad API keys, per-call human approval, coarse role-based access — each fail on one of safety, scale, or auditability. The buyer is any organization operating agent fleets; the initial wedge is security for agent platforms and the fast-growing MCP tool ecosystem, where tool invocation is already the unit of action and a capability model drops in naturally.

You supply — NSF weights commercial evidence heavily Insert your customer-discovery evidence (who you spoke to, what they said), a defensible TAM/SAM framing with sources, and any letters of interest or design partners. NSF is funding research with commercial potential; this section should be concrete and specific, not aspirational.
04

The Company & Team

“Why you, why now.” The strongest asset here is that a working sovereign implementation already exists — the “can they build it?” risk is largely retired; Phase I is about proving and generalizing it.

Company
[Legal entity — the R&D C-corp from the funding roadmap; must be US-based, ≤500 employees incl. affiliates, majority US-owned] You supply
PI
[Name — must be ≥51% employed by the company at time of award, with legal right to work in the US: citizenship, permanent residency, or appropriate visa] You supply
Team
[Technical team and the specific expertise that makes this credible] You supply
Why now
Agent fleets are moving into production this cycle; the action-security gap is open and urgent, and a deployed implementation de-risks execution.
One structural reminder from the roadmap

This company is Track A — the R&D entity. Keep the import/export business and any Belarus relationships in the separate Track B/C entities. NSF requires the PI’s primary employment to be here, runs foreign-risk due diligence, and asks senior personnel to document foreign affiliations; the firewall keeps that disclosure clean and the 51% math workable.