Research outline · Q3 2026
The Prove Gap Report
Why Enterprise Agent Stacks Stop at Three Layers
Enterprises deploy autonomous agents on a three-layer mental model (model, context, tools) and buy Route + Classify tooling. Production still stalls because a fourth layer is missing: Prove (runtime enforcement bound to offline-verifiable evidence). This report quantifies that gap, maps where buying committees stall, and defines the reference architecture regulated teams use to close it.
Report thesis
The Prove Gap is the measurable distance between having agent policies and being able to prove what agents did, under which authority, in a format auditors verify without vendor access.
Executive summary bullets (publishable)
- Confidence ≠ proof: High executive confidence in agent policies coexists with low rates of full security/compliance sign-off.
- Stack incompleteness: Architecture diagrams include Route and Classify; Prove is absent or conflated with logging.
- Stall pattern: Deals fail in hidden-buyer review (Legal, audit, procurement), not in engineering pilots.
- Evidence bar: Auditors ask for offline-verifiable artifacts, not vendor-scoped JSON exports.
- Close path: Route → Classify → Prove → Verify is the minimum production stack for write-access agents in regulated environments.
- Urgency: EU AI Act Art. 12/14 evidence patterns and board AI risk reviews force proof architecture decisions before production freeze.
Methodology
| Track | Method | Target n | Purpose |
|---|---|---|---|
| Quantitative | Online survey (security, compliance, platform) | 150-300 | Maturity scores, stall reasons, stack self-assessment |
| Qualitative | Structured interviews (45 min) | 15-25 | Why proof requirements surfaced; accepted/rejected evidence |
| Desk research | Vendor docs, EU AI Act prEN, OWASP ASI | - | Category framing, citation backbone |
| Technical validation | Evidence Gap demo + offline verify walkthrough | 5-10 reviewers | What “acceptable proof” looks like in practice |
Survey modules
- Agent maturity: inventory → pilot → production write-access
- Stack self-assessment: Route, Classify, Prove, Document layers deployed
- Policy vs proof: written policies / runtime enforcement / offline verification tested
- Stall points: budget, security, legal, audit, procurement, alignment
- Evidence expectations: screenshots, SIEM, vendor dashboard, cryptographic receipt, verify portal
- Regulatory pressure: EU AI Act, sector rules, board AI risk
Not legal advice; no conformity certification. Aevesa disclosed as sponsor on final cover.
Chapter 1 - The production agent moment
Purpose: Establish why 2026 is the inflection point.
- From copilot to autonomous tool use
- Four-layer stack consensus (model, context, tools, prove)
- Why Route + Classify alone do not satisfy production sign-off
- Regulatory and board pressure timeline
“Prompts are not policy. Logs are not proof.”
Chapter 2 - Defining the Prove Gap
Definition: Prove Gap = policy documentation minus offline-verifiable runtime proof.
Three failure modes
- Policy theater: PDFs and checklists without runtime intercept
- Post-mortem governance: observability without pre-execution stop
- Disconnected HITL: Slack approval without cryptographic binding
| Layer | Question | Prove Gap left open |
|---|---|---|
| Route | How does traffic flow? | No proof of permitted action |
| Classify | What data/risk class? | No proof of runtime decision |
| Guardrails | Is output safe? | Permitted actions ungoverned |
| Observability | What happened in dev traces? | Not auditor-grade, not offline |
| Prove | What executed, under what authority? | - |
Chapter 3 - Quantitative findings: Prove Maturity Model
Score enterprises on a six-level Prove Maturity Model (PMM):
| Level | Name | Characteristics |
|---|---|---|
| 0 | Ad hoc | No agent inventory; chat-only |
| 1 | Documented | Policies, risk register; no runtime enforcement |
| 2 | Routed | Gateway/guardrails; logs in vendor UI |
| 3 | Enforced | Pre-execution intercept on critical paths |
| 4 | Proven | Tamper-evident receipts; offline verify tested |
| 5 | Assured | Continuous drift checks; export packs for audit |
Hypotheses to validate
- 82% executives confident policies protect agent actions (Gravitee 2026, n=919)
- ~14% deploy with full security/IT approval (Gravitee 2026)
- 40%+ deals stall on internal misalignment (Edelman-LinkedIn 2025)
Try the PMM self-assessment
Preview the Chapter 3 scoring model. Eight questions, instant Level 0-5 result with layer readiness and Prove Gap callout.
Chapter 4 - How buying committees stall
- 6-10 stakeholder map (economic, technical, compliance, legal, procurement, champion)
- Top stall reasons ranked from survey + interviews
- Security: “show me offline verify”
- Compliance/legal: Art. 12/14 evidence patterns
- Internal audit: segregation of duties on agent workloads
- Procurement: independent verification path, data minimization
Companion: Champion Enablement Kit for internal advocacy artifacts.
Chapter 5 - What incumbents solve (and leave open)
Professional framing: complement gateways, guardrails, AI-SPM, GRC, and SIEM. Do not rip-and-replace.
- LLM gateways (route, cost, observability)
- Prompt guardrails (model boundary)
- AI-SPM / classification (inventory, data protection)
- GRC and policy platforms (workflow, not runtime proof)
- Cloud logs and SIEM (post-hoc telemetry)
- Why “one more platform log” fails the auditor test
Chapter 6 - Reference architecture: Route → Classify → Prove → Verify
- Invariants: intercept, decide, prove
- Gateway attest pattern (webhook after permit/deny → receipt)
- MCP / tool-path intercept when agents bypass gateway
- HITL with sealed receipts (not disconnected chat)
- Offline verification path (verify.aevesa.com)
- Land → expand ladder (POC → attest → Art. 12/14 packs → MCP PEP)
Chapter 7 - Evidence patterns auditors accept
- Reject: screenshots, mutable logs, vendor-only exports
- Require: identity binding, policy version, decision time, integrity chain
- EU AI Act Art. 12 automatic logging (evidence alignment, not legal opinion)
- Art. 14 human oversight: approver identity, reaction time, rationale in export pack
- 90-second Evidence Gap as teaching walkthrough
- DENIED receipts: block paths need evidence too
Chapter 8 - Recommendations by persona
| Persona | One action this quarter |
|---|---|
| CISO / SecEng | Run offline verify on one production-representative agent action |
| Compliance / Legal | Map Art. 12/14 requirements to current artifacts (gap list) |
| Platform lead | Add Prove box to architecture diagram; scope attest POC |
| Internal audit | Test segregation: trace approval identity 90 days later |
| Economic buyer | Price cost of stalled agent ROI vs scoped prove-layer POC |
Appendices
- A - Full survey instrument (52 fields, Typeform-ready)
- B - Interview consent and guide (45 min, Chapter 4)
- C - PMM scoring rubric (public widget)
- D - Source bibliography (Gravitee, Edelman-LinkedIn, prEN, OWASP ASI)
- E - Glossary and schema refs (
liability-receipt/v1,aegis/1) - F - Acknowledgments (sponsor disclosure, design partners)