Appendix A · Research instrument v0.1

Prove Gap Report Survey

52 fields · 12-14 minutes · Target n 150-300 · Typeform / Qualtrics build

Full quantitative instrument for The Prove Gap Report (Q3 2026). Section 5 (S5A-S5H) matches the public PMM self-assessment widget. Field IDs are export column headers for analysis.

Section 0 - Screening

Terminate if S0A = none, or unqualified role/size combo.

IDPromptResponse
S0AOrganization’s use of AI agentsChat only · Planning · Pilots w/ tools · Production w/ tools · None (terminate)
S0BYour roleSecurity · Compliance/Legal · Platform · Audit · Procurement · Executive · Other
S0COrganization size<500 · 500-4,999 · 5,000-19,999 · 20,000+
S0DRegulated industry?Yes · No

Section 1 - Respondent profile

IDPromptResponse
S1APrimary industryFinancial · Healthcare · Tech · Manufacturing · Retail · Public · Services · Other
S1BHQ regionNA · EU/UK · APAC · MEA · LATAM
S1CDeployment contexts (multi)Customer-facing · Employee · IT/ops · Dev/coding · Back-office · R&D
S1DInfluence on production decisionsFinal approver · Strong influencer · Contributor · Observer
S1EAgent workload count (optional)Integer or Unknown

Section 2 - Agent maturity

Feeds Chapter 1; validates 14% sign-off hypothesis via S2C.

IDPromptResponse
S2AFurthest stage any workload reachedIdeation · Policy only · Non-prod pilot · Limited prod · Broad prod w/ write
S2BFormal agent inventory?No · Informal · Yes, quarterly+
S2CFull security/compliance sign-off (Likert 1-5)Validates ~14% approval stat
S2DBlocked or rolled back after pilot?No · One · Multiple · N/A

Section 3 - Stack self-assessment

Multi-select per layer. Derived: STACK_LAYER_COUNT.

IDLayerOptions (multi-select)
S3ADocumentPolicy · Risk register · Architecture diagrams · Vendor DD · None
S3BRouteLLM gateway · API mgmt · MCP routing · Direct API only · None
S3CClassifyAI-SPM · Guardrails · DLP · Classification tags · None
S3DEnforceModel guardrails · Pre-exec intercept · Policy engine · HITL gate · None
S3EProveVendor logs · SIEM · Decision records · Tamper-evident receipts · Prove product · None
S3FVerifySecurity offline test · Audit offline test · Third-party portal · Runbook · None

Section 4 - Policy vs proof

Likert 1-5. Derived: CONFIDENCE_GAP = S4A minus mean(S4B,S4C,S4D). Validates 82% confidence hypothesis via S4A.

IDStatement
S4AExecutive leadership is confident our agent policies protect against harmful actions.
S4BWe can prove what each production agent did, under which policy version, at decision time.
S4CRuntime enforcement exists on high-risk tool paths before side effects.
S4DAn external reviewer could verify our agent evidence without vendor logins.
S4EOur observability stack satisfies internal audit for agent workloads.

Section 5 - PMM core

Same questions as the PMM widget. Calculates PMM_LEVEL 0-5 and PROVE_GAP flag.

IDWidgetPrompt
S5Aq1Autonomous agents at your organization (0-3)
S5Bq2Governance documentation maturity (0-2)
S5Cq3Route layer: traffic direction (0-2)
S5Dq4Classify layer: inventory and protection (0-2)
S5Eq5Runtime enforcement on high-risk paths (0-3)
S5Fq6Proof artifacts for decisions (0-3)
S5Gq7Offline verification without vendor login (0-2)
S5Hq8Ongoing assurance / drift monitoring (0-2)

PMM gates: Level 4 requires S5F≥2 and S5G≥1; Level 5 requires S5F≥3, S5G≥2, S5H≥1. See outline Chapter 3 and Appendix C rubric.

Section 6 - Production stall points

Feeds Chapter 4. S6B ranks top 3 from S6A selections.

IDPrompt
S6AFactors that stalled sign-off (multi): Security · Legal · Compliance · Audit · Procurement · Budget · Integration · Alignment · No stall · Other
S6BRank top 3 stall factors
S6CLongest single review cycle: <4wk · 1-3mo · 3-6mo · 6mo+ · N/A
S6DHidden buyers drive more delays than engineering (Likert 1-5)

Section 7 - Evidence expectations

IDPrompt
S7AEvidence requested for sign-off (multi): Policy PDFs · Architecture · Pentest · SIEM · SOC2 · Screenshots · Crypto receipts · Offline demo · Third-party verifier · None
S7BEvidence rejected as insufficient
S7CCould audit re-verify a decision 90 days later?

Section 8 - Regulatory and board pressure

IDPrompt
S8APressures apply (multi): EU AI Act · US sector · UK/other · Board AI risk · Insurance · Customer clauses · None
S8BArt. 12-style logging influences architecture (Likert 1-5)
S8CArt. 14-style oversight influences approval flows (Likert 1-5)
S8DLegal asked for proof of what agent did (not just policy)

Section 9 - Tooling map

IDPrompt
S9ATool categories in stack (multi): Gateway · Guardrails · AI-SPM · GRC · Observability · SIEM · IAM · Prove layer · DIY · None
S9BWhat is missing from architecture diagram? (open, optional)
S9CPrimary governance integration: Gateway webhook · Sidecar/PEP · Post-hoc export · Manual only · Not defined

Section 10 - Incidents

IDPrompt
S10AAgent security/compliance incident in last 12 months
S10BIncident type if applicable (multi, optional)

Section 11 - Human oversight (HITL)

IDPrompt
S11AWhere human approval lives: None · Email/chat · Ticketing · Slack w/ partial log · Crypto-bound receipt
S11BCan identify approver, timestamp, policy version months later (Likert)
S11CDENY decisions recorded with same rigor as approvals

Section 12 - Architecture completeness

IDPrompt
S12AArchitecture diagram includes distinct Prove/Evidence layer?
S12BRoute + Classify alone will satisfy production sign-off (Likert)
S12CTime to add offline-verifiable proof to one prod path

Section Z - Opt-in

IDPrompt
SZ1Contact with aggregated benchmark results?
SZ2Work email (if opt-in)
SZ345-minute research interview? (optional)
SZ4Anonymous quote attribution? (optional)