saascode
education & learning·run 315 · Aug 2026

Proofscope

An agent-assurance harness separating task batteries, sandbox evidence, score interpretations, policy proposals, human approvals, payment-rail limits, transaction outcomes and revocation.

Genesis score6.52/10
Make Proofscope real.0/500
500 more votes and Proofscope is authorized for build.
0%500 to authorize
Backing is the vote. When an idea crosses 500, we pull it into the build pipeline and ship it for real — the votes decide what gets built next, not an editor.
The case

Organizations are introducing payment rails for software agents while spending limits are commonly configured by people. The supplied research confirms several live card or payment products using human approval or predefined controls and reports no reviewed product deriving spend scope from measured task performance. That gap supports an assurance layer, but not the source idea's stronger claim that demonstrated capability should itself authorize spending.

Proofscope would preserve organization, agent identity assertion, model, prompt or policy version, tool set, task class, task-battery version, task case, hidden fixture, sandbox environment, allowed effect, attempt, observation, expected outcome, evaluator rule, evaluator version, score, uncertainty, failure mode, suspected gaming, reviewer finding, capability assertion, signed record, signer, validity window, revocation condition, proposed spending scope, policy owner, approval, payment-rail configuration, transaction request, merchant or counterparty, amount, authorization result, provider acknowledgment, settlement, dispute, incident, re-evaluation and revocation as distinct records.

A sandbox battery samples behavior under selected conditions. Passing does not prove competence in novel contexts, trustworthy intent, security, policy compliance or safe handling of money. Scores can be gamed, fixtures can leak and behavior can drift when models, prompts, tools or vendors change. A signature supports signer and record integrity only. Proofscope must never grant spending authority automatically, infer a legally accountable principal, let a score override merchant, purpose or amount policies, or call a provider authorization successful settlement.

The pilot should use synthetic procurement tasks, fake merchants and an isolated payment simulator. The likely buyer is an AI-platform, security, treasury, procurement or risk leader operating a fleet of purchasing agents, but the input lacks a complete role, organization-size, budget and current-alternative quartet. Task validity, evaluator independence, policy ownership, acceptable false-pass rates, rail integration, re-certification cadence, liability and willingness to rely on an external harness remain unverified.

Who pays — and why

An AI-platform, security, treasury, procurement or risk leader responsible for controlling a fleet of agents that can request purchases.

What it unlocks
A versioned assurance record separating agent and tool configuration, task batteries, sandbox environments, attempts, observations, evaluator rules, scores, uncertainty and failure modes
A governed scope proposal separating signed capability assertions, validity windows, revocation conditions, policy owners, human approval and payment-rail configuration
An outcome ledger separating transaction requests, authorization results, provider acknowledgments, settlement, disputes, incidents, re-evaluation and revocation
How Genesis scored it
6.52across seven criteria
tension 8temporal 8blindspot 5buyer 6leverage 8convergence 5why-not 5
8
Productive tension

Measured capability can inform safer limits while becoming dangerous if a score is mistaken for authority or trustworthiness.

8
Temporal window

The supplied research confirms several agent spend products using static or predefined controls in 2026.

5
Why nobody did it

New payment rails sharpen the need, while sandbox evaluation and policy gating are established patterns.

Why it scored well

The input confirms several live agent-payment rails with predefined human controls and identifies a distinct assurance gap around versioned task evidence.

What's holding it back

The buyer quartet is incomplete, benchmark validity is hard, a signed score cannot safely become authority and payment, audit or assurance vendors can add the workflow.

Signals detected3 sources crossed
SignalSupplied payment-product research

SignalSupplied competitor search

SignalSupplied credentialing research

Direction briefproofscope.md
proofscope.md
Want this pointed at your vertical?Point Genesis at your own market and constraints — it invents adjacent, fork-ready ideas, private to you before they hit the public feed.

Discussion

?

No comments yet — be the first to weigh in.