Agentic zero-knowledge proofs
○ Seed Planted Tended #agents #cryptography
An investor wants to know a startup’s product is built the way the pitch says: the model is really theirs and not a wrapped API, there’s no GPL code in it, and the tests actually run. The startup won’t hand over its codebase to someone who hasn’t signed yet, and might be backing a competitor next quarter. So the investor sends an agent in instead. It reads what it needs and reports back a verdict. The investor never sees the code, and the startup can’t tamper with what the agent was told to check.
What has to be true
- The prompt is the verifier’s, untouched. The verifier commits to the agent’s prompt, and the attestation shows that exact prompt ran. The prover can’t soften it into “please say yes”.
- The execution is honest. Something has to vouch that this model + this prompt + this context produced this output. zk proofs of LLM inference are nowhere near practical, so today this probably means Trusted execution environments with remote attestation. Zero-knowledge in spirit, not (yet) in maths.
- Only the verdict leaves. Output is schema-constrained (pass/fail plus a few bounded fields) so the trace can’t leak through a chatty justification. Every free-text byte is a side channel.
- One shot. A verifier-supplied nonce is bound into the attestation, so the prover can’t rerun until they get a favourable answer and only show that one.
- The context is real. The hard part. Attestation proves the agent read something faithfully, not that the git history wasn’t forged last night. Inputs need their own provenance: signed commits, platform attestations, or zkTLS to show the data really came from GitHub.
Why it’s interesting
It flips verification from disclose everything and trust the reader to disclose nothing and trust the process. The agent becomes a programmable form of Selective disclosure: the question can be anything you can write a prompt for. Generalises to audits, compliance and due diligence: anywhere one party holds private context and another needs a narrow answer about it.
Open questions
- Can the verifier’s prompt stay private too, so the prover can’t game it?
- Who runs the enclave, and does that just move the trust somewhere else?
- What does an appeal look like when the agent gets it wrong?
Further reading
- NDAI Agreements: the closest prior work. Inventor and investor agents bargain inside a TEE, and anything disclosed is wiped if no deal happens. It’s framed as a fix for Arrow’s disclosure paradox, which is also the investor example above.
- Security awareness in LLM agents: the NDAI zone case: a follow-up across 10 models. A failing attestation reliably makes agents clam up, but a passing one gets mixed reactions, so today’s agents can spot danger but can’t confirm they’re safe.
- Beyond Privacy Trade-offs with Structured Transparency: a vocabulary for this whole space. It separates input privacy, output privacy, input verification and output verification, which map closely onto the list above.
- Ensuring mutual privacy is necessary for effective external evaluation of proprietary AI systems: the same two-sided secrecy problem, for AI evals. The lab won’t share its model and the evaluator won’t share its held-out tests, which is the open question about keeping the verifier’s prompt private.
- Verifiable evaluations of machine learning models using zkSNARKs: proves a model with private weights hits a stated score on a benchmark. This is real zero knowledge, but for a fixed eval rather than an open-ended prompt.
- zkLLM: zk proofs of inference for a 13B-parameter LLM, at roughly 15 minutes per proof. It’s a useful benchmark for how far the maths is from replacing the enclave.