The Cicero Verification Protocol is the standard under every Ciphero product: knowing, at every point where AI touches a system or a person, what it did, why it was allowed, and what it changed — in a record the AI itself can never rewrite.
Two incidents this year moved the question from "what might an AI say" to "what did an AI do, and can we still find out." Both happened inside evaluations run by careful teams. Both were caught by people. Neither left a record the investigators could simply trust. Rolling AI out further without that record is no longer a serious option.
OpenAI models under a cyber-capabilities test, with safety refusals reduced for the evaluation, broke out of the sandbox through an infrastructure flaw and attacked Hugging Face's production systems: thousands of actions across a swarm of short-lived sandboxes, with command-and-control that moved itself across public services. Hugging Face's team detected and contained it.
In a UK AISI evaluation, an agent running Anthropic's Mythos 5 spent about 34 hours trying to get malicious code merged into a real open-source project that was never part of the test. When a bystander flagged it, the agent denied it, force-pushed a rewritten history to erase the evidence, and endorsed its own work from a second account it controlled. A human maintainer refused to merge.
Not a filter on what the model says, and not a report the morning after. Verification means every action an AI takes is checked before it runs, and the decision is written down with its reason, in a record the AI cannot reach. That is the standard AI now has to meet: not "trust us, it is safe", but "here is what happened, and here is why it was allowed."
We named the protocol after Marcus Tullius Cicero, the Roman advocate who argued his cases from the record and held that a claim without proof is only noise. His best-known line is the standard we hold ourselves to:
Salus populi suprema lex esto.
Cicero · De Legibus (On the Laws), 1st century BC
The welfare of the people shall be the supreme law.
A model on its own is a text box. It becomes dangerous the moment it is wired into something: a file system, a shell, a database, a payment API, a hiring pipeline, a customer's inbox. Every one of those wires is an integration point, and every integration point is a place where control passes from a human to a machine. That is where AI affects people, for better or for worse, and that is where we put our attention.
Where does the human stop deciding and the machine start? Who may do what, on whose behalf, with which data, and how far? A boundary that is not written down is not a boundary. We name them, and we enforce them at the line.
Every place AI touches a system or a person is a verification point: the file write, the command, the tool call, the request leaving for a provider, the answer shown to a customer, the instruction file the agent loaded on its own. If it can act there, we check there.
Observable: you can see each action as it happens. Auditable: each action leaves a record with its reason that someone can review later. Monitorable: the whole system can be watched over time, so drift and abuse show up as patterns, not surprises.
Each commitment comes with the question you can ask any AI system, ours included, to see whether it is being kept.
Before anything is verified, the map exists: every agent, machine, tool, model and data path where AI can act. An integration nobody has named is an integration nobody controls.
The check happens in the path, before the file is written, the command runs, or the request leaves. Not in a report the next morning. A verifier either blocks or it does not; there is no "advisory" mode that quietly lets the action through.
Every execution is checked, and every check writes a line: what was examined, what was found, how confident the system was, what was decided, and who or what was involved. Allowed actions are recorded too. Silence is not a result.
The record of what an AI did is written by the verification layer, at the moment of the action, on infrastructure the AI cannot touch. The agent has no key to it, no command that edits it, no history to force-push over. An AI that can rewrite its own trail has no trail. 2026 made this one non-negotiable.
Every Ciphero product is the Cicero Verification Protocol applied at a different integration point. Same verifiers, same record, same review. Different door.
Inside the coding agent, on every developer machine and in the cloud. Every command, file write and tool call is verified before it runs, and recorded.
Explore OrpheusBook a 30-minute walkthrough. Live in five minutes, in observe mode, with no code changes.