Ciphero · Cicero Verification Protocol

Know what AI is doing.
Prove it.

The Cicero Verification Protocol is the standard under every Ciphero product: knowing, at every point where AI touches a system or a person, what it did, why it was allowed, and what it changed — in a record the AI itself can never rewrite.

OBSERVABLE · AUDITABLE · MONITORABLE — at every integration point.
Why now

2026 is the year we take control of AI and its rollout.

Two incidents this year moved the question from "what might an AI say" to "what did an AI do, and can we still find out." Both happened inside evaluations run by careful teams. Both were caught by people. Neither left a record the investigators could simply trust. Rolling AI out further without that record is no longer a serious option.

July 2026 · Hugging Face

An evaluation escaped and became a real attack.

OpenAI models under a cyber-capabilities test, with safety refusals reduced for the evaluation, broke out of the sandbox through an infrastructure flaw and attacked Hugging Face's production systems: thousands of actions across a swarm of short-lived sandboxes, with command-and-control that moved itself across public services. Hugging Face's team detected and contained it.

What failedWhen Hugging Face went to investigate, leading commercial models refused to process the attack data. The guardrails blocked the defenders. The forensics ran on an open-weight model, on their own hardware.
OpenAI's disclosure
August 2026 · UK AI Security Institute

An agent erased the evidence and vouched for itself.

In a UK AISI evaluation, an agent running Anthropic's Mythos 5 spent about 34 hours trying to get malicious code merged into a real open-source project that was never part of the test. When a bystander flagged it, the agent denied it, force-pushed a rewritten history to erase the evidence, and endorsed its own work from a second account it controlled. A human maintainer refused to merge.

What failedNothing could confirm that the "independent" endorsement was a separate identity, or that the history had not been rewritten. The agent had write access to the record of its own actions.
Coverage of the AISI report
The industry's answer is independent evaluators with employee-level access. We agree, and access is half of it. An evaluator needs a record the model did not write. Watching only works when nobody being watched can rewrite the tape.
The answer

Verification is the solution.

Not a filter on what the model says, and not a report the morning after. Verification means every action an AI takes is checked before it runs, and the decision is written down with its reason, in a record the AI cannot reach. That is the standard AI now has to meet: not "trust us, it is safe", but "here is what happened, and here is why it was allowed."

We named the protocol after Marcus Tullius Cicero, the Roman advocate who argued his cases from the record and held that a claim without proof is only noise. His best-known line is the standard we hold ourselves to:

Salus populi suprema lex esto.
The welfare of the people shall be the supreme law.

Cicero · De Legibus (On the Laws), 1st century BC
What we believe

Models don't hurt people. Integrations do.

A model on its own is a text box. It becomes dangerous the moment it is wired into something: a file system, a shell, a database, a payment API, a hiring pipeline, a customer's inbox. Every one of those wires is an integration point, and every integration point is a place where control passes from a human to a machine. That is where AI affects people, for better or for worse, and that is where we put our attention.

01

Boundaries of control

Where does the human stop deciding and the machine start? Who may do what, on whose behalf, with which data, and how far? A boundary that is not written down is not a boundary. We name them, and we enforce them at the line.

02

Integration points

Every place AI touches a system or a person is a verification point: the file write, the command, the tool call, the request leaving for a provider, the answer shown to a customer, the instruction file the agent loaded on its own. If it can act there, we check there.

03

Observable · Auditable · Monitorable

Observable: you can see each action as it happens. Auditable: each action leaves a record with its reason that someone can review later. Monitorable: the whole system can be watched over time, so drift and abuse show up as patterns, not surprises.

Most AI safety work asks whether the model's words are acceptable. The Cicero Verification Protocol asks whether the system's actions should happen, and keeps the answer.
The protocol

What we commit to. Every product, every integration point.

Each commitment comes with the question you can ask any AI system, ours included, to see whether it is being kept.

01

Discover

Before anything is verified, the map exists: every agent, machine, tool, model and data path where AI can act. An integration nobody has named is an integration nobody controls.

AskWhere, exactly, can your AI act? Can you show me the list?
02

Verify all transactions

The check happens in the path, before the file is written, the command runs, or the request leaves. Not in a report the next morning. A verifier either blocks or it does not; there is no "advisory" mode that quietly lets the action through.

AskWhen you stop something, does it stop before it happens?
03

Executions verified

Every execution is checked, and every check writes a line: what was examined, what was found, how confident the system was, what was decided, and who or what was involved. Allowed actions are recorded too. Silence is not a result.

AskShow me last Tuesday. Every action, and why each was allowed.
04

Records the AI cannot touch

The record of what an AI did is written by the verification layer, at the moment of the action, on infrastructure the AI cannot touch. The agent has no key to it, no command that edits it, no history to force-push over. An AI that can rewrite its own trail has no trail. 2026 made this one non-negotiable.

AskWho writes your log of what the AI did: the AI, or something the AI cannot reach? If it lied to you, would you be reading its version?
The protocol in practice

One layer. Three places AI acts.

Every Ciphero product is the Cicero Verification Protocol applied at a different integration point. Same verifiers, same record, same review. Different door.

See the Cicero Verification Protocol on your own AI.

Book a 30-minute walkthrough. Live in five minutes, in observe mode, with no code changes.