EU AI Act logging: what agents must record
Article 12 has applied to high-risk systems since 2 August 2026. It requires automatic, lifetime event logs, which rules out a self-reported trace.
If you operate a high-risk AI system in the EU, Article 12 requires it to automatically record events across its whole lifetime, in a form that lets someone reconstruct what happened afterwards. The obligations for high-risk systems have applied since 2 August 2026. The practical consequence for agent builders is narrow and sharp: a log the agent produces about its own behaviour is unlikely to satisfy a rule whose entire purpose is independent reconstruction.
We are not lawyers, this is not legal advice, and whether your system is high-risk at all is the question a lawyer should answer first. What follows is what the text says and what it implies for how you build a trace.
What does Article 12 actually require?
The operative sentence is short. Article 12 states that "high-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system."
Three words in that sentence do the work. Technically allow makes it a property of the system, not of your process: a policy saying staff will keep records does not satisfy it. Automatic excludes anything a person compiles afterwards. Over the lifetime rules out logging that begins when an investigation does.
The article then says what the logging has to make possible: identifying situations where the system may present a risk or undergo substantial modification, supporting post-market monitoring, and letting deployers monitor operation. Those are reconstruction jobs. They describe a record somebody reads later to work out what occurred.
Why does self-reported telemetry sit badly with this?
Because the log and the thing being logged have the same author.
When an agent writes its own trace, the record of what it did is produced by the component whose behaviour is in question. It is usually accurate. It is not independent, and independence is the property the whole obligation exists to create. If an agent misunderstood its task, its self-report will faithfully record the misunderstanding as though it were the intent.
There is a second, more mundane problem: an agent that crashes, hangs or is killed writes nothing about the moment that matters most. A record kept by the path the call travelled through survives the failure of the thing making the call.
This is why a gateway position matters for compliance and not only for control. The component that mints the token and forwards the call is the one that can record the call regardless of what happens to the agent afterwards.
What does "appropriate" logging look like in practice?
The regulation does not hand you a schema, and anyone selling you one as mandated is overreaching. What the surrounding requirements imply is reasonably consistent, though:
- Append-only. A record that can be edited after the fact answers no question anybody serious will ask. Chaining each entry to the previous one by hash makes tampering detectable rather than merely forbidden.
- Refusals as well as actions. A blocked call is exactly the kind of event that indicates a system presenting a risk. Logging only what succeeded discards the interesting half.
- Retention you can actually meet. The Act sets a floor for keeping automatically generated logs, at least six months where the provider controls them, and other law may require longer.
- Redaction at capture, not afterwards. Filtering personal data out of a trace store later means it was in there. Article 12 sits beside GDPR rather than replacing it, so a log that solves traceability by keeping everything has traded one problem for another.
- Identity on the record. Which agent, acting for whom, under what authority. "The system did X" is not reconstructable; "this key, minted for this contract, called this tool" is.
Does human oversight change the design too?
Yes, and it is the requirement people notice second.
Article 14 requires high-risk systems to be designed so that people can effectively oversee them, with measures proportionate to the risk and the level of autonomy. Effective is doing a lot of work in that sentence. Oversight that exists as a dashboard nobody is accountable for reading is not obviously effective; oversight that can actually stop the thing is.
In practice that pushes toward two mechanisms. A checkpoint before anything irreversible, so a person approves rather than reviews after the fact. And a halt that is one action rather than a deployment, because an oversight capability with a release cycle in front of it is not available at the moment you need it.
What should you do if you are not sure you are in scope?
Ask a lawyer, and instrument anyway.
The honest position is that scope is genuinely uncertain for a lot of agent deployments, and confidently telling you otherwise would be the least trustworthy thing on this page. But the cost asymmetry is stark. Building a tamper-evident, gateway-written trace is a design decision that is cheap at the start and expensive to retrofit, because retrofitting means you have no record covering the period before you started.
It is also worth doing when no regulator is involved at all. Every reason Article 12 gives for wanting reconstructable logs is a reason you already have: working out why an agent did something strange, showing a customer what happened, and answering the question of who authorised a change six months after whoever knew has left.
In Cognibl the trace is written by the gateway rather than by the agent, entries are append-only and hash-chained so the chain re-verifies end to end, refusals are recorded alongside calls, and redaction happens at capture according to the policy map rather than as a later cleanup.
Keep reading
- Where to put a human checkpoint on an agentFull autonomy is the wrong goal. Let an agent run through reversible steps and stop it before anything that cannot be undone: sending, deleting, charging.
- What replaces velocity when agents write codeStory points measured effort, and an agent can emit a hundred in an hour. What survives is verified throughput, first-pass rate and where the time went.
- What is AGENTS.md, and what belongs in it?AGENTS.md is a README for agents: the build commands, conventions and boundaries an agent needs. Over 60,000 repositories ship one. What to put in yours.