Alex Vakhitov

software engineering4 min read

Why I treat AI agents as identities first

By Alex Vakhitov

A model that gets something wrong gives you a bad answer. An agent that gets something wrong may already have acted on it. That difference is why I think about governing agents separately from governing models.

Agents act

When people talk about AI governance they usually mean the model: how often it's wrong, what it may say, how you test it. That work matters, and I still run the evals I've written about. An agent adds a different problem. It decides while it runs which tools to call, so a single request can turn into a chain of reads, writes and messages that crosses several systems. Nobody wrote that sequence down in advance.

Most agents also work for someone. A person asks, and the agent goes off and does things in their name. So every step raises a question worth asking out loud: whose authority is this?

Access control was designed around two kinds of actor. People sign in and do things. Services run fixed code with fixed permissions. An agent fits neither. It runs like software but chooses its next step more like a person, and it takes direction from whatever it reads, including text written by people outside the organisation. A permission granted when someone signed in might be used hours later, in a situation nobody looked at. And if several agents share a person's login or a single long-lived key, the logs can't tell you which of them did what, or for whom.

The question I start from

Those are identity and authority questions first, and model questions second. The rest of security already has open standards for identity and authority, which is a relief. So the test I hold any agent to is short. For every action, I want to know who acted, for whom, under which rule and with what result. And I want to be able to stop it.

Six layers

I've written this up as a framework with six layers, each with controls you can test.

  1. Identity and authentication. Every agent gets its own identity, registered with an owner and a purpose, and proves who it is with credentials that expire quickly. It never borrows a person's login or a shared key.
  2. Authorisation and delegated authority. An agent acting for someone may do only what both of them are allowed to do, and that is checked again each time it calls a tool. Anything irreversible waits for a named person to approve it.
  3. Control plane and lifecycle. Every agent reaches models, tools and data through one governed route, so there is one registry, one set of policies and one off switch. Whoever builds an agent doesn't also sign off its production permissions.
  4. Accountability. Each agent, and each kind of decision it makes, has a named individual who answers for it. A team alias doesn't count.
  5. Auditability and evidence. One trace ties together the original request, every policy decision, every tool call and any approval, in logs that reveal any tampering.
  6. Runtime security. Assume the agent will read hostile input. Prompt injection can't be solved inside the model alone, so the defence is containment: sandboxed tools, limited outbound access, and secrets kept out of prompts entirely.

Why I built it this way

Three reasons, all from the work itself.

First, I wanted something engineers can build and risk teams can check. Each control has someone who owns it, a way to test it and a record it produces. A control I can't test is one I don't expect to hold.

Second, boundaries. I've written about bounded contexts as the way I structure code for AI coding agents. In production, those same boundaries become identity, permissions and audit. The framework is that idea applied to agents. It is also vendor-neutral, so it holds whichever model or provider you pick.

Third, habit. At Cloud KB I was CTO for seven years, changing software that energy-market participants relied on. That kind of change is slow and checked: know the boundary, prove the new path, keep a way back. With agents, the way back is an off switch you have tested before you need it.

What it maps to

The framework maps to ISO/IEC 42001, to the NIST AI RMF and, where it applies, to the EU AI Act. The Act has no rules written specifically for agents, and whether an agent counts as a high-risk system depends on what it's used for. Where it does, the trace and oversight controls support the Act's duties on automatic event logging (Article 12) and human oversight, including the ability to stop the system (Article 14). Mapping a control to a standard is a long way from certification, and the legal reading belongs to your lawyers.

Where I'd begin

You don't need all six layers on day one. I'd start small: register every agent with an owner, give each its own identity, keep permissions tight, require approval for anything irreversible, add trace IDs, and test the off switch. The rest can follow as you run more agents, or riskier ones.

The full framework, with the controls for each layer and a staged rollout, is on the Comonad site. How I apply it in practice is described on the AI governance page.