Welcome to part three and four of our multi-part series on AI Invariants. In Part 1 we described the stack and how AI works. In Part 2 we discussed the service models of AI, how we use it, and where the threats are. In this part we’ll discuss the available enforcement points and then start crafting the invariants themselves.
Comments are available on the Google Doc version.
Every generative AI interaction is, at its simplest, sending a prompt to a model and getting a result. That said, there are a range of possibilities for how these interactions route, and for enhancing the process by allowing the model to extend its capabilities and interact with other services. This diagram is a simplified model for tracking these communications, and it helps us describe where invariants can be enforced.

If you are familiar with our work on cloud security invariants you’ll notice one important difference in how we describe AI invariants. In the cloud, nearly all of our invariants have a wider, organization-level, scope. But with AI our scope could be as low as an individual agent on an employee’s laptop. Why? Because the scope in the cloud is typically at the provider level: AWS / GCP Organizations, or Azure Enrollment / EntraID . With AI we may want a security property that always holds true at the level of an individual agent or application.
At the core of it, we have the agent or the user. Both have an objective (goal) and will iterate until they achieve the goal. For humans we have the enforcement points of conscience, consequence and common sense (sometimes). Anything resembling that for the agent happens in the classifiers just before they reach the model.
The first enforcement point is at the harness. The problem is there are dozens, if not hundreds, of harnesses in widespread use. Claude Code CLI, Codex, OpenClaw, and Cursor are just a few examples.
The harness may run in a sandbox. This could be a local container runtime, or something like AgentCore. The runtime sandbox is designed to limit the access an agent has to credentials, tools, and data. It should be noted that the harness itself has some level of “sandbox” capability, but it’s paired with the model and can be somewhat non-deterministic. For our purposes we’ll only consider the runtime-sandbox as an enforcement point.
The harness and its runtime-sandbox run somewhere on a host. This gives us another enforcement point with MDM and EDR (if that host is an endpoint). We can ensure that the harness is configured in a specific way, that only specific tools are available, and that the path from the local harness to the remote harness passes though some sort of proxy, firewall or CASB. Server-based harnesses, like you would build in your cloud provider, won’t use MDM or EDR but provide similar levels of control since you own the entire operating system and platform configurations. This is where your cloud invariants can support your AI invariants.
From there, the agent or user’s context will travel over a network (including localhost) to where the inference will happen. There, an additional part of the harness, what we’re calling the inference-harness, will take the entire context, potentially augment it with additional instructions, process it through classifiers, tokenize it, and then pass it to the model. The model produces its output, output classifiers are run, and the result is returned to the local harness.
The process could decide that the use of tools is needed to achieve the goal. The local or inference harness could make a call out to an MCP or CLI to perform an action or gather more data. What tools can be called can be enforced by the runtime-sandbox, the host, or the local or remote harness.
How an agent is identified by the various systems a tool may call is the final enforcement point. The distinct identity an agent has, the method of authentication, and the actions it’s authorized to perform are the final three pieces to the puzzle.
Invariants can be enforced at different points depending on the security property, and the scope. In our next section we’ll use the numbering in the diagram to identify possible enforcement points for each invariant we describe. But, keep in mind, these aren’t absolute since AI is a rapidly evolving technology and different vendors and platforms have wildly different architectures and capabilities.
Now that we’ve covered how AI is built, used, and the broad AI- specific threat landscape, we can finally start digging into the actual list of Invariants an organization might implement. This will vary, depending on the regulatory environment, risk profile, and technical capabilities. This is not intended to be a one-size-fits-all checklist to implement, or a complete list of all possible invariants, but rather a series of example invariants that you could implement to meet your specific needs.
As a reminder, these are just common examples and you will likely only use some of these, and add your own. For example, one of the authors recently implemented an AI platform with an invariant “for any given inference call or AI agent, only the current customer’s data is accessible during a given session”. This was enforced using a combination of IAM, code, and data structures, plus specific adversarial testing in the CI/CD pipeline to ensure future code changes don’t reduce the invariant (blocking test in CI/CD).
Invariants must always be true, so where an exception is needed, it is built into the wording. “No one can create a VPC” is unworkable because that action needs to occur. The correct wording is “only the networking administration team can create a VPC.” An invariant gets a named exception owner only when the absolute form is impractical to live with; the bracketed placeholders ([security], [legal], [privacy], [platform team]) are each organization’s exception owner. Invariants without one are absolute by design.
While writing this paper, we brainstormed a large number of possible invariants we’d want to implement in the organization s we support. However in many cases the “will always hold true” requirement could not be met in a way that is defensible to auditors, regulators, management, or the board. What we were writing were guardrails, not invariants.
An AI guardrail is a point solution intended to stop a misconfiguration or misbehavior. They have to be applied wherever there is a danger zone. For guardrails that protect against misbehavior, they often have to be baked directly into the agentic AI application. For misconfigurations, a guardrail that relies on a human isn’t an invariant. The best guardrails to use for invariants should be:
An AWS Service Control Policy to prevent the creation of unauthorized IAM Access Keys fits these three criteria. It’s deployed at the AWS Organization, so it’s broadly applied. It’s centrally managed by the team that manages the AWS Organization. And it’s enforced via AWS IAM which has a long track record and artifacts to back it up.
“The [platform team] must always be able to halt any high-risk company AI system immediately, including mid-task” enforces an important regulatory requirement under the EU’s AI Act. It needs to be implemented as a guardrail in the scope of the high-risk system.
Our next set of posts will look at what enforcement points are available in the major cloud providers and AI Labs. Stay Tuned.
The regulatory driver is GDPR Chapter V (Articles 44 through 49): transfers of personal data outside the EU/EEA require an adequacy decision or appropriate safeguards. Inference on EU personal data using US-hosted infrastructure is a transfer. ↩︎
GDPR Article 22: the data subject has the right not to be subject to a decision based solely on automated processing which produces legal or similarly significant effects. EU AI Act Annex III classifies most of these decision contexts as high-risk, and Article 26(2) requires deployers to assign human oversight to competent, trained persons. ↩︎
EU AI Act Article 12 requires high-risk systems to be technically capable of automatic event logging with traceability across the system’s lifetime. Article 26(6) is the deployer half: logs automatically generated by a high-risk system must be kept, to the extent they are under the deployer’s control, for at least six months. ↩︎