This is the second part of a multi-part (we’ll know when we’re done) series on creating Security Invariants for AI. Comments are available on the Google Doc version. You can find Part 1: Understanding the Stack here
Part 2: How AI is Used, and the Risks
There are a number of Service Deployment Topologies with modern AI applications. These can be thought of as similar to the Cloud’s Service Models. We group them by how the human and deploying organization interacts with the AI.
- SaaS (or Agent / AI as a Service if we want to break free of the cloud) - The organization does not manage or control the user interface, the underlying inference, or even the harness or cloud infrastructure with the possible exception of limited user-specific application configuration settings. One example is Claude.ai, where you access Claude only through a web browser, or the add-on AI agents in SaaS platforms like ServiceNow and Salesforce.
- PaaS (or Harness as a Service) - A third party provides the inference and harness primitives like memory, context window management, etc, but doesn’t provide an end-user interface - the deployer has to provide the human interface. Examples include Claude Platform or the OpenAI (not ChatGPT) interface. It’s not intended for end users, you need to write code to use this topology.
- Personal/Local - The human interface runs on the local machines the user interacts with directly. The deployer either writes their own interface, deploys opensource, or uses one from the inference provider. Claude Code CLI, Claude Desktop, Cursor IDE, etc.
- Shared - The human interface runs in the organization’s managed environment (public/private cloud) that is distant from the user. An AI Enterprise or Customer facing application would fall into this category.
- Autonomous - The human interface exists solely to monitor or configure (define the goal) the agentic harness - but the system operates on its own with minimal human interaction. Ex: Skynet, W.O.P.R.
For enterprises, we see a common set of AI use cases matched to those topologies:
- Knowledge worker support: AI to answer questions, write reports, manage spreadsheets, etc. If you see AI slop in the workplace, it usually starts here.
- Coding support: AI writes code, edits code, builds applications, and assists developers and administrators. This includes “citizen developers” which is a fancy way of saying a sales exec making their own version of Flappy Bird with your company logo on it.
- SaaS agents: Various SaaS platforms adding everything from useful analysis agents to support agents that are only slightly less frustrating than the call center.
- AI in enterprise applications: Building AI tooling into your own applications. This can be one of the more powerful use cases when there are well defined desired outcomes, or it can be a massive boondoggle that results in a misaligned agent sending your financials to that landscape service with the similar name to your favorite hotel. Much of our invariant discussion focuses on this use case.
AI Brings New Risks
The OWASP GenAI Security Project publishes twenty numbered risks across two lists, one for LLM applications and one for agentic applications. Grouped by what actually goes wrong, they collapse into five broad themes.
- Model Provenance and Supply Chain - This is related specifically to the model you choose (or the harness chooses for you). Issues like factual inaccuracies and bias fall into the threats from the wrong model. Other supply chain issues relate to the skills and harnesses you choose. Just like an npm package can include a malicious installation script, a skill can also get the harness or model to disclose sensitive information.
- Malicious Inputs - The model cannot tell your instructions apart from the attacker’s because, to it, they are the same pile of tokens. This theme includes prompt injections, and more harness-level techniques like memory or context poisoning. It also includes poisoning training data (grounding attacks), which is now a major focus of SEO firms as AI replaces search engines.
- Malicious Outputs - Hallucinations fall into this category and can take the form of disinformation “Yes, you can get a refund later”, or more consequential mistakes like “the best solution here is to delete the production databases”. The LLM can also output harmful code, potential prompt injects to other LLMs, deliberate misinformation, and other output that can cause damage.
- Excessive Agency and Permissions with Lack of Supervision - Destructive malicious outputs are just bad advice without agency, and so one major category of threat is giving the agents the ability to impact the real world, letting the agents act unsupervised. This is closely tied to identity related issues (often misconfigurations) and includes confused deputy attacks. Go ask OpenAI and Hugging Face about this one.
- Data Exfiltration or Destruction - Lurking in the weights or data stores accessed by the AI may be sensitive information, and the right prompt could output those relationships in a way you don’t intend. Additionally the context window contains data from multiple sources. The model doesn’t know the difference, and can return information from the prompt you might not like.
The Lethal Trifecta
Simon Willison gave us the cleanest framing of agentic risk: the lethal trifecta. An agent becomes dangerous when it has all three of: access to private data, exposure to untrusted content, and a channel to communicate externally. Any two are survivable. All three mean an attacker can plant instructions in something the agent reads, have it gather your sensitive data, and exfiltrate the results. The model will do it cheerfully because it cannot reliably distinguish instructions from data. This has bitten companies like Microsoft, Salesforce, and ServiceNow.
Look back at the list of capabilities. RAG is untrusted content. Memory and retrieval are private data. Browser use, MCP connectors, and code execution are all exfiltration channels. Most production agents ship with the full trifecta on day one because each hand was added for a legitimate business reason. As we discuss invariants, you’ll see we focus a lot on minimizing the risk of the trifecta.
Regulatory Factors
Governments are stepping in to regulate how AI can be used. How you build your product, or how your employees are using AI can expose you to regulatory risks. The EU’s AI Act and GDPR limit where and how you can use AI.
Additionally, models are becoming a national security concern. Access to frontier models may become unavailable, and access could be restricted as part of a geo-political conflict.
Other Factors
There are other factors beyond threats from using AI and regulatory threats. The AI landscape is moving fast, and new tools are constantly emerging. This creates a Shadow AI problem: employees sign up for consumer-grade or free accounts that lack the same level of data protection as enterprise accounts.
Additionally, there is the cost factor. AI usage is measured in tokens (see Part 1), and these tokens are currently subsidized by a lot of speculative venture capital. When token costs start to reflect the actual costs of electricity, cooling, and capital return on GPUs, the number of cat memes will drastically decrease.
From Threat Assessment to Enforcement
Now that we’ve established how AI works, how we use it, and what the threats are we’re ready to move on to parts 3 and 4 which will cover the enforcement points available to us and the invariants we can build with those enforcement points.