Agentic AI vs Generative AI: What Changes for Governance

Agentic AI vs generative AI: a pantograph arm drawing on its own while a hand guides the stylus

Key takeaways

  • Generative AI produces an output you review. Agentic AI takes an action that has already happened by the time you look at it. That single shift is what moves the compliance work.
  • Under the EU AI Act, agents are not a new category. Article 3(1) already covers systems operating with varying levels of autonomy, so the question is never whether the rules apply, only which ones bite harder.
  • The agentic AI vs generative AI boundary is a gradient, not a switch. Govern it by degree of agency: tool access, memory, self-directed planning, unreviewed effect, agent-to-agent handoff.
  • Your risk register changes in kind, not only in severity. Prompt injection stops being a bad-output problem and becomes an unauthorised-action problem.
  • Your evidence pack changes too. Prompt and output logs are no longer enough, and an auditor will ask for action logs, the identity the agent acted under, and proof that the stop mechanism works.

Agentic AI vs generative AI: the difference that changes your obligations

Ask a vendor to explain agentic AI vs generative AI and you get the same sentence everywhere: generative AI creates content, agentic AI takes actions. It is accurate. It is also useless to anyone who has to sign off on a deployment. The useful version is narrower. Generative AI puts a draft in front of a person, and that person decides whether it becomes real. Agentic AI closes that gap. It plans a sequence, calls a tool, writes to a system of record, and reports back. The output and the effect arrive together. Everything governance quietly depended on was sitting in that gap: the review step, the approval record, the chance to catch a wrong answer before it cost anything. Remove the gap and the controls do not fail loudly. They simply stop being where the risk is. Regulators did not miss this. The EU AI Act defines an AI system in Article 3(1) as a machine-based system that operates with varying levels of autonomy and infers outputs that influence physical or virtual environments. Autonomy sits inside the definition. An agent is not a new regulatory object, it is the same object with the dial turned up, which is why obligations do not change identity so much as change weight.

The capability stack that creates the delta

The OWASP Agentic Security Initiative describes an agent through four capabilities: planning and reasoning, memory and statefulness, tool use, and action. Each is a governance surface as much as a technical one.

  • Planning means the system chose steps you never enumerated, so a risk assessment built on an exhaustive list of expected behaviours no longer holds.
  • Memory means state persists between runs, so a poisoned input can change behaviour long after the session that introduced it has closed.
  • Tool use means the blast radius is whatever the credentials allow, not whatever the text box allows.
  • Action means the effect is real before it is reviewed.

A generative AI model has the first capability in a weak form and none of the other three. That is the entire delta. Every section below is a consequence of it.

The five-question test: has your deployment crossed into agentic?

Most teams do not decide to build an agent. They add a tool call to a chatbot, then a retry loop, then a scheduler, and one sprint later the thing writes to production. The crossover is rarely a decision, which is why it rarely triggers a governance review. Run these five questions against any deployment you own.

  1. Can it call a tool, an API or a database without a human approving that specific call? If yes, your control point moved from the output to the permission.
  2. Does it keep state across steps or sessions that changes later behaviour? If yes, you have an asset that can be corrupted, and it needs integrity controls, not just access controls.
  3. Did it decompose the goal into steps you did not enumerate? If yes, your risk assessment cannot be behaviour-by-behaviour and has to become boundary-by-boundary.
  4. Can its output take effect without a person reviewing it first? If yes, human oversight is now a design problem rather than a process one.
  5. Does it hand work to, or accept work from, another agent? If yes, you are governing a system, and the interesting failures live in the interaction rather than in any single component.

One yes is not a crisis. Five is a different system than the one your documentation describes. The useful framing comes from UC Berkeley’s Center for Long-Term Cybersecurity, whose Agentic AI Risk-Management Standards Profile argues that governance should scale with the degree of agency rather than treat autonomy as a binary attribute. Score the five answers, record the score in your inventory, and let the score drive how heavy the controls get. A registry field that reads autonomous yes or no will not survive contact with a real portfolio. This is also where shadow AI stops being a discovery problem and becomes a classification problem. Teams that registered a tool as a writing assistant eighteen months ago are not going to refile it because someone enabled function calling.

What changes in risk classification

Here is where the agentic AI vs generative AI distinction produces its first concrete legal consequence, and it arrives on two axes at once. The Future Society’s analysis of how the AI Act reaches agents, Ahead of the Curve (Oueslati and Staes-Polet, 2025), sets out both. Chapter V reaches the general-purpose model underneath the agent, and providers of models with systemic risk must assess and mitigate the risks that arise when those models are integrated into agent systems. Chapter III reaches the agent system itself through high-risk classification. Three consequences follow, and each one catches somebody. First, high-risk classification for a general-purpose agent is genuinely unsettled. Annex III was drafted before agent risks were well understood, and an agent that can be pointed at many tasks may land in scope by default unless the provider deliberately excludes the high-risk uses. Deliberate exclusion means documented and technically enforced, not a line in an acceptable use policy. If you are working through this, start with how to classify a high-risk AI system and apply it to the agent’s widest reachable use, not its intended one. Second, the scaffolding you add may change who you are. Wrapping a general-purpose AI model in retrieval, an orchestration framework and a toolset can amount to a substantial modification, which pulls provider duties onto a team that had budgeted for deployer duties. That is a materially different compliance bill, and it usually lands on an engineering group that never read Chapter V. Third, the classification has to be redone. An agentic capability added to an existing generative deployment is not a version bump. It is a new system for classification purposes, and the conformity work restarts from the risk assessment rather than resuming from the last sign-off.

What changes in human oversight

Article 14 requires high-risk AI systems to be designed so that natural persons can effectively oversee them, with measures commensurate with the risks, the level of autonomy and the context of use. Read that clause with an agent in mind and the difficulty is obvious: the provision scales its demands with exactly the property that makes the demands hard to meet. Generative oversight is cheap because it is synchronous. Someone reads the draft, then acts. Agentic oversight is expensive because the action has already occurred, so oversight has to be designed into the execution path rather than layered after it. In practice that means three changes to human oversight under the EU AI Act. The checkpoint replaces the review. Oversight moves from reading output to authorising categories of action in advance and interrupting specific ones in flight. That is a permission model, and it needs to be designed with the same care as the prompt. The oversight mode shifts, but not everywhere. Moving from human-in-the-loop to human-on-the-loop is the right call for high-volume, low-consequence work. It is the wrong call where a decision produces legal or similarly significant effects for a person, because GDPR Article 22 still constrains solely automated decision-making regardless of how the AI Act classifies the system. Oversight has to survive its own volume. OWASP catalogues this as T10, Overwhelming Human-in-the-Loop, and it is the most underrated item on the list. An approval queue that a reviewer clears at 400 items an hour is not oversight, it is a rubber stamp with an audit trail. If your control depends on attention that the throughput makes impossible, the control has already failed and your evidence will show it failing.

What changes in the risk register

This is the section most comparisons skip, and it is where the agentic AI vs generative AI difference stops being conceptual. The failure modes do not get worse. They change category. A generative register is mostly about what the system says: hallucination, confidentiality leakage, discriminatory or toxic output, and prompt injection as a content problem. An agentic register is about what the system does. OWASP’s taxonomy of fifteen agentic threats names memory poisoning, tool misuse, privilege compromise, cascading hallucination, intent breaking and goal manipulation, repudiation and untraceability, and rogue agents in multi-agent systems. The single sentence worth taking to your risk committee: prompt injection stops being a bad-sentence problem and becomes an unauthorised-transaction problem. Same attack, different consequence class, different owner. <table header-row=”true”> <tr> <td>Failure mode</td> <td>In generative AI</td> <td>In agentic AI</td> <td>What the control becomes</td> </tr> <tr> <td>Prompt injection</td> <td>Off-policy or embarrassing text</td> <td>Unauthorised tool call, data exfiltration, code execution</td> <td>Input trust boundaries plus least-privilege tool scoping</td> </tr> <tr> <td>Hallucination</td> <td>A wrong statement a reviewer catches</td> <td>A wrong action taken, then built upon by later steps</td> <td>Action validation and reversibility, not just output review</td> </tr> <tr> <td>Data leakage</td> <td>Sensitive text in a response</td> <td>Sensitive data moved between systems by the agent itself</td> <td>Egress control at the tool layer</td> </tr> <tr> <td>Memory corruption</td> <td>Not applicable</td> <td>Poisoned state steers behaviour across sessions</td> <td>Memory validation and session isolation</td> </tr> <tr> <td>Identity</td> <td>One service account, one call pattern</td> <td>A non-human identity acting continuously across systems</td> <td>Agent identity, scoped credentials, full attribution</td> </tr> <tr> <td>Compounding</td> <td>Contained to one response</td> <td>Cascading failure across steps or across agents</td> <td>System-level assessment, circuit breakers, containment</td> </tr> </table> The Berkeley profile adds the tail risks that matter at higher degrees of agency: loss of control, oversight subversion, deceptive alignment, and cascading multi-agent failures. It also makes a point worth carrying into design reviews, which is that a multi-agent system has to be evaluated both per agent and collectively, because emergent interaction effects do not show up in component testing. Its blunt recommendation is to treat a sufficiently capable agent as untrusted and contain it accordingly. Two practical consequences for existing programmes. Your AI risk management methodology has to assess at system level, since assessing each agent alone will miss the interaction failures. And AI red teaming has to target the tool boundary and the memory store, not only the prompt, because that is where an agent actually breaks.

What changes in the evidence you must produce

Compliance is not what you believe about your system. It is what you can show. Here the two paradigms diverge sharply, and teams usually discover the gap during an audit rather than before one. For a generative deployment, the evidence pack is familiar: model documentation, evaluation results, prompt and output logs, and the human review record. For an agentic deployment, Article 12 record-keeping has to carry considerably more. An auditor reconstructing a single agent decision needs the goal it was given, the plan it produced, every tool call with parameters and results, the identity it acted under, the permission that authorised the call, whether the action was reversible and whether it was reversed, and where the human checkpoint sat. OWASP files the failure to produce this as T8, Repudiation and Untraceability, and treats it as a security threat. It is equally a governance failure. If you cannot attribute an action to an agent, a version and an authorising permission, you cannot answer a regulator, an auditor or a claimant, and the absence of the log becomes the finding. Three artefacts to add to the pack that a generative deployment never needed:

  • A documented go or no-go decision before deployment, which the Berkeley profile places at Manage 1.1. Written down, with the conditions that would reverse it.
  • Proof the disengage path works, tested rather than asserted. A stop button nobody has pulled in a drill is a design intention, not a control.
  • A non-human identity inventory, so every agent action resolves to a credential, a scope and an owner.

This is also the point where your AI system documentation and your incident reporting procedure need rewriting rather than extending, because both were written around a system that produces output rather than one that produces effects.

What changes in accountability across the value chain

The hardest problem is not technical. The Future Society calls it the many hands problem: by the time an agent acts, so many parties have contributed that accountability disperses unless somebody allocates it deliberately. Three actors, with asymmetric resources, expertise and context. The model provider builds the underlying capability and the controls that make governance possible at all. The agent system provider adapts it to a purpose and sets the boundaries appropriate to that purpose. The deployer runs it against real data, real users and real consequences. Monitoring is the clean illustration. The model provider has to build configurable monitoring infrastructure. The system provider has to set alert thresholds that suit the use case. The deployer has to actually watch the alerts and act on them. Any one of the three acting alone produces something that looks like oversight in a diagram and is not oversight in operation. For most organisations reading this, the practical question is what to demand from a vendor selling an agent. Four things belong in the contract: the log schema and retention you will receive, the granularity at which you can scope the agent’s permissions, the disengage mechanism and its tested latency, and notification duties when the model or the scaffolding changes underneath you. Treat these as AI vendor due diligence items with the same weight as security questions, because a capability change shipped silently is a classification change you did not make. What you cannot delegate is the deployer duty. Context-specific risk, the fundamental rights position and the oversight design stay with the organisation running the system, which is the recurring theme of AI accountability once agents enter the estate.

The crossover checklist

Ten actions when a deployment crosses the line. Each one is verifiable, which matters more than whether it is comfortable.

  1. Re-run classification against the agent’s widest reachable use, not its intended use.
  2. Record a degree-of-agency score in the inventory, and stop using a binary autonomy flag.
  3. Re-do the risk assessment at system level, including agent-to-agent interactions.
  4. Register every agent as a non-human identity with a named owner.
  5. Scope tool permissions to least privilege, and set the blast radius deliberately rather than inheriting it.
  6. Instrument per-action logging that carries goal, plan, tool call, identity and reversibility.
  7. Define and test the disengage path, and keep the test record.
  8. Redesign the oversight checkpoint so it survives production throughput.
  9. Re-paper the vendor agreement for logs, permissions, kill switch and change notification.
  10. Extend the incident procedure to cover actions the agent took, not only outputs it produced.

If your organisation manages this in spreadsheets today, the agentic step is usually where that stops working, because the evidence is continuous rather than periodic. That is the problem an AI risk management platform exists to solve.

FAQ

Is ChatGPT generative AI or agentic AI? Both, depending on how it is configured. Used as a chat interface that returns text for a person to read, it is generative. Given tools, browsing, code execution or connectors, and allowed to chain steps toward a goal, the same product behaves agentically. This is precisely why the agentic AI vs generative AI question cannot be answered at the level of a product name. It has to be answered at the level of your configuration, your enabled tools and your permissions, which is also why two companies using the same subscription can owe very different obligations. What is the difference between an AI agent and agentic AI? The Berkeley profile draws the line usefully. An AI agent is a single model equipped with tools to complete a well-defined task end to end. Agentic AI describes a system of multiple agents coordinating toward broader goals. The distinction matters for assessment: a single agent can largely be evaluated on its own, while a multi-agent system has to be evaluated both per agent and collectively, because the failures that matter most emerge from how the agents interact rather than from any one of them misbehaving. Can you give an example of agentic AI? A procurement assistant that reads an incoming invoice, checks it against the purchase order, queries the supplier record, flags a discrepancy and posts an approved payment instruction to the finance system. Each step is unremarkable. The combination is agentic because the system planned the sequence, used several tools and produced a financial effect without a person approving that specific payment. The generative version of the same assistant would have drafted a summary for an accounts clerk to act on. Is agentic AI high-risk under the EU AI Act? Not automatically, and not by virtue of being agentic. High-risk status depends on the use case, whether the system is a safety component or falls under an Annex III area. The complication for agents is that a general-purpose agent may reach many use cases, and analysts have flagged that such systems could fall in scope by default unless the provider deliberately and demonstrably excludes high-risk uses. Classify against what the agent can reach, and treat the exclusion as a technical control you can evidence rather than a policy statement. Does agentic AI need a different risk assessment than generative AI? Yes, and not merely a longer one. A generative assessment is largely output-centric, asking what the system might say and who could be harmed by it. An agentic assessment has to be action-centric and system-level: what the agent can reach, what it can change, what happens when a step fails midway, and how failures compound across steps or agents. The Berkeley profile maps this onto the NIST AI RMF functions, which means most organisations can extend an existing method rather than adopt a new one. Do I need a separate AI policy for agents? Usually not a separate document, but the existing one needs new clauses. At minimum: who may grant an agent tool access and under what approval, which action classes always require a human checkpoint, mandatory logging and identity requirements, and a trigger that forces reclassification when an agentic capability is switched on. Folding these into your AI policy keeps one enforceable document rather than two competing ones.

Conclusion

The comparison everyone publishes is correct and stops one step early. Generative AI creates, agentic AI acts, and the interesting question is what that costs you. It costs a reclassification, because the reachable use set widened. It costs an oversight redesign, because Article 14 asks for more exactly as autonomy makes it harder to give. It costs a new risk register, because the failure modes changed category. It costs a heavier evidence pack, because the auditor now asks what the system did and under whose authority. And it costs a renegotiated vendor relationship, because the party that changes the capability is rarely the party that carries the consequence. None of that is a reason to avoid agents. It is a reason to know precisely when you acquired one. Run the five questions against your portfolio this quarter, and see how much of it has already crossed the line. If you want the classification, oversight design, risk register and evidence trail held in one place rather than five spreadsheets, that is what AI Sigil is built for.

HIPAA Compliance Software: The 2026 AI-Era Buyer’s Guide

HIPAA compliance software was built for systems that store PHI, not for systems that infer from it. What a 2026 tool must cover, and what to demand.

AI Inventory: What Regulators Expect to Find in It

An AI inventory is the artefact every AI rule assumes. See which clauses compel one (EU AI Act, NIST, ISO 42001, OMB) and the fields each expects.

China AI Regulation in 2026: Filings, Labels and Liability

China AI regulation explained for foreign firms: CAC filings, AI content labels, companion AI rules, 2026 enforcement and the evidence to keep ready.

California SB 243: Companion Chatbot Duties After Adam’s Law

California SB 243 set the first companion chatbot rules. Adam's Law (SB 1119) added risk assessments, parental controls and independent audits from 2027.

California AI Transparency Act: What SB 942 Now Requires

The California AI Transparency Act took effect on 2 August 2026. What SB 942 requires, what SB 1000 would change, and the records you must keep.

Agentic AI vs Generative AI: What Changes for Governance

Agentic AI vs generative AI is not just a technical difference. See what changes in classification, oversight, risk and evidence when AI acts.