News|Articles|August 27, 2026

Pharma's Next AI Problem Is Authority

Author(s)Brian Drapeau

Agentic AI changes the question from what a model can produce to what a system is allowed to decide.

In December 2025, FDA announced that it was deploying agentic AI capabilities across the agency. FDA defined these systems as AI designed to achieve goals through planning, reasoning, and multi-step actions, and said its deployment would allow employees to create more complex workflows using multiple models. The agency specifically identified applications ranging from premarket review and postmarket surveillance to inspections and compliance.1

Four months later, FDA issued a warning letter to a pharmaceutical manufacturer with an unusually direct subsection: “Inappropriate Use of Artificial Intelligence in Pharmaceutical Manufacturing.” The manufacturer had used AI agents to create drug-product specifications, procedures, and master production or control records. FDA's objection was not simply that AI had been used. The firm had failed to adequately review what the agents produced, and FDA stated that AI outputs used for current good manufacturing practice (cGMP) activities must be reviewed and cleared by an authorized human representative of the quality unit.2

There is something useful in the proximity of those two events. The same regulator deploying agents internally had just demonstrated where accountability remained when agents entered a regulated workflow. The agent did not receive the warning letter. The manufacturer did.

For the first wave of generative AI, pharma spent considerable time asking whether models could hallucinate, whether their outputs could be trusted, and whether humans needed to review them. Those questions do not disappear with agentic AI, but they are no longer enough. A chatbot produces an answer. An agent can participate in the process that determines what happens next.

Once that happens, the more important question becomes: what did we give it permission to decide?

A Chatbot Answers. An Agent Acts.

Most generative-AI governance was built around a relatively simple interaction. A person asks a model a question, the model generates something, and the person reviews it. Prompt, output, review. Whatever uncertainty existed inside the model, there was an obvious control point downstream where a human could evaluate the artifact before it entered the regulated process.

Agentic AI changes that architecture. FDA's own definition captures the distinction: these systems can plan, reason, and execute multi-step actions toward a defined goal. Recent work on agentic systems in regulated environments similarly distinguishes systems that reason over supplied information from systems capable of committed writes to authoritative records.1,3

Put that inside a pharmaceutical quality system and the distinction stops being theoretical. Imagine an AI assistant drafting a deviation investigation from information a human provides. Now imagine an agent that retrieves the deviation itself, searches previous investigations, identifies relevant procedures, selects potentially related events, develops possible causal relationships, drafts the investigation, recommends a classification, and routes the record to the next workflow step.

Both might appear on an AI inventory as “AI supporting deviation investigations,” but they are not remotely the same use case. The first generates content. The second participates in work, and participation creates something output accuracy alone cannot describe: authority.

Specificity Comes Before Autonomy

There is an upstream problem here that is easy to miss. The industry talks about whether “AI works in pharma” as though AI were a capability rather than a category of technologies. It is not particularly useful to ask whether an LLM is good at pharmaceutical research, deviation management, quality review, or drug development in the abstract. The useful question is narrower: which part of the task matches what this particular system has demonstrated it can do?

An LLM may perform well at retrieving and summarizing scientific literature. That does not by itself establish that it can independently reason about the underlying biology. An agent may reliably retrieve previous deviation investigations and summarize potentially related events. That does not establish that it can determine root cause. Same system, same records, completely different claim.

Draft EU GMP Annex 22 already moves in this direction by requiring the intended use of an AI model and the specific tasks it assists or automates to be described in detail, based on an in-depth understanding of the process in which the model operates. The draft also requires personnel involved with AI systems to have adequate qualifications, defined responsibilities, and appropriate access.4

This is why successful AI use cases often look less ambitious than failed ones. The task is narrow. The expected output is defined. Failure can be detected. Human intervention occurs before the consequential decision. There is some way to determine whether the system actually performed acceptably. That is not a lack of ambition; it is specificity.

Specificity has to come before autonomy, because an organization cannot responsibly delegate authority for a function it has never demonstrated the system can perform.

“Human in the Loop” Tells Us Less than We Think

Draft Annex 22 explicitly contemplates human-in-the-loop controls for non-critical generative-AI applications, assigning qualified personnel responsibility for determining whether outputs are suitable for their intended use.4 The problem is that saying a human is “in the loop” tells us that a person exists somewhere in the workflow. It does not necessarily tell us what that person still controls.

Consider two workflows:

  • Observe → Human decides → AI executes.
  • AI observes → AI interprets → AI decides → AI executes → Human approves.

Both contain a human. They do not contain the same human control.

A reviewer who enters after an agent has already selected the evidence, excluded alternatives, interpreted the information, classified the event, and initiated the workflow may technically remain in the loop. But many consequential choices have already occurred. Human review can only operate on what reaches the human.

If an agent searches 10 years of deviation records and returns five “relevant” events, a reviewer can competently assess those five records and still have no idea that the sixth, most consequential record was excluded upstream. The final output can be correct relative to the evidence presented and still be wrong relative to the evidence that existed. The model did not hallucinate anything. It made a selection decision.

Human-in-the-Loop Is Not a Control Strategy by Itself. It Is a Description of Topology

The control question is where the human sits relative to the consequential decision, what information reaches them, and whether intervention still occurs early enough to change the outcome.

GMP Already Governs Authority

This is where agentic AI collides with something pharmaceutical quality systems have been defining for decades: decision rights.

Regulation 21 CFR 211.22 does not merely assign the quality unit responsibility. It assigns responsibility and authority. The quality control unit has authority to approve or reject components and drug products, review production records, and approve or reject procedures and specifications affecting product identity, strength, quality, and purity.5

That language matters. GMP deliberately distinguishes between performing work and possessing authority over the consequential decision attached to that work. A manufacturing operator may understand why a batch should be rejected and still lack authority to disposition it. An investigator may recommend a corrective action and preventive action (CAPA) without having authority to approve closure.

Agentic AI introduces a new actor into that architecture, yet many AI inventories still describe systems primarily by function: documentation assistant, investigation copilot, quality agent, regulatory assistant. Those labels tell us what the system is supposed to help with. They tell us almost nothing about what the system is permitted to do.

FDA's Purolea letter makes the distinction unusually concrete. FDA did not merely require that AI-generated cGMP documents eventually encounter a person. It said outputs or recommendations from an AI agent must be reviewed and cleared by an authorized human representative of the quality unit.2

Authority was already part of the regulatory architecture before the agent arrived.

Map the Decision Rights

The Table maps the distinct decision rights an AI agent may support, from observing information through approving an accountable GMP determination.

Function

Agent Activity

Example

Observe

Retrieve or monitor information

Pull deviations associated with an equipment train

Interpret

Assign meaning to information

Identify a possible recurring failure pattern

Recommend

Propose an action

Recommend opening a corrective action and preventive action (CAPA)

Decide

Select an outcome

Classify an event against predefined categories

Execute

Change the state of a regulated workflow

Open the CAPA and assign actions

Approve

Make an accountable GMP determination

Approve closure or disposition

Retrieving five related deviations is not equivalent to deciding that those deviations constitute a recurring quality event. Recommending that a CAPA be opened is not equivalent to opening one, and opening a CAPA is not equivalent to approving its closure.

There is emerging technical precedent for thinking about agents this way. Recent work separates what an agent is technically capable of doing from the autonomy it should actually be allowed to exercise,6 while other work on regulated agentic architectures distinguishes low-agency reasoning from actions that commit writes to authoritative records.3

Yet an implementation containing any of those activities could still be summarized internally as “AI used in CAPA management.” That description is nearly useless for governance because the meaningful unit is not simply the use case. It is the decision right embedded inside the use case.

Capability and Permission Are Different Questions

Capability asks what the agent can do. Permission asks what the organization has decided it may do.

A sufficiently capable agent might technically be able to retrieve an investigation, classify it, modify the record, open a CAPA, assign actions, and route the record. Nothing requires the organization to permit all of those actions simply because the technology can perform them.

Recent agent-governance research makes this distinction by separating an agent's technical capability from the autonomy it should actually be allowed to exercise, based on risk, reversibility, oversight, and accountability.6

That distinction maps well onto GMP. Capability should establish the outer technical boundary. Risk and accountability should determine how much of that capability the organization actually authorizes.

Define the Authority Envelope

Agent governance is beginning to develop language for this bounded permission structure. The useful idea is an authority envelope: the set of actions and permissions within which an agent may operate without additional authorization.

For GMP, I would define that envelope at the level of decision rights. It should specify what the agent can observe, interpret, recommend, decide, and execute; which systems and data it can access; which actions require human authorization; which actions are prohibited; when the system must stop and escalate; and who remains accountable for the outcome.

The narrower the envelope, the easier the assurance problem. As the envelope expands, the evidence supporting it should expand with it.

This is also where cybersecurity provides a useful precedent. The National Institute of Standards and Technology's (NIST) 2025 preliminary Cybersecurity Framework Profile for AI recommends treating AI systems as separate entities with their own permissions and authorization policies. For AI agents specifically, it recommends least privilege: agents should receive only the permissions necessary to perform their assigned role.7

GMP does not currently use the term “authority envelope” as a regulatory requirement. The point is not to pretend it does. The point is that existing GMP decision rights and emerging agent permission models are converging on the same control problem.

The Dangerous Decision May Happen Before the Final Answer

This becomes particularly important when organizations rely on final-output review as the primary control. Suppose an investigation agent produces an excellent final draft. The facts presented are accurate, the language is appropriate, and the causal argument is reasonable. A qualified QA reviewer reads it and approves it.

But the reviewer sees the artifact, not necessarily everything that produced it. Did the agent search every relevant prior event? Did terminology differences cause it to exclude an older investigation? Did it select the correct SOP revision? Did it discard an apparently unrelated result that would have changed the causal hypothesis? Did it decide that an intermediate finding did not warrant escalation?

The dangerous agentic failure can therefore occur several steps before the output anybody reviews. An upstream retrieval, selection, or interpretation can constrain everything downstream while leaving the final artifact looking perfectly reasonable. By the time the human enters the loop, the decision space may already have been narrowed.

This is also why 21 CFR 211.68 deserves more attention in the agent discussion. The regulation allows computerized and automated systems to perform pharmaceutical manufacturing functions, but requires appropriate controls so that changes to master production and control records or other records are made only by authorized personnel. It also requires computer-system inputs and outputs to be checked for accuracy, with the degree and frequency of verification based on the system's complexity and reliability.8

The regulation predates agentic AI by decades, so applying it to agent trajectories requires interpretation. But the control logic is strikingly relevant: authorization, accuracy, complexity, and reliability determine how much assurance the computerized process requires.

Audit Trails Record Events. Agents Raise a Provenance Problem

Pharma already understands audit trails. If User X changed field Y from A to B at 2:32 pm, the audit trail tells you what happened.

An agent introduces a different evidentiary problem because the consequential information may include what objective the agent received, which information it retrieved, which information it excluded, what tools it called, which intermediate conclusions affected subsequent actions, what model and configuration were operating, where human intervention occurred, and what happened after that intervention.

There is already a name for the broader concept: decision provenance. Singh, Cobbe, and Norval proposed it in 2018 as a way of exposing decision pipelines: the chains of inputs, decisions, actions, and downstream effects that operate across complex automated systems. Their argument was explicitly tied to accountability, audit, compliance, and oversight.9

The GMP application is straightforward. Logging that an agent changed a field establishes the event. Decision provenance would attempt to establish the information and intermediate actions that produced it.

That distinction becomes particularly important when the final output does not reveal the path that generated it.

Least Privilege Is Not Theoretical

Suppose a QA director has broad permissions across the quality management system because the role requires them, and the director launches an agent. Should the agent automatically inherit those permissions?

There is little reason it should. The director's authorization reflects role, training, qualification, organizational responsibility, and accountability. The agent possesses none of those merely because an authorized person initiated it.

NIST's emerging guidance points in the same direction: AI agents should have their own permissions and authorization policies, constrained by least privilege.7 And the technical evidence suggests this is not merely administrative housekeeping.

ToolPrivBench, published in June 2026, tested whether LLM agents select higher-privilege tools when a lower-privilege option would suffice. Across multiple domains and models, unnecessary high-privilege tool selection was common and became worse after transient tool failures. General safety alignment did not reliably prevent the behavior.10

Earlier work on Progent approached the same problem architecturally, proposing deterministic privilege policies around agent tool calls so that agents can perform necessary actions while unnecessary actions remain unavailable.11

That suggests a useful GMP principle: do not ask the agent to exercise restraint over authority it never needed in the first place. A deviation-triage agent may require read access across deviation records without requiring modification rights. An investigation-support agent may need to draft content without being able to approve the record containing that content.

The control is stronger when unnecessary authority is unavailable rather than merely discouraged.

Segregation of Duties Gets Strange Very Quickly

Agentic systems also complicate one of the oldest controls in quality systems: segregation of duties. A human investigator drafts an investigation. QA reviews it. Another authorized role may approve a consequential action. Those separations exist because concentrating incompatible responsibilities in one actor creates risk.

Now consider an agent that retrieves the evidence, determines which evidence is relevant, interprets it, drafts the investigation, recommends the classification, and routes the record. How many roles is that agent performing? More importantly, can the same system generate an analysis and then evaluate whether its own analysis is sufficient?

Using a second agent does not automatically solve the problem. If both agents depend on the same foundation model, retrieval architecture, or underlying data, the second system may reproduce the same blind spot rather than provide genuinely independent review.

Current GMP regulations do not prescribe an agent-specific segregation-of-duties architecture. This is an inference from existing quality and access-control principles, not a regulatory requirement. But it is precisely the kind of question manufacturers should resolve before automation quietly collapses boundaries that were deliberately separated when humans performed the work.

Qualification Changes Again

This extends the qualification model I have argued for previously. For generative AI, the qualified state is not “AI trained.” It is a relationship:

Person × System × Version × Use Case

The use case itself has to be bounded by what the system has demonstrated it can actually do. Agentic AI adds another variable because the same system can perform the same nominal use case while possessing very different permissions.

The qualified state therefore becomes:

Person × Agent × Version × Use Case × Authority Envelope

Suppose a QA reviewer demonstrates competent use of an investigation agent that retrieves historical records and drafts recommendations. Six months later, the organization enables the same agent to initiate workflows automatically. Same reviewer, same product, same model, same apparent use case, but different authority.

That is not the same qualified state. The system crossed a decision boundary, and the assurance case should cross with it.

Draft Annex 22 already provides part of the regulatory foundation for this argument. It requires the specific tasks an AI system assists or automates to be defined, assigns responsibilities and access levels to qualified personnel, and makes human responsibility explicit for generative AI used in non-critical GMP applications.4 What it does not yet provide is an agent-specific model for connecting those requirements to progressively greater decision authority.

That is the gap.

The Regulator Is Already Using Agents

There is an irony in all of this. FDA is not standing outside the technology telling industry to avoid it. The agency deployed agentic capabilities in December 2025 and has continued integrating AI more deeply into its own workflows.1

Meanwhile, FDA's first explicit manufacturing enforcement language around AI agents made human quality-unit authorization unmistakable.2

Those developments should not be collapsed into a claim that FDA has established a comprehensive regulatory framework for agents. It has not. But together they make something else clear: agentic technology is entering regulated work while the underlying accountability architecture remains human and organizational.

That matters because the value proposition of agents comes precisely from delegation. If a person still performs every retrieval, interpretation, decision, and action, there is not much agency left. The economic value comes from delegated work, but the regulatory problem changes when delegated work becomes delegated judgment.

The Alternative Is Bounded Deployment

The answer is not to slow everything down until another governance committee produces another framework. Pharma already knows how to deliberate. Nor is the answer to deploy broadly and trust that human review will catch whatever goes wrong.

The better approach is smaller and more concrete: narrow tasks, demonstrated capability boundaries, separate agent permissions, observable actions, reversible consequences where possible, named owners, predefined acceptance criteria, and explicit conditions that stop the system or force escalation.

This is where capability and authority need to remain separate. A system may be capable of acting autonomously without being authorized to do so. Emerging governance research, NIST guidance, and agent-security research independently support limiting permissions according to the actual task.6,7,10,11

Move quickly where failure is bounded. Demand more evidence as authority expands.

Specificity comes before autonomy.

The Next AI Failure Will Look Ordinary

The first major agentic-AI quality failure probably will not look like science fiction. There will be no rogue machine independently deciding to release a batch. It will look like a normal workflow: a deviation gets classified, a record gets routed, a trend gets missed, a CAPA does not get opened, or a specification moves forward. Each individual action may look reasonable, and the final reviewer may even approve the result.

Somewhere three steps upstream, an agent will have made a decision nobody realized had become a decision. That is the frontier pharma needs to govern now, not because agents are uniquely dangerous, but because they are unusually easy to mistake for tools after they have started exercising authority.

The first wave of AI asked whether the model could be trusted. The next requires a harder question: What did we give it permission to decide?


References

1. U.S. Food and Drug Administration. FDA expands artificial intelligence capabilities with agentic AI deployment. Published December 1, 2025. Accessed Aug 26, 2026. View source

2. U.S. Food and Drug Administration. Warning letter: Purolea Cosmetics Lab, 722591. Published April 2, 2026. Accessed Aug 26, 2026. View source

3. Safin D, Balta D. Autonomy and agency in agentic AI: architectural tactics for regulated contexts. Preprint. Posted online May 12, 2026. View source

4. European Commission. Draft EU GMP Annex 22: artificial intelligence. Stakeholder consultation draft. Accessed Aug 26, 2026. View source

5. 21 CFR § 211.22. Responsibilities of quality control unit. Electronic Code of Federal Regulations. Accessed Aug 26, 2026. View source

6. Zheng H, Dong Q, Depena RK, et al. Separating capability from permission: a governance framework for agentic AI autonomy levels. Preprint. Posted online July 26, 2026. View source

7. National Institute of Standards and Technology. Preliminary cybersecurity framework profile for artificial intelligence (NIST IR 8596, initial public draft). Published 2025. Accessed Aug 26, 2026. View source

8. 21 CFR § 211.68. Automatic, mechanical, and electronic equipment. Electronic Code of Federal Regulations. Accessed Aug 26, 2026. View source

9. Singh J, Cobbe J, Norval C. Decision provenance: harnessing data flow for accountable systems. Preprint. Posted online April 2018. View source

10. Yang K, Bu Y, Yi J. When lower privileges suffice: Investigating over-privileged tool selection in LLM agents. Preprint. Posted online June 2026. View source

11. Shi T, He J, Wang Z, et al. Progent: Securing AI agents with privilege control. Preprint. Posted online April 2025. View source