
Lessons For AI Governance in Quality Systems
Key Takeaways
- Pre-deployment due diligence should interrogate signal precedent, enterprise dissemination controls, and node-level inspection/litigation exposure while explicitly mitigating confirmation bias and sunk-cost thinking triggered by model outputs.
- A two-layer architecture can separate RAG-driven document interrogation of inspection PDFs from structured-data exploration, enabling traceable month/site/authority trend views tied to responses and after-action reviews.
Karthik Iyer, Eli Lilly, details an AI system linking GMP findings, audits, and deviations for cited, inspection-ready regulatory intelligence.
At the PDA/FDA Joint Regulatory Conference 2026 in Washington, DC, Karthik Iyer, Associate Vice President, Compliance and Post Market Reporting, Eli Lilly, laid out a working model for how a global manufacturer is using AI to connect inspection findings, audit observations, and deviation data that traditionally live in separate systems.1 His presentation, "From Data Noise to Inspection Readiness: AI-Powered Regulatory Intelligence Across Your Quality Landscape," treated AI adoption less as a technology rollout and more as a governance exercise, this is a recurrent theme PharmTech tracked from the FDA's
What Questions Should a Quality Organization Ask Before Deploying AI?
Iyer claimed the most important questions were: Has a given signal occurred before, internally or externally? How does the enterprise share that signal without creating new exposure? What are the possible adverse inspection and possible litigation risks embedded in each data node? And how does a team guard against confirmation bias or sunk-cost thinking once a model starts producing answers it likes?1 That final point lines up with concerns raised during the FDA's push toward
How Does the System Turn Documents into Defensible Answers?
Iyer walked through track findings by month, site, and health authority alongside Lilly's response and after-action review, with two AI layers sitting underneath it: a Document Explorer that applies retrieval-augmented generation to unstructured files such as inspection PDFs, and a Data Explorer for structured data sets.1
What Did Lilly's Team Learn?
Iyer was candid about the tradeoffs.1 Connecting the system to Lilly's data lake meant wrestling with translation issues, global stop-words, and parameter tuning to balance document volume against accuracy. He urged discipline in weighing outside expertise against internal judgment, and cautioned against using the tools to generate predictions or restate what a reviewer already knows, keeping the focus instead on surfacing a defensible range of outcomes.
Why the Emphasis on Hallucination Research?
Iyer cited the AA-Omniscience benchmark, which found hallucination rates ranging from 22% to 94% across 26 leading models, and noted a sharper finding: accuracy on identical false claims collapsed once a statement was framed as the user's own belief rather than a third party's, with GPT-4 falling from 98.2% to 64.4% and DeepSeek R1 from over 90% to 14.4%. He also pointed to a growing body of research showing that even long-context models lose accuracy as inputs grow. For a compliance function built on getting the source citation right, that's less a footnote than the whole point.
References
Parenteral Drug Association. PDA/FDA Join Regulatory Confernce 2026 Agenda. Available at




