
PharmTech AI Pulse Check: Availability Bias Threatens AI Regulatory Reviews
API-gated journals may skew pharma regulatory research as AI agents favor accessible sources, raising citation integrity and pharmacovigilance risks.
As publishers and research platforms such as
Florin Muraru, an independent regulatory advisor specializing in EU and US regulatory strategy and AI governance, expects the shift to affect regulatory work more than science. "If the literature became an API, the evidence nobody put behind an API quietly stops counting," he says. Many regulatory documents are public, he notes, but are locked in PDFs that models do not read. A pharmacovigilance literature search that covers only API-enabled publishers, he argues, will not be considered complete, which makes the limitation potentially dangerous from a regulatory standpoint, since regulators look for information well beyond journal articles.
Richard Jaenisch, senior director of education, outreach, and digital experience at Open Biopharma, stands by his prediction, pointing out that FDA already offers an API-accessible data dashboard. "If the FDA is ahead of you, I'm just saying, that's not usually a good sign," he says. He describes AI-native systems as more plug-and-play and, after visiting Springer Nature's booth at BIO, credits publishers with building curated, peer-review-informed tools tailored to specific literature search use cases. APIs let developers pull content as precisely or as broadly as they choose, but the PDF status quo offers little real protection: generative tools can already extract text from any document visible through a viewer, regardless of its security, and copyright concerns have not historically slowed model training. The open question, in Jaenisch’s view, is whether FDA will care about data provenance, since the agency may ultimately be the only entity that checks.
Gourav Pandey, R&D quality lead at Takeda, frames the risk differently. "The danger of API era is that we will quietly start optimizing our research more from what is ingestible instead of what is correct," he warns. He points to individual research articles now being converted into Model Context Protocol (MCP) formats that serve ready-to-use data to AI agents. In his own study comparing citations against what AI agents return from regulatory documentation, a problem he calls citation laundry, he finds that the retrieval architecture built around a model matters more than the model itself, because it determines whether an answer can be trusted. The result, he says, is availability bias at scale: under pressure to respond quickly, agents favor easily accessible sources over high-value articles that lack API or MCP access. Because companies rely on literature citations in regulatory justifications, including arguments to FDA when in-house data or stability reports are unavailable, he expects added human verification of citations to become necessary within the next few years, even as he welcomes more journals moving toward API access.
The panel ties the concern to Goodhart's law, the principle that a measure loses its value once it becomes the target, with citation scores offering a ready example of a metric agents could game. Pandey adds that latency compounds the problem: with deep research agents expected to finish their work in roughly 30 minutes, journals will need to keep pace so that quality research is actually available to them.
Related to this article








