Large language models are moving into clinical trial documentation fast, but they're not ready to work alone. A 2026 study in JAMIA found that off-the-shelf AI produced factual errors in 18 to 43 percent of informed consent forms when compared against source protocols. With structured human review and retrieval-augmented generation, that error rate dropped dramatically, with accuracy exceeding 90 percent.
The problem isn't that AI makes mistakes. It's that those mistakes look polished and plausible, creating automation bias where reviewers trust outputs they haven't truly verified. Protocol errors are especially costly: the average amendment now takes 260 days to implement and can cost up to $500,000, while 76% of protocols now require at least one amendment.
Regulators are paying attention. FDA's January 2025 draft guidance establishes risk-based credibility requirements for AI in drug development, and ICH E6(R3) requires validated computerized systems and documented review for essential trial documents. The path forward isn't rejecting AI, it's designing oversight that actually catches the subtle errors baseline models produce.