AutoIND: LLM Drafting for IND Summaries
- AutoIND is an LLM-based drafting tool that systematically compiles extensive pharmacology, pharmacokinetics, and toxicology reports into first-draft nonclinical IND summaries.
- The system follows a structured workflow that ingests, classifies, and processes source documents with AWS Textract and customized prompts, before expert review refines the content.
- Efficiency evaluations indicate a time reduction of over 97%, boosting productivity from 0.2 to 12.1 pages per hour while highlighting areas for quality enhancement.
AutoIND is a LLM-based drafting capability within the Weave Platform, built to generate first drafts of nonclinical written summaries for Investigational New Drug applications. It is positioned as a human-in-the-loop acceleration tool rather than a fully autonomous submission engine: AI performs the initial “systematic compilation and synthesis” of Module 4 nonclinical study reports, while expert regulatory writers review, refine, verify, and mature the draft to submission-ready quality (Eser et al., 10 Sep 2025).
1. Scope and regulatory role
AutoIND targets one of the most burdensome steps in early regulatory writing: turning large sets of pharmacology, pharmacokinetics, and toxicology reports into coherent eCTD narrative summaries. In the reported evaluation, the drafting tasks were the nonclinical written summaries in eCTD Modules 2.6.2, 2.6.4, and 2.6.6—respectively pharmacology, pharmacokinetics, and toxicology written summaries. Source materials came from Module 4 and consisted of pharmacology, pharmacokinetics, and toxicology study reports, typically ranging from about 50 pages to several hundred pages per report (Eser et al., 10 Sep 2025).
The system is framed as a response to a schedule-critical bottleneck. The paper argues that IND preparation remains highly manual, expertise-dependent, and especially difficult when late-arriving data must be integrated under compressed timelines. AutoIND is therefore not presented as a replacement for regulatory writers. The paper explicitly frames it as a Regulatory Automation Management Platform-style process in which AI performs initial drafting and expert writers contribute scientific judgment, strategic narrative shaping, regulatory emphasis, and final QC.
The reported use case is specific rather than general. The study evaluated two historical Takeda INDs that had been cleared by FDA within the prior four years, and both were large-molecule INDs. This suggests a demonstrated application in nonclinical IND summary drafting rather than across the full regulatory dossier.
2. Architecture and drafting workflow
In the workflow described by the paper, AutoIND sits after source document availability and before expert regulatory polishing. Its role is to ingest study reports, extract content, classify it into the appropriate IND section, and draft section-specific narrative summaries in Takeda’s style. Human writers then review, refine, verify, and mature the output. The implementation described in the paper identifies the deployed system as AutoIND V2.3 using gpt-4-turbo (release 2024-10), running inside a VPC on AWS; PDFs were extracted using AWS Textract, classified into IND sections, and inserted into prompts customized to Takeda’s style guide (Eser et al., 10 Sep 2025).
The paper describes the platform as processing both structured and unstructured source documents. It does not provide detailed retrieval architecture, chunking strategy, context-window handling, or whether embeddings or RAG were used. It also notes supplementary prompt templates for AutoIND generation, but the prompt text is not reproduced in the main article. This suggests that the reported contribution is primarily system-level and workflow-level rather than a detailed exposition of internal prompt engineering.
The implied drafting pipeline is:
The conclusion further proposes a five-step scalable framework: ingest/extract secure documents and data; draft section-specific content via LLM in minutes; refine/review with SME judgments and systematic error correction; verify by tracing information flow to sources and checking completeness; and publish/monitor eCTD build and AI-drafted responses.
3. Evaluation design and measurement framework
The evaluation used two previously cleared Takeda INDs and a calibration phase before formal testing. For each IND, five representative study reports were used by an experienced regulatory writer and an AI engineer to iteratively refine AutoIND prompts; those calibration reports were excluded from quality assessment to avoid bias. The actual evaluation was therefore conducted on unseen source documents. Time measurement for the AI workflow used Toggl, and the recorded drafting time included PDF upload and content extraction processing time, not only pure text generation latency (Eser et al., 10 Sep 2025).
Manual drafting was not measured side by side. Instead, manual time was estimated from the experience of Takeda regulatory writers with years of experience, using historical manually written INDs as industry-standard comparators. The paper emphasizes that these manual times were contextual benchmarks only, whereas the AutoIND times were actual measured times.
Quality assessment used seven pre-specified categories: correctness, completeness, conciseness, consistency, clarity, redundancy, and prominence/emphasis. Each category had multiple sub-criteria scored on a 0–3 scale, where , , , and . Category scores were normalized as:
and overall quality was normalized across all sub-criteria as:
A critical regulatory error was predefined as any misrepresentation or omission likely to alter regulatory interpretation of safety, efficacy, or compliance; the paper gives concrete examples such as incorrect NOAEL attribution and omission of mandatory GLP dose-formulation analysis. The paper does not describe the assessment as blinded, and it later identifies the use of a single assessor as a limitation.
4. Efficiency outcomes
The quantitative result reported most prominently is drafting-time reduction. For both evaluated INDs, manual drafting time for comparable written-summary work was estimated at approximately 100 hours. Against that contextual benchmark, AutoIND produced complete first drafts in a few hours, including document upload and extraction time, and the figure caption reports statistical significance as , (Eser et al., 10 Sep 2025).
| IND | Source corpus | First-draft time |
|---|---|---|
| IND-1 | 61 reports, 18,870 pages | 3.7 h vs 0 h manual |
| IND-2 | 58 reports, 11,425 pages | 2.6 h vs 1 h manual |
For IND-1, the paper reports a 2 time reduction. For IND-2, it reports a 3 reduction. The abstract summarizes the result as approximately 4 overall reduction. The paper also reports a productivity metric: mean pages per hour increased from 0.2 to 12.1. These gains exceeded the authors’ pre-study hypothesis of 5 drafting-time reduction.
The importance of these numbers is interpretive rather than absolute. Because the comparator was estimated rather than contemporaneously measured, the results establish a large practical acceleration signal rather than a fully controlled head-to-head benchmark. Even so, the reported throughput—handling 61 reports and 18,870 pages, or 58 reports and 11,425 pages, in a few hours—indicates that AutoIND functioned as a large-corpus drafting accelerator rather than as a narrow summarization demo.
5. Quality profile, systematic deficiencies, and error structure
The quality results were mixed but structured. Overall normalized quality scores were 6 for IND-1 and 7 for IND-2. The strongest category was correctness, which was reported as 8 across all modules and both INDs, and no critical regulatory errors were detected in either IND. Clarity and redundancy were also described as relatively strong and consistent compared with other dimensions (Eser et al., 10 Sep 2025).
The weakest category was emphasis. Emphasis scored 9 for IND-1 toxicology and 0 for IND-2 toxicology. The paper interprets this as a systematic failure to foreground the most important findings: drafts often over-described methods and under-weighted actual results and their regulatory implications. Section-level analysis also showed that pharmacokinetics tended to have higher completeness—1 for IND-1 and 2 for IND-2—while toxicology had the poorest completeness at 3 and 4, respectively.
The detailed deficiency profile is unusually specific. Conciseness showed “3–5x word count inflation” relative to expert-written documents. Completeness analysis found that 35% of summaries were missing essential study design elements, and there was a 100% omission rate for GLP-required dose-formulation analysis. Clarity review found 32% structural/organizational issues and 28% AI-characteristic language patterns. The paper also reports that more than 60% of mentions in the prominence/emphasis analysis were method-heavy rather than result-focused. In refinement analysis, 60–70% of cases required structural reorganization for logical flow, and 70–80% required content reduction to reach acceptable conciseness.
The error analysis indicates that high correctness did not imply perfection. Reported noncritical scientific mistakes included claiming dose-response relationships where none existed, confusing parameters such as “latency” versus “duration,” assigning NOAEL values to in vitro studies, and introducing numerical inaccuracies such as “11 h vs 7 h onset” or stating that a parameter “increased” when the source said it “decreased.” Common omissions included missing 5, vehicle composition, dose regimens, toxicokinetic data, control-group results, and quantitative values such as 6 and 7. Redundancy errors included repeated endpoint lists, repeated buffer compositions, and repeated dose and route details. Clarity analysis recorded 156 comments, including structural issues, language/style problems, and technical deficiencies; a representative AI-language artifact was the recurring phrase “The study aimed to evaluate ...”, which appeared 12 times.
Taken together, these findings suggest that AutoIND’s weaknesses were systematic rather than random. The paper repeatedly emphasizes an “AI signature”: overlong prose, repeated boilerplate, structural disorder, incomplete capture of mandatory regulatory elements, and poor narrative emphasis.
6. Governance, limitations, and development trajectory
The system operated under explicit governance constraints. The study states that data remained encrypted in transit and at rest, that AutoIND parameters were not updated using sponsor data, and that the implementation adhered to WHO 2023 ethical AI principles and GAMP 5 guidelines (Eser et al., 10 Sep 2025).
The paper is explicit that AutoIND is not submission-ready on its own. Expert regulatory writers remain necessary not merely for copyediting, but to verify numbers, restore omitted mandatory details, remove repeated and unnecessary procedural text, restructure the document for regulatory logic, standardize terminology, and re-weight the narrative so that critical findings receive appropriate prominence. The paper characterizes the system as handling the “70% cognitive lift,” while reserving final scientific and regulatory judgment for human experts.
Several limitations constrain generalization. The study covered only two INDs at a single company, in one modality and therapeutic context. Manual benchmarks were estimated rather than directly measured. Quality assessment relied on a single assessor. The AI drafts were not submitted to FDA, so the study does not connect rubric scores to downstream regulatory outcomes such as review comments, information requests, or filing success. The evaluation also covered only nonclinical written summaries, not tabulated summaries, clinical sections, or CMC content.
The future roadmap is correspondingly concrete. The authors recommend domain-specific fine-tuning, explicit verbosity penalties, structured templates with mandatory field validation, domain-specific evaluation metrics, automated AI-signature and deficiency detection, and more standardized source reports. They also identify broader applicability beyond IND nonclinical summaries to marketing applications, regulatory responses, safety reports, and protocol amendments. A plausible implication is that the paper treats AutoIND less as a finished authoring system than as an early regulatory drafting substrate whose value lies in speed, source-grounded first-pass synthesis, and a controlled division of labor between LLM generation and expert review.