Drug mechanisms with genetic support have a 2 to 3-fold higher probability of clinical success in clinical development compared to those without. We built an automated pipeline capable of transforming basic asset details into a structured genetic insights report in minutes, allowing our BD team to generate genetic insights for an entire quarterly pipeline in the time it once took to evaluate a handful of candidates.
Jul 8, 2026 · 4:59 PM UTC
2
3
385
The pipeline consists of seven stages – some are deterministic, while others use LLMs to handle ambiguity or synthesize findings. All LLM-generated outputs require human review before the pipeline proceeds, ensuring scientific judgment determines the final deliverable.
1
153
Consider “Asset X”, an anti-IL-18 antibody in development for Crohn's disease. Since IL-18 is a potent pro-inflammatory cytokine, we can hypothesize that blocking its activity could be efficacious in a chronic inflammatory condition like Crohn’s. To test this, we investigate individuals carrying genetic variants that reduce IL-18 function. In effect, their genetics mimic what the drug does. We then assess whether they are:
• Protected from Crohn's and related diseases
• Protected from other diseases
• At increased risk for any adverse conditions
1
93
Step 1: Extracting the MoA. First, the pipeline receives a string describing the asset’s MoA: "anti-IL-18 antibody." Using LLM-based parsing, we extract the following:
• Target gene: IL-18
• Modality: Antibody
• Modulation direction: Negative (since the drug is an antagonist antibody)
• Confidence score: 100 (indicating high conviction in the extraction)
Here, the input is well-structured, so the extraction presents no issues. However, not every input arrives so cleanly formatted. When the pipeline encounters ambiguity, it flags uncertainty for human review.
1
27
Step 2: Validating the MoA (Human-in-the-Loop). Since Step 1 relies on LLM parsing, we introduce a human checkpoint to verify the extraction. A scientist reviews the output to confirm the mapping is correct and that subsequent analysis should focus on variants that mimic the drug’s effect. This checkpoint catches any errors before they affect downstream steps.
1
35
Step 3: Querying genetic databases. With the target gene confirmed, the pipeline focuses on individuals whose genetics mimic the drug’s effect. To identify relevant outcomes in these individuals, we scan thousands of phenotypes for statistical associations with the target gene. The pipeline queries three complementary biobanks: UK Biobank, FinnGen, and All of Us.
The initial output contains thousands of associations which we filter aggressively, requiring both statistical significance and meaningful effect size. What remains are the signals most likely to reflect real biology.
1
28
Step 4: Generating the narrative. The filtering yields associations that still require interpretation, but expert review of each association doesn’t scale. An LLM synthesizes the data into a structured narrative, grouping findings by therapeutic area and distinguishing between signals that suggest new opportunities, warrant caution, or validate the existing strategy.
1
20
Step 5: Reviewing the narrative (Human-in-the-Loop). Since Step 4 relies on LLM synthesis, which can hallucinate claims or misgroup findings, we introduce a second human checkpoint. A scientist reviews the output, which includes asset context, extracted associations, and a full narrative with claims and therapeutic area groupings. Here's a condensed view of what that looks like:
1
31
The reviewer confirms that:
• The groupings are logical (e.g., “Renal failure” and “Chronic kidney disease” belong together)
• Overlapping categories are merged rather than duplicated
• Claims are grounded in the filtered data rather than hallucinated. If the narrative mischaracterized a signal (for instance, overstating a weak association or miscategorizing a phenotype), the scientist would flag it here.
This checkpoint ensures the final output reflects what the genetic evidence actually supports.
1
24
Step 6: Delivering the report. The approved narrative is packaged into two deliverables: a presentation-ready HTML report summarizing the genetic evidence, and an Excel file with the underlying data for verification.
1
18
The report for Asset X reveals two themes:
The first is encouraging: individuals carrying variants that mimic the drug’s effect show lower risk of several disease categories, including kidney and renal diseases (renal failure: β = -2.03; chronic kidney disease: β = -1.62) and degenerative joint conditions (arthrosis: β = -1.09 to -1.74). While no direct Crohn's disease signal was observed, the asset-specific analysis suggests IL-18 antagonism may increase liver disease risk (β = 2.41 to 3.85) — relevant given IBD-liver comorbidity.
The second is cautionary: the same variants are associated with increased risk for conditions including kidney stones (calculus of kidney: β = 3.54 to 3.70), liver disease (chronic hepatitis: β = 3.83 to 3.85), and certain malignancies. These associations flag potential adverse effects to monitor in clinical development.
1
97
Producing this output manually would take hours; the pipeline generates it in minutes. Human input totals roughly ten minutes across two checkpoints, preserving scientific judgment for where it matters most. Full blog on how we built the pipeline and how we’re continuing to develop it: formation.bio/blog/scaling-t…
87





