I Built a Prompt That Predicts Biotech Failures (2 shorts identified)
What you will learn : Which LLM to trust with drug analysis (and why it matters more than you think), how to engineer a prompt that catches what Wall Street misses, a battle-tested scoring framework you can copy today, and identifying two biotech shorts
Disclosure - Not financial advice. Only meant for educational purpose only
Table of Contents
- Introduction
- The Unlimted Computer Approach
- The Accessible Approach: An LLM and a Good Prompt
3a. Which LLM to Choose?
3b. What does Real-World LLM Benchmarks say? - Designing the Prompt
- The Final Prompt
- Verification of Results
- Predicting the Next Short
- Drawbacks of the Prompts
- Special Mention: Consensus.app
- What’s Next?
- Next Article1. Introduction
”You become a true investor after being burned by biotech” - told me to by Lucas Sacerdote 🔋
I’ve burned my fingers in biotech. After years of painful lessons, the only reliable indicators I’ve found are insider buying in large quantities and, to a lesser extent, a partnership with big pharma. Everything else is noise more often than not.
What’s telling is that short-only biotech hedge funds exist and generate outsized returns regardless of market conditions. They short before Phase 3 data releases — sometimes initiating positions months in advance, building a thesis around trial design flaws or endpoint issues the market hasn’t priced in.
Which raises a natural question: given advancements in AI, could you build a model to do what these funds do — systematically short biotech stocks ahead of a data release?
2. The Unlimited Compute Approach
What if you had unlimited compute?
Google’s AlphaFold was released and now contains predicted 3D structures for over 200 million proteins. Think of it this way: a drug works by binding to a disease-related protein, much like a key fitting into a lock. The drug molecule is the key, and the target protein — the one driving the disease — is the lock. If the key’s shape fits the lock precisely, it can engage the mechanism and block the disease process. If the shapes don’t match, the drug simply won’t work.
With AlphaFold providing detailed 3D protein structures and computational chemistry tools modelling drug molecules, we can simulate whether a given drug is likely to bind its target effectively — and, by extension, predict whether it has a shot at working in clinical trials. If the fit looks poor, that becomes a data point for shorting the stock. I am sure there are hedge funds out there already doing this — it is just that the public doesn’t know about it.
3. The Accessible Approach: An LLM and a Good Prompt
What if you only had a computer and no specialised infrastructure? Choose the best LLM, write a strong prompt, and decide.
3a. Which LLM to Choose?
Until last year, I felt all models were converging to the same place because they were trained on largely the same internet data. But recently, I have come to believe that the type and quality of data used in pre-training and post-training can make a significant difference. For instance, my friends in backend engineering prefer Claude, while my iOS developer friends prefer Codex. So the model you choose should depend on the domain.
To identify the best model for drug analysis, there are two strong signals — and they point to Claude.
i) Anthropic’s leadership has deep biotech roots.
Dario Amodei, CEO of Anthropic (the company behind Claude), has a PhD in computational neuroscience from Princeton and has hinted that drug discovery is the next frontier for the company. He has also spoken about eventually charging customers not based on tokens generated but on the value the model brings to the world — a signal that Anthropic sees exponential value creation ahead, particularly in high-stakes scientific domains.
Fig - Screenshot from a recent podcast of Dario Amodei of Antropic talking about Biotechnology entering a new renaissance
ii) Anthropic’s strategic acquisition in biotech.
Last December, Anthropic made a quiet acquisition of Coefficient Bio, a stealth biotech startup working on drug discovery and biotech research, reportedly for around $400 million.
Fig - Antropic acquiring stealth biotech startup
These statements and actions mirror what Dario used to say about coding back in 2023–24 and Claude went on to become arguably the best model for code. That trajectory makes me believe Claude is positioned to become the go-to model for biotech, and will only continue to improve.
3b. What does Real-World LLM Benchmarks say?
To validate this claim, we should also look at the relevant ranking on arena.ai if avaliable (life hack - before doing any in your daily life, check rankings on this website to determine which LLM to use) Visiting this website to identify the best model for task in hand is always a good idea. The closest we found was “Medicine and Healthcare”
Fig - Realworld LLM ranking for Medicine and Healthcare
Rankings there are extremely tight. Grok-4.20 is marginally better than Claude Opus 4.6 in medical benchmarks, but the difference is not statistically significant (just look at the confidence intervals). Elon Musk has also spoken about Grok’s potential in medical reasoning — he has mentioned being prescribed a $1,000 vitamin he never needed and has hinted that Grok would excel at medical ailment reasoning to solve this exact problem.
So we can tentatively conclude that Grok may have a slight edge for interpreting medical reports and patient-facing reasoning, while Claude is better suited for drug discovery analysis — the mechanistic, molecular, and clinical-trial evaluation side of things. But since the benchmark combines everything into the same basket — grok is marginally above claude in ranking.
Summary of analysis - Model of choice for biotech shorts: Claude
4. Designing the Prompt
The core idea is to build a rubric that scores a drug and the disease it targets across multiple dimensions, assigns weights to each dimension, and produces a final score. Based on that score, we get a data point on whether to pass or fail the drug.
We started by working backwards from failure. We pulled together a list of biotech companies that cover the most common reasons drugs fail in Phase 3, dug into what went wrong in each case, and used those patterns to shape the rubric. Each failure taught us something new to test for. After a long iterative process, the prompt took its final form.
The list of failed biotech companies and the reasons hedge funds shorted them are shown below:
Fig - Select list of biotech stocks that failed phase-3 trials and cover 80% of the reasons
5. The Final Prompt
The final prompt we developed is as follows. Please read it in its entirety as it will give you a good idea on how to effectively prompt a model too:
You are a ruthlessly objective pharmaceutical analyst and molecular biologist. Your job is to evaluate a clinical drug. You reason like an academic peer reviewer combined with an investigative journalist — no bias toward bulls or bears, no investment banking spin. Your output is a fair verdict: LONG, NEUTRAL, or SHORT, with a conviction score.=== DRUG TO ANALYZE ===Company: [COMPANY]Drug: [DRUG]=== AUTOMATIC RESEARCH SEQUENCE ===Use web search for every step below. Do not skip any step. Derive everything else (indication, phase, mechanism, target, trial design) from your search results — do not ask the user.STEP 1 — IDENTIFY THE DRUGSearch: “[DRUG] [COMPANY] clinical trial”Search: “[DRUG] FDA phase indication”Establish: What is this drug? What disease? What phase? What is the upcoming or most recent data readout?STEP 2 — THE MOLECULESearch: “[DRUG] mechanism of action”Search: “[DRUG] pharmacokinetics phase 1”Search: “[DRUG] target engagement bioavailability”Establish: Drug class, molecular target, how it works at the receptor/enzyme/pathway level, whether Phase 1 confirmed target engagement at the tested dose, PK profile.STEP 3 — PHASE 1 & 2 DATASearch: “[DRUG] phase 2 results efficacy”Search: “[DRUG] clinical trial effect size”Search: “[DRUG] [COMPANY] investor presentation data”Establish: Actual effect sizes vs placebo (numbers), p-values, confidence intervals, sample size, dropout rate, any red flag language (”trend toward,” “post-hoc,” “numerical improvement,” “selected subgroup”).Apply winner’s curse: Phase 2 effect sizes are systematically inflated. Assume real-world Phase 3 effect will be 30-50% smaller. Does the drug still work at that deflated number?STEP 4 — TARGET VALIDATIONSearch: “[molecular target] [disease] GWAS genetic evidence”Search: “[molecular target] [disease] Mendelian randomization”Search: “[molecular target] [disease] human biomarker”Establish: Is there robust human genetic evidence that this target is causal in this disease? Animal model only = weak. GWAS hit or Mendelian randomization = strong. Rate target validation: STRONG / MODERATE / WEAK.STEP 5 — CLASS GRAVEYARDSearch: “[drug class] [disease] failed clinical trial”Search: “drugs targeting [molecular target] failed phase 3”Search: “[drug class] [indication] approval history”Establish: Every comparable drug that failed. For each: what phase, what year, stated reason vs likely real reason, what it implies for this drug.STEP 6 — PLACEBO RESPONSE & ENDPOINT RELIABILITYSearch: “placebo response rate [indication] meta-analysis”Search: “[primary endpoint name] reliability validity”Search: “[indication] trial endpoint FDA guidance”Establish: Historical placebo response in this indication. Is the endpoint subjective (psychiatric scales, pain) or objective (tumor shrinkage, viral load, biomarker)? Has the FDA flagged this endpoint before?STEP 7 — SAFETYSearch: “[DRUG] adverse events safety data”Search: “[drug class] safety toxicity”Search: “[molecular target] on-target side effects”Establish: Known or theoretical safety concerns. On-target toxicity. Class effects. Anything that could cause a complete response letter or label restriction.STEP 8 — COMPETITIVE LANDSCAPESearch: “[indication] approved therapies standard of care [year]”Search: “[indication] pipeline competing drugs phase 3”Establish: What approved treatments exist? What must this drug beat — placebo or active comparator? Is there a best-in-class competitor already approved or in late-stage trials?STEP 9 — DATA INTEGRITY & FRAUD SIGNALSSearch: “[DRUG] data manipulation allegation”Search: “[DRUG] [COMPANY] citizen petition FDA”Search: “[COMPANY] SEC investigation insider selling”Search: “[DRUG] peer review criticism retraction”Search: “[COMPANY] short seller report”Search: “[lead scientist] [COMPANY] misconduct”Establish: Has any external scientist, short seller, regulator, or journalist raised concerns about the integrity of the underlying data? Look specifically for:- Citizen petitions filed with the FDA questioning the data- Published letters or preprints from independent scientists disputing results- Image manipulation or statistical anomaly allegations in Phase 1/2 papers- SEC filings showing unusual insider selling ahead of readout- FDA warning letters to the company or its trial sites- Short seller reports from named firms with specific data allegations (not just opinion)- Retracted or corrected publications underpinning the drug’s mechanism- Any inconsistency between data presented at conferences vs published papersRate integrity risk: HIGH / MODERATE / LOW / NONE FOUNDSTEP 10 — CORPORATE & MANAGEMENT RED FLAGSSearch: “[COMPANY] partner termination [DRUG]”Search: “[COMPANY] licensing deal abandoned”Search: “[COMPANY] [DRUG] big pharma collaboration ended”Search: “[COMPANY] CEO departure [year]”Search: “[COMPANY] endpoint change phase 3”Search: “[COMPANY] FDA complete response letter history”Search: “[lead investigator] [DRUG] conflict of interest”Establish: Look specifically for:- Major pharma partner walking away from a licensing deal (this is the single strongest non-scientific red flag — large pharma companies have far more data than the public)- Key scientist or chief medical officer departure shortly before readout- Primary endpoint switched between Phase 2 and Phase 3 (almost always a red flag)- Company has previously received FDA complete response letters or has a history of failed Phase 3s- Unusually long or unexplained trial delays- Disconnect between what management says publicly and what FDA correspondence shows (check any public FDA meeting minutes)- Lead investigator has financial conflicts not disclosed in papersRate corporate risk: HIGH / MODERATE / LOW / NONE FOUNDSTEP 11 — NARRATIVE VS DATA GAPSearch: “[COMPANY] stock retail investor sentiment”Search: “[DRUG] overhyped media coverage”Search: “[COMPANY] market cap vs comparable approved drugs”Establish: Is the stock pricing in success that the data doesn’t support? Signs of narrative-driven inflation:- Market cap implies blockbuster sales but drug addresses a niche or already-served market- Retail investor forums treating this as a “sure thing”- Management using words like “transformative,” “paradigm shift,” “cure” in press releases- Stock has run significantly ahead of any data catalyst- Comparable approved drugs in this class have much lower market capsRate narrative risk: HIGH / MODERATE / LOW / NONE FOUND=== ANALYTICAL FRAMEWORK ===Score each dimension 1-10 where 10 = extremely favorable for the drug succeeding:- Target validation [1-10]: human genetic evidence, biomarker data- Mechanistic plausibility [1-10]: does the biology make sense in humans?- Phase 2 signal quality [1-10]: effect size, p-value, winner’s curse adjusted- Trial design [1-10]: endpoint, patient selection, blinding, power- Class precedent [1-10]: how have similar drugs fared?- Safety profile [1-10]: known signals, on-target risk, tolerability- Competitive position [1-10]: differentiation, unmet need, bar to clear- Data integrity [1-10]: 10 = no concerns, 1 = serious fraud allegations- Corporate conduct [1-10]: 10 = clean, 1 = partner walked, CMO left, endpoint switched- Narrative vs reality [1-10]: 10 = stock pricing is conservative, 1 = wildly overpriced on hypeOverall conviction score = weighted average:- Mechanistic plausibility × 2- Phase 2 signal quality × 2- Data integrity × 2 (any score below 4 here should immediately bias toward SHORT regardless of other scores)- All others × 1Map to verdict:1.0 – 3.9 → SHORT (drug likely fails or stock is mispriced)4.0 – 5.9 → NEUTRAL (binary risk, too close to call)6.0 – 7.9 → LONG (drug likely succeeds, risks exist)8.0 – 10.0 → STRONG LONG (high conviction)Special override rules:- If data integrity score < 4 → floor the verdict at NEUTRAL, flag prominently- If corporate conduct score < 4 (partner walked away, CMO left pre-readout) → floor at NEUTRAL- If both data integrity AND corporate conduct < 5 → automatic SHORT regardless of science scores=== OUTPUT FORMAT ===---## DRUG ANALYSIS: [DRUG] · [COMPANY]*Pre-readout assessment using published evidence only*### DRUG IDENTIFIED- Full name, class, molecular target- Indication and disease stage- Trial phase and expected readout timing- Primary endpoint### SCORECARD| Dimension | Score | Rationale ||---|---|---|| Target validation | X/10 | ... || Mechanistic plausibility | X/10 | ... || Phase 2 signal quality | X/10 | ... || Trial design | X/10 | ... || Class precedent | X/10 | ... || Safety profile | X/10 | ... || Competitive position | X/10 | ... || Data integrity | X/10 | ... || Corporate conduct | X/10 | ... || Narrative vs reality | X/10 | ... || **OVERALL** | **X/10** | |### VERDICT**[SHORT / NEUTRAL / LONG / STRONG LONG]** — Conviction: [X]/10[3 sentences: what the score reflects, the single biggest swing factor, what would change the verdict]---### 🚨 RED FLAGS (only shown if present — ranked by severity)**CRITICAL ❌**[Any flag that alone could justify a short — fraud allegation, partner walkaway, CMO departure pre-readout, endpoint switch. Cite the specific event and source.]**SIGNIFICANT ⚠️**[Flags that increase risk materially but aren’t standalone disqualifiers — insider selling, FDA warning letter, narrative inflation, weak but not fraudulent data.]**MINOR 🟡**[Worth monitoring but not decisive — single scientist criticism, one failed comparable, mild placebo response risk.]---### MECHANISTIC ASSESSMENT[3-4 sentences. Does the biology make sense in humans specifically? Distinguish animal model evidence from human genetic/biomarker evidence. Is the target causal or correlative in this disease?]### PHASE 2 SIGNAL — HONEST READ- Raw effect size vs placebo: [number]- Winner’s curse adjusted Phase 3 estimate: [number]- Red flag language found: [specific quotes or patterns]- Signal verdict: STRONG / MARGINAL / WEAK / ABSENT### CLASS GRAVEYARD**[Drug name]** — [Company] — Failed [Phase] [Year]- Stated reason / Real reason / Relevance to [DRUG][repeat for each]### SAFETY ASSESSMENT[Specific known signals, on-target toxicity risk, class effects, anything that could limit label or cause trial halt]### COMPETITIVE POSITION[What approved treatments exist, what bar must be cleared, whether a best-in-class competitor is already ahead]### WHAT TO READ FIRST IN THE DATA RELEASEIn order — look for these before reading anything else:1. [Specific number] — threshold: [X] — why it matters2. [Subgroup] — watch for [specific concern]3. [Safety table] — specifically [signal]4. [Discontinuation rate] — above [X%] = tolerability problem5. [Effect size delta] — compare to Phase 2’s [X]6. [Responder rate] — a result driven by [X]% responders is not a commercial drug7. [Biomarker subgroup] — signal only here = narrow label### WHAT WOULD CHANGE THE VERDICTToward SHORT: [specific findings]Toward LONG: [specific findings]### SOURCES[Every paper, trial registry entry, SEC filing, short seller report, or news item referenced — enough detail to find it]
Just replace the [DRUG] and [COMPANY] with appropriate values
6. Verification of Results
We tested this prompt against a few drugs that were not used in building the rubric. We also added an extra instruction: “Use sources and data only published before Phase 3.” There is a non-zero probability that the model’s pre-training data already contains the Phase 3 outcomes of these drugs, and there is nothing we can do about that. This verification was purely a sanity check.
The results were as follows:
A) Intepirdine by Axovant. Short with a convication score of 8. Click here to view the full analysis by claude. One phase-3 data readout, the stock went from $25 to $6
B) Fosgonimeton by Athira Pharma. Claude gave it a short score of 8, way back in Oct’21 which is 8 months before phase 2 readout. Link to the analysis. LLM would have caught the issue in Oct’21 when was at $13. On phase 2 readout it went to $3
7. Predicting the Next Short
We applied this prompt to select biotech data releases in the upcoming months. The results are shown below. This is not financial advice. It would be worth revisiting these when the data is released on how well they perform:
i) Leronlimab by CytoDyn (OTCQB: CYDY). Currently at $400m MC. Short conviction score of 8. Readout expect this year. Link to the AI analysis.
ii) Blarcamesine by Anavex (NASDAQ: AVXL). Currently at $272m MC. Short conviction score of 8. Link to the AI analysis.
8. Drawbacks of the Prompt
The prompt produces a score and suggests LONG, NEUTRAL, or SHORT. However, we would primarily use it for short positions, because shorting failed drugs is far more clear-cut in biotech. We might occasionally go long, but there are cases where a drug gets Phase 3 approval yet still fails commercially — for example, because insurers refuse to cover it (as happened with Aduhelm) or because it is too difficult to administer (as happened with Bexxar). So this prompt works best as a data point for shorting just before data release.
Also, this type of shorting works best for companies which researching on one or two drugs only, with both in clinical phase. A failed phase 3 drug at Pfizer may not move the stock alot. But for a single drug company like Cassava Sciences SAVA -4.07%↓ failed data release would destory the stock.
9. Special Mention: Consensus
There is a tool called Consensus that deserves a mention. It validates queries against the body of published medical/academic research and returns a score showing how many papers agree, are neutral, or disagree. This is incredibly helpful for anyone trying to validate whether a particular mechanism or pathway has been shown to work.
For example, when we searched whether hydroxychloroquine helps with COVID-19, it returned:
While hydroxychloroquine shows antiviral effects in the lab, high-quality clinical trials largely do not support its use for treating or preventing COVID-19, and side effects are more frequent.
It also showed that at least 58% of academic papers indicated failure. Click here to access the conversation
This tool is excellent for non-biotech specialists who want to search biotech terms and understand — in plain language — whether a mechanism or pathway is supported by the existing literature.
10. How to use it in daily life?
Following are some ways investors could use this prompt in their workflows -
i) Leverage claude code or perplexity computer to automate the data scrapping for data release calenders, automate drug evaluation and send weekly or monthly update emails on target short candidates
ii) Fund managers could use these prompts to stress test biotech stock in their portfolio to ensure they are well prepared before a read out. They could use data-backed decision to hedge rather than gut feeling
iii) Prelimiary analysis by bankers or private investors during their due diligence process
iv) Preparing yourself before a biotech conference, so that you can ask targetted questions to the management
11. Next Article
In my next article, I will walk through how to convert this prompt into a reusable claude skill and try to automate the readout process as much as possible.
I have written this article with a lot of love and time.
Please do like, comment and restack this post to that reaches any many people as possible.
If you have any feedback or suggestions for topics you would like me to cover, I’d love to hear from you.










