The policy simulation engine, Forecasts that get graded. Powered by @Agentese_ai

Yarrow retweeted
🏆 Your BUIDL_QUESTS 2026 Top 10 are officially in. Huge thanks to the 18 judges from @fenbushi, @IOSGVC, @HivemindCap, @vertexventures, @decasonic, @archetypevc, @1kxnetwork, @TheSpartanGroup, @marine_vc, @primitivecrypto, @amber_ac_ , and @ambergroup_io for taking the time to review the teams during our online pitch round and help shape this year's Top 10. Now, the final 10 move forward. 🔹Afu Pocket AI for operators’ decisions, built to help business owners think through day-to-day choices faster. 🔹Alano @leonchen3a Sports understanding + sports-data assets, turning game footage into structured, usable intelligence. 🔹Aria Studio @AriaStudioHQ A one-person media company run by AI Agents, compressing an entire creative workflow into a much smaller team. 🔹Astrail @astrailxyz Turns scattered travel inspiration into routes you can actually take, connecting discovery with planning. 🔹FlowIn AI autocomplete and reply drafts across every Mac app, designed to make everyday communication faster and less fragmented. 🔹Mercatus AI @mercatus_ai An exchange where AI inference tokens are priced, creating a market around the cost and access of compute. 🔹Mercha @MerchaNET A merch Agent for teams, communities and events, handling more of the process from idea to execution. 🔹Pawa Robotics @pawa_sci Wearable pet-health intelligence, combining hardware and data to better understand animal health. 🔹Whale Dance @whaledotdance Trading infrastructure for AI Agents, built for autonomous systems participating in financial markets. 🔹Yarrow @Yarrow_ai A policy simulation and strategic foresight platform using Agents to model scenarios and support better decision-making. 🔒 Next up: an internal Pitch Day, where the teams will make their final case before we head to Singapore. 👉Coming soon: 📍 BUIDL_QUESTS 2026 Final Day 📅 October 6 🏛️ CHIJMES Hall, Singapore 🎤 Final Day will bring together the Top 10, investors, founders and ecosystem partners for a full afternoon of keynotes, panels, final demos and awards. 👀 The Final Day VC panel will be announced soon on @OpenArena_To. 🎟️RSVP for Final Day: luma.com/y9g4tduu?tk=0AcdzN
2
3
7
982
Say you want to know: will the Fed raise rates next month? If you're a regular person, you Google it. You find three headlines — one says yes, one says no, one says maybe. You pick the one that sounds most convincing, or the one from the source you trust most. That's not a forecast. That's a vibe check. If you ask ChatGPT, it gives you a confident paragraph. "Based on current economic indicators, it appears likely that..." It sounds smart. But it has no track record, no accuracy score, and if you ask the same question tomorrow with slightly different wording, you might get a different answer. It doesn't know how often it's been right before — because nobody's counting. If you check Polymarket, you get a number — say 55%. That's better. Real money is behind it. But it's just a price. There's no explanation for why it's 55% and not 40%. No evidence trail. No reasoning you can walk through. And if the number changes overnight, you have no idea what drove it. ☘️ Here's what happens inside Yarrow. The same question goes to many independent analysts. They're running on three different AI models from two different providers. None of them can see the market price. None of them can see each other's work. Each one gets the same frozen evidence pack — official data, dated sources, verified statistics. No live search. No Twitter. No vibes. They each have to produce a complete chain: what mechanism drives the outcome → what the data says → what probability they assign. Before any of that ships, the evidence pack runs through five checks: are the numbers internally consistent? Is anything missing that should be there? Is there any information that could only be known after the fact? Has an independent red team tried to break the argument? The final number is the median of the valid analysts — not a vote, not a discussion, not a compromise. Just the middle of six independent judgments. Then it gets frozen. Locked in git. Published. And when the outcome is known, scored — right next to every other prediction we've ever made. Google gives you headlines. ChatGPT gives you confidence. Polymarket gives you a price. Yarrow gives you a probability with the evidence behind it, the reasoning you can audit, and a track record you can check. That's the difference between looking for an answer and building one.
1
4
83
On June 23, 2016, the UK held a referendum: remain in the EU, or leave. Betting markets gave Remain an 85–90% probability. Pollsters predicted Remain. The UK government had no contingency plan for Brexit, and the prime minister had effectively staked his political career on an outcome he himself could not predict with certainty. Brexit won. 51.9% to 48.1%. The pound plunged 10% overnight, its largest single-day drop since 1985. David Cameron resigned the following morning. Global markets lost $2 trillion in value within two days. And Britain spent the next four years trying to figure out what “Brexit” actually meant. The data telling a different story was already there. Online polls had consistently been closer to the truth than telephone polls. The turnout models had never been properly validated against a nationwide referendum. And anti-immigration sentiment in northern England had been systematically underestimated. Betting odds were anchored to the same telephone polls that were later shown to be wrong. No one stress-tested the consensus. An 85% probability became a “fact,” and every plan was built around it. ☘️ Now imagine if the UK government had used a system like Yarrow before the vote. The system receives one question: “Will the UK vote to leave the EU?” Yarrow has independent analysts approach the question from different bodies of evidence: polling data using two different methodologies, regional turnout models, historical precedents from referendums versus elections, and sentiment data from regions that conventional polling struggled to reach. Three analysts flag online polls showing Leave ahead. Two build turnout scenarios in which older, non-urban voters who leaned toward Leave turn out at higher rates than assumed by the polling models. One points out that betting markets are pricing the same telephone polls—not independent information. The system ultimately returns: Remain: 58%, Leave: 42% And it adds: The 85% consensus probability for Remain depends heavily on a telephone-polling methodology that had never been validated in a nationwide referendum. Cameron might still have held the referendum.But the Treasury might have had contingency plans ready. The Bank of England could have positioned itself in advance. And perhaps $2 trillion would not have vanished from global markets within 48 hours simply because everyone was blindsided by an outcome that, if someone had looked at the full evidence base, was clearly possible. The most expensive predictions aren't the ones that turn out to be wrong. They're the ones that are wrong—and that nobody prepared for.
1
4
101
What should a prediction actually look like? Most AI tools give you a number. "There's a 70% chance this happens." Maybe a paragraph of explanation. That's it. No source trail, no what-if analysis, no way to check later whether 70% was the right call. 👀 Here's what we think a prediction should contain — at minimum — before anyone should take it seriously: 1️⃣ a probability that means what it says. When the system outputs 70%, events like that should happen about 70% of the time. Not "I'm pretty sure" dressed up as a number. Actual calibration, measured over hundreds of questions. 2️⃣ a reasoning chain you can walk through. Not "based on our analysis" — the actual evidence, the actual logic, step by step. Which data points support the conclusion? Which ones cut against it? If you can't trace from evidence to number, it's not a forecast — it's a vibe. 3️⃣ the conditions under which the prediction breaks. Every good forecast comes with its own kill switch. "We'd revise if X happens." "This falls apart if Y turns out to be true." If a prediction doesn't tell you what would prove it wrong, it's not brave — it's unfalsifiable. 4️⃣ a frozen record before the outcome. Anyone can claim they predicted something after the fact. The test is whether the prediction was locked in — with all its reasoning — before the answer was known. No edits, no revisions, no "well what I actually meant was." 5️⃣ a public track record. Not just the wins. The misses too. Side by side. Scored by the same rules. Because a system that only shows you its highlights isn't building trust — it's building a marketing deck. ☘️ This is the standard Yarrow holds itself to. Every prediction frozen in git. Every outcome scored. Every miss published alongside every hit. Not because we're always right. But because a prediction you can't audit is just an opinion with better formatting.
1
4
115
We actually tested how Jev performs in a prediction workflow. Yarrow used 48 real-world cases across 31 independent documents. Total model cost: less than $0.10. Yes, we have a cost advantage too. ☘️ Here’s what we found — what worked, what didn’t, and what remains inconclusive. Where Jev did well: Jev can read complex corporate announcements and turn them into structured, verifiable event records. 48/48 cases — it correctly identified the current status. It also caught distinctions that simple keyword matching would miss. For example, one announcement was titled “Withdrawal of Bank Regulatory Application”, while the body also said the company “reaffirmed its full-year guidance.” A keyword-based system could mistakenly interpret this as a withdrawal of the guidance. Jev correctly distinguished the two. Where it failed: Understanding the current status is not the same as determining whether a contract has already resolved. In 10 of 48 cases, Jev correctly recognized that the available material did not mention the outcome — and then incorrectly concluded that the outcome was “No.” For anyone trying to use it in production, this is a critical issue. Overall judgment accuracy:Earnings guidance: 22/25 HBM supply: 16/23 Neither cleared our 90% threshold. Stability is still an issue. We ran the exact same inputs repeatedly. In 8 out of 48 cases, at least one field changed. We also changed only the order of the input paragraphs. In 20 out of 48 cases, the result changed. That means it is not yet suitable for automated resolution or fully unattended workflows. Can Jev predict? We’ve already frozen 12 forward-looking questions — covering Fed meetings, weather dates, and geopolitical events — along with 6 corporate disclosure predictions. The outcomes are still pending. But there is already one signal worth investigating: For companies with regular quarterly disclosure patterns, Jev assigned 93–99% probability to the statement “the disclosure has not been released.” We’re watching this closely. There’s another interesting finding: When we asked about the same event in two different formats — binary Yes/No versus multiple-choice — the probabilities shifted by an average of 12–13 percentage points. The probability depends on how you ask the question. Our interim conclusion Jev is genuinely useful for one thing: Turning messy corporate announcements into structured, checkable records. It’s fast. It’s cheap. And it reads surprisingly well. But reading accurately is not the same as judging accurately. Understanding is not prediction. We’ll keep experimenting and keep testing. And if you’re currently researching Jev — especially its applications in prediction — we’d be happy to compare notes and discuss the results together. 👀 ☘️
Everyone is talking about Jev this week. From the perspective of a prediction system, here’s how we see it. Jev is not another chatbot. It’s a “System 1” model — essentially a classification engine that returns calibrated probabilities rather than generating text. We noticed something interesting: Jev achieves accuracy comparable to GPT-5.6 Terra, while being 25× faster and 75% cheaper. Zero hallucinations — not because the model is smarter, but because its output space is predefined. It physically cannot generate content it wasn’t asked to produce. We’ve also seen some impressive demos recently: searching for flights in 7 seconds for $0.004, running 50 simultaneous games of Subway Surfers for less than a cent, and rebuilding a Tesla FSD prototype in an hour. Jev’s core training method is called RLCD — reinforcement learning for calibrated decision-making. Its optimization target is simple: when the model says 70%, the event should actually happen roughly 70% of the time. That’s calibration — and it’s exactly the same core objective Yarrow is built around. ☘️ The difference is scope. 🤖 Jev answers: “Which button should I click right now?” — in milliseconds. ☘️ Yarrow answers: “If this policy passes, how will the corresponding market price change?” Jev is System 1 — fast, intuitive, cheap. Yarrow is System 2 — slower, more deliberate, evidence-dense. Here’s the interesting possibility: What happens if we combine Jev and Yarrow? 👀 Use Jev for thousands of micro-decisions throughout the prediction pipeline — scoring evidence relevance, classifying sources, verifying claims — cutting costs by 100×, while keeping the full reasoning engine for the final probability judgment and scenario distribution. Small decisions get fast judgment. Big decisions get deep simulation. Jev and Yarrow are different — but as a technology stack, they could be highly complementary. Stay tuned for Yarrow’s upcoming research. 👀 ☘️
1
6
860
Everyone is talking about Jev this week. From the perspective of a prediction system, here’s how we see it. Jev is not another chatbot. It’s a “System 1” model — essentially a classification engine that returns calibrated probabilities rather than generating text. We noticed something interesting: Jev achieves accuracy comparable to GPT-5.6 Terra, while being 25× faster and 75% cheaper. Zero hallucinations — not because the model is smarter, but because its output space is predefined. It physically cannot generate content it wasn’t asked to produce. We’ve also seen some impressive demos recently: searching for flights in 7 seconds for $0.004, running 50 simultaneous games of Subway Surfers for less than a cent, and rebuilding a Tesla FSD prototype in an hour. Jev’s core training method is called RLCD — reinforcement learning for calibrated decision-making. Its optimization target is simple: when the model says 70%, the event should actually happen roughly 70% of the time. That’s calibration — and it’s exactly the same core objective Yarrow is built around. ☘️ The difference is scope. 🤖 Jev answers: “Which button should I click right now?” — in milliseconds. ☘️ Yarrow answers: “If this policy passes, how will the corresponding market price change?” Jev is System 1 — fast, intuitive, cheap. Yarrow is System 2 — slower, more deliberate, evidence-dense. Here’s the interesting possibility: What happens if we combine Jev and Yarrow? 👀 Use Jev for thousands of micro-decisions throughout the prediction pipeline — scoring evidence relevance, classifying sources, verifying claims — cutting costs by 100×, while keeping the full reasoning engine for the final probability judgment and scenario distribution. Small decisions get fast judgment. Big decisions get deep simulation. Jev and Yarrow are different — but as a technology stack, they could be highly complementary. Stay tuned for Yarrow’s upcoming research. 👀 ☘️
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
5
537
The biggest threat to AI forecasting isn't that the model isn't smart enough. It's that the evidence isn't clean enough. We tested this ourselves. We planted one piece of information into the evidence pack that looked completely real—but was entirely fabricated. We ran it 8 times.All 8 times, the Brier score deteriorated from 0.14 to 0.52—almost no better than a coin flip. One false statement. Total contamination. Use a smarter model? It doesn't help. Give it a larger context window? Still doesn't help. Because we deliberately made the false information sound plausible, every model treated it as true. The solution isn't intelligence. It's clean evidence. ☘️ In Yarrow's evidence packs, every data point must cite an official source or be explicitly labeled as an estimate. Every conditional conclusion must be calculated under the condition that its premise actually holds. Before any forecast is published, the evidence pack goes through five layers of audit—including a red team whose only job is to find what everyone else missed. Clean evidence in. Good forecasts out. Garbage in, garbage out. What model sits in the middle? --- It doesn't matter.
1
4
161
Three thousand years ago, someone picked up fifty yarrow stalks and invented the first prediction algorithm. ☘️ Not a prayer. Not a sign from the sky. An algorithm. The I Ching's yarrow stalk method — Da Yan — was humanity's earliest systematic approach to forecasting. You take 50 stalks, set one aside, and run the remaining 49 through a series of divisions. Repeat the process eighteen times, and you get a hexagram — a structured output with a built-in probability distribution. Here's what makes it remarkable: the four possible line types don't appear equally. They show up at a ratio of 1:5:7:3. Stable lines appear about 75% of the time. Changing lines about 25%. That's not random — it's a designed distribution encoding a specific worldview: change is possible, but stability is the default. And that one stalk you set aside at the start? It's never used. It represents what can't be modeled — irreducible uncertainty. The system's designers built in an acknowledgment, three millennia before anyone formalized the idea, that no prediction framework captures everything. The jump from oracle bones to yarrow stalks was the jump from "reading signs" to "computing probabilities." Oracle bones crack once and you interpret the pattern. Yarrow stalks run a repeatable process and produce a structured result. Same question, different day, same method — that's an algorithm. The plant itself carries the same duality across cultures. In China, it's the sacred tool of the Book of Changes. In Greece, it's named after Achilles — Achillea — who used it to heal wounds on the battlefield. Celtic and Norse traditions tied it to prophecy and dreams. Wherever it shows up, yarrow sits at the intersection of knowledge and uncertainty. That's why we named the company Yarrow. Not because we do divination — but because we believe the core idea was right all along: take uncertainty, run it through a structured and repeatable process, and produce a probability you can act on. The stalks are different now. The process is different. But the principle is the same. ☘️
3
2
4
363
Will an international court find Israel or its leaders guilty of genocide by December 31, 2027? 🏦 Market: 14% VS Yarrow: 2% ☘️ We're seven times lower than the market. Here's why this isn't a political call — it's a calendar call. The ICJ genocide case against Israel has specific procedural deadlines. South Africa's Reply is due November 2027. Israel's Rejoinder is scheduled for May 2029. Oral arguments won't begin before late 2029 at the earliest. A final ruling producing a genocide finding before December 31, 2027 is procedurally near-impossible. The court hasn't even finished the written phase. Eight of our analysts independently looked at this. Seven of them cited the same concrete court dates. The consensus wasn't built on geopolitical opinions — it was built on a legal calendar that nobody can accelerate. This is actually the same pattern we see on other emotionally charged markets: people price how they feel about the issue, not how the institution actually works. The ICJ won't rule on whether something is genocide based on public pressure — it will follow its own timeline, and that timeline extends well past the resolution date. What would change our view: an extraordinary expedited proceeding that skips the written phase, or a separate tribunal convened outside the ICJ with a faster mandate. Neither is currently scheduled.
1
7
177
On October 16, 1929, the most famous economist in America stood up and said: "Stock prices have reached what looks like a permanently high plateau." His name was Irving Fisher. Yale professor. The most respected economic mind of his generation. Twelve days later, the stock market crashed. Black Monday — down 13%. Black Tuesday — down another 12%. Over the next three years, the Dow fell from 381 to 41. An 89% drop. Fisher personally lost around $10 million — roughly $175 million in today's money. He went bankrupt. The crash triggered the Great Depression: US GDP fell 30%, unemployment hit 25%, and 9,000 banks failed. ☘️ Now imagine Fisher had access to something like Yarrow before making that call. He gives the system his thesis: "stocks have reached a permanent plateau." The system doesn't argue. Instead, it runs it through six independent analysts — each working from the same frozen data, but on different models, blind to each other. Three of them flag that stock prices have outpaced earnings for years. Two note that margin lending is at record levels. One builds a scenario where a single wave of margin calls triggers a cascade. The system returns: "Probability that stocks remain at current levels through Q1 1930: 22%. Probability of a decline exceeding 30%: 41%. Here's the evidence for and against — frozen before the outcome." Fisher might still have disagreed. That's his right. But he would have seen the other side — with evidence, not just opinions — before staking his reputation and his fortune on a plateau that turned out to be a cliff. That's the difference Yarrow is designed to make. Not to replace the expert. But to make sure that before any big call, someone — or something — has already asked: "what if you're wrong, and here's specifically how."
1
6
167
In 1992, Britain had a problem. They'd promised to keep the pound tied to the German mark — but their economy was too weak to hold that promise. Keeping the peg meant keeping interest rates high, which was crushing British businesses and homeowners. Everyone in the government said the peg would hold. The Bank of England said the peg would hold. The whole establishment was committed. George Soros looked at two numbers: what the British economy could actually support, and what the peg required. They didn't match. So he bet $10 billion that Britain would have to give up. On September 16, the Bank of England tried everything. They raised interest rates twice in one day — from 10% to 12%, then to 15%. They burned through $6 billion in reserves buying pounds. None of it worked. By evening, Britain quit. The pound dropped 15%. Soros made about $1.5 billion in a single day. Soros didn't have insider information. He had the same data as every economist at the Bank of England. He just asked the question they refused to ask: "what if the policy can't survive its own economy?" That's the kind of question that gets buried in institutions. Not because people are stupid, but because there's no system for forcing it to the surface. No one at the Bank of England ran the scenario: "what happens if we can't hold the peg?" If they had, the answer would have been obvious — and $6 billion in reserves wouldn't have been burned defending an indefensible position. ☘️ This is exactly what Yarrow is designed to do. Give it a policy. It simulates the scenario where the policy works — and the scenario where it doesn't. Both get a probability. Both get a reasoning chain. And both get published before the outcome is known. Soros made his fortune because one institution refused to stress-test its own policy. Yarrow exists so the next one doesn't have to learn that lesson the hard way.
1
1
98
In March 2024, Venezuelan President Nicolás Maduro was indicted by a U.S. federal court on charges of narco-terrorism, drug trafficking, and corruption. Prosecutors allege he co-led a cocaine smuggling network that moved over 250 metric tons into the United States, using the proceeds to fund his political operations. It's one of the rare cases where a sitting head of state has been charged by a foreign government with crimes that could carry a life sentence. Now there's a market on Polymarket: will Maduro be sentenced to at least 60 years? 🏦 Market: 38.5% VS Yarrow: 19% ☘️ Our probability is half the market's. Here's why. We don't treat this as a single prediction. It's a chain of four things that must happen in sequence: Step 1: Maduro's immunity gets denied. Legal proceedings are moving forward, and precedent from other authoritarian leaders suggests this clears. Around 80%. Step 2: Trial and sentencing finish before the December 2027 deadline. Federal sentencing in complex narco-terrorism cases typically comes 3-6 months after conviction. The timeline is genuinely tight. Around 65%. Step 3: Conviction. Given the weight of evidence in cases like this — around 75%. Step 4: The sentence reaches 60 years or more. Typical narco-terrorism sentences run 20-40 years. 60+ is the upper end of judicial discretion. Conditional on conviction, around 45%. Multiply them: 0.80 × 0.65 × 0.75 × 0.45 ≈ 0.18. The market gives this nearly 40%. That means either they think one of these steps is easier than the evidence suggests, or they're not running the chain math. Our view is that when four independent conditions all have to clear, the combined probability drops fast. When we'd adjust: if the trial moves to an expedited schedule with sentencing before mid-2027, or prosecutors explicitly signal they're seeking consecutive maximum terms.
2
5
640
The 2026 IPO crown: SpaceX or Anthropic? 🏦 Market picks Anthropic — 59% chance of the highest first-day-close market cap this year VS ☘️ Yarrow picks SpaceX — 65.7% We only give Anthropic 32.3%. A 25-point gap. Same question, completely opposite answer. 👀 Here's Yarrow's case: As of today, Anthropic has not filed a public S-1 on EDGAR. If Anthropic wants the 2026 IPO crown, it needs to file, roadshow, price, and list — all before December 31. That's weeks, not months. Meanwhile, SpaceX's private valuation already sits above Anthropic's, backed by Starlink revenue, launch contracts, and Starshield defense work. One detail the market may be overlooking: xAI merged into SpaceX in February 2026 and can no longer list independently. But the venue still shows a 25.5% quote on xAI — that's a midpoint on an empty order book, not a real contender. Remove that, and SpaceX's effective odds look even stronger. We started out the same way the market did — running the two legs as separate predictions. The probabilities added up to over 100%, which is impossible for a "who wins" race. So we rebuilt it as a single joint forecast: SpaceX, Anthropic, OpenAI, xAI, and everyone else. Probabilities sum to 100% by design. After the correction, Anthropic moved down. SpaceX moved up. The market is pricing the AI hype cycle. We froze the evidence chain and priced the filing calendar and the balance sheet. When we'd adjust: Anthropic files an S-1 with a price range implying more than $2.1 trillion, or SpaceX still hasn't filed by end of October. 📅 December 31 settles it 👋
Will SpaceX have the highest first-day-close market cap among 2026 IPOs? 🏦 Market: 39% VS Yarrow: 64% ☘️ 😺 What this market is really trading: if SpaceX goes public in Q4 as expected, can its first-day close beat every other IPO this year? The main rival is Anthropic, targeting an October listing. SpaceX's private valuation already sits above $350 billion. To win, its first-day close needs to top Anthropic's debut. We see three advantages SpaceX has that Anthropic doesn't: First, diversified revenue. Starlink, launch contracts, Starshield defense work. Anthropic is growing fast but relies mostly on API services. Wall Street tends to give a higher multiple to companies that aren't built on a single revenue line. Second, retail demand. SpaceX is one of the most anticipated IPOs in a decade. Musk's retail following will drive first-day volume in a way that Anthropic — known mainly within the tech community — probably can't match. Third, defense premium. Government contracts and national security relevance give SpaceX a valuation floor that pure AI plays don't have. The market prices this at 39% — below a coin flip. We think SpaceX is the clear favorite. That said, our own analysts are split on this one, but we're sharing our reasoning for others to evaluate. What could flip this: Anthropic lists first and its first-day close blows past $350B. Or SpaceX pushes its listing to 2027. Either event could reshape the odds entirely.
2
6
14,765
In October 2002, the U.S. Intelligence Community published a National Intelligence Estimate: "high confidence" that Iraq possessed stockpiles of biological and chemical weapons and was rebuilding its nuclear program. This assessment was presented to Congress and the UN Security Council. It became the basis for one of the most consequential military decisions of the 21st century. After the invasion, the Iraq Survey Group searched the country for over a year. Weapons stockpiles found: zero. Active weapons programs found: zero. The 2004 Duelfer Report concluded that Iraq had dismantled its WMD programs back in the 1990s. This prediction was completely wrong. What did it cost? - Over $2 trillion in direct spending — some estimates reach $4 trillion including long-term veteran care and interest - 4,488 U.S. service members killed - An estimated 134,000 to 300,000 Iraqi civilians dead across two decades of conflict - Damage to U.S. credibility in the Middle East that still hasn't fully recovered The cost was devastating. So what's the lesson? At the national level, predictions don't go wrong just because there isn't enough intelligence. They go wrong because the process of turning intelligence into a forecast lacks discipline — and lacks truly useful tools. Analysts were under political pressure to reach a specific conclusion. Dissenting views were buried in footnotes and lost to history. Nobody stress-tested the most fundamental assumption: "what if these weapons simply don't exist?" Twenty years later, this lesson still hasn't been fully absorbed: a confident prediction without a system for challenging it isn't intelligence — it's just an expensive guess. ☘️ Yarrow was born to fix exactly this.
2
6
300
Will Russia capture the Kupiansk-Vuzlovyi railway station by September 30? 🏦 Market: 9.6% VS Yarrow: 18.5% ☘️ Both sides agree it's unlikely. But we see nearly twice the probability the market does. Kupiansk-Vuzlovyi is a railway junction in Kharkiv Oblast — a logistics node that matters. Russian forces have been grinding toward it for months. The ISW map, which is how this question resolves, shows them within a few kilometers. The market is looking at this and thinking: "they haven't taken it yet, so they probably won't by month-end." We looked at the same map and asked a different question: at the current rate of advance, can they get there in 28 days? The math says it's tight but possible. Russian forces have been moving slowly in one direction for six months without a major reversal. September is historically the last operational window before autumn mud season — when both sides push hardest because they know the ground is about to freeze movement for weeks. This isn't a prediction that Russia will win the war. It's a prediction that the market is underpricing how far a slow-moving front line can travel when it's been trending one direction for half a year. What would change our view: a successful Ukrainian counterattack that pushes the line back, or a ceasefire that freezes current positions.
5
171
Nike's market cap peaked at $280 billion in November 2021. Recently, it's just $57 billion. Dropped from the S&P 100. 🤔 What happened in between is a textbook case of what goes wrong when a company makes a major decision based on an unverified forecast. ❌ Nike made the wrong bet: GO all-in on direct-to-consumer (DTC). Cut ties with hundreds of wholesale partners. Pour everything into Nike's own digital channels and apps. The underlying forecast: "DTC growth will continue, and we won't need retail partners anymore." Nobody predicted the consequences of that decision. Nobody scored it. Nobody asked: "what if consumers go back to physical stores after COVID?" They went back. And Nike's products were nowhere to be found. The retailers who'd been cut gave their shelf space to On and Hoka. A former Nike exec who spent 20 years in basketball resources put it this way: "Nike lost its product category experts and their insights." The DTC pivot also shifted Nike's marketing from "create demand" (attract new customers) to "serve and retain demand" (keep existing customers). For a brand that spent decades building desire through brand advertising, this was a fundamental miscalculation. The result: Nike's sales slowed, advertising quality declined, and competitors took ground in running — a category Nike had dominated for decades. One unverified forecast. One unchecked assumption. And a former giant lost $223 billion in market value. Now imagine a different timeline — one where Nike's strategy team had a system like Yarrow telling them: ☘️ "Your DTC forecast has only a 30% chance of holding if physical retail rebounds. Here's the scenario distribution. Here's what you lose if your wholesale partners don't come back." ☘️ That's what it means for companies to have prediction infrastructure. And that's why Yarrow exists. Yarrow doesn't replace decision-makers — it makes sure that every major decision you make has been tested and forecasted before the bad outcome happens.
BREAKING: After falling -80% from its record high, Nike, $NKE, will be removed from the S&P 100 at the end of this month, ending a near 18-year run in the index. The stock has now erased -$230 billion in market cap since its all time high. A collapse for the history books.
1
5
4,535
Will SpaceX have the highest first-day-close market cap among 2026 IPOs? 🏦 Market: 39% VS Yarrow: 64% ☘️ 😺 What this market is really trading: if SpaceX goes public in Q4 as expected, can its first-day close beat every other IPO this year? The main rival is Anthropic, targeting an October listing. SpaceX's private valuation already sits above $350 billion. To win, its first-day close needs to top Anthropic's debut. We see three advantages SpaceX has that Anthropic doesn't: First, diversified revenue. Starlink, launch contracts, Starshield defense work. Anthropic is growing fast but relies mostly on API services. Wall Street tends to give a higher multiple to companies that aren't built on a single revenue line. Second, retail demand. SpaceX is one of the most anticipated IPOs in a decade. Musk's retail following will drive first-day volume in a way that Anthropic — known mainly within the tech community — probably can't match. Third, defense premium. Government contracts and national security relevance give SpaceX a valuation floor that pure AI plays don't have. The market prices this at 39% — below a coin flip. We think SpaceX is the clear favorite. That said, our own analysts are split on this one, but we're sharing our reasoning for others to evaluate. What could flip this: Anthropic lists first and its first-day close blows past $350B. Or SpaceX pushes its listing to 2027. Either event could reshape the odds entirely.
2
5
942
Will Trump's approval hit 35% on Silver Bulletin this year? 🏦 Market: 26.5% chance VS Yarrow: 11.5% ☘️ Here's the setup. Silver Bulletin tracks Trump's approval as a daily trend line — one decimal, any single day counts. Right now it sits at about 38.6%. This summer it drifted from 39.7% down to a 2026 low of 38.1%. 👀 To hit 35%, it needs to fall another 3.6 points in four months. For context, 35% would be three points below the lowest it's been all year. The market seems to think: "it's been falling all summer, so it'll keep falling." We looked at the same trend and asked: when was the last time a sitting president's approval dropped 3.6 points in four months without a major crisis? The answer is basically never — not without a recession, a government shutdown, or some kind of shock event. So this is a real disagreement about how floors work. The market is betting the slide continues. We're betting that approval ratings have gravity — they fall until something catches them, and 38% has been catching this one all year. What would make us reconsider: if Silver Bulletin drops below 37% for several days running AND something concrete drives it — a recession signal, a December shutdown, $4+ gas. Without a trigger, trends don't break floors. They bounce.
1
3
306
Something interesting happened before the 2008 financial crisis. In 2005, a hedge fund manager named Michael Burry did something nobody else on Wall Street would do — he sat down and read thousands of individual mortgage loan documents, traced the data, and built his own forecast. Here's the thing: the data was public. The mortgages were public. The default rates were sitting right there. But every major bank, every rating agency, every regulator looked at the same market and reached the same conclusion: it's safe. Burry was screaming from his office that the market was about to collapse. He put $1.3 billion of real money behind that call, betting against subprime mortgages. All of Wall Street laughed at him. His own investors wanted out. He nearly got fired. Then 2008 happened. The market collapsed exactly as he predicted. Burry personally made $100 million and returned $725 million to his investors. Meanwhile, the crisis wiped out $16 trillion in household wealth and 8.7 million jobs. So what's the lesson? That Burry was smarter than everyone else? No. He did the homework that a good forecasting institution is supposed to do — study the real situation, make an honest call, and hold the line when the entire market disagrees. That's discipline. Good predicting comes from seeing the evidence others overlook, combining data attribution with scenario simulation, and of course, a system that reflects on itself and keeps evolving.
2
8
12,859
Will OpenAI IPO by December 31, 2026? 🤖 🏦 Market: 16.5% VS Yarrow: 8.5% ☘️ The company filed its S-1 confidentially in June. Then on August 19, its finance chief publicly said the plan is to list in 2027. Since then — no public filing, no updated timeline, no reversal. Here's what our six independent analysts saw when they looked at this: The signaling math doesn't add up for 2026. A premature listing means absorbing roughly $60 billion in burn against the stated 2027 window. The CFO publicly committed to 2027 — reversing that within 128 days creates a credible commitment cost that most companies avoid. SEC review alone takes 3-6 months minimum. Management's incentive right now points toward valuation protection and filing-to-listing lead time, not a rushed debut. One analyst gave this a 25% chance, noting it's not impossible if OpenAI decides to absorb very large expected losses. The other five ranged from 3% to 15%. The median: 8.5%. As of August 25, reports still describe the company as "ahead of IPO" — not "going public this year." Amazon just closed a $50 billion investment commitment. The market is giving one-in-six odds that a company will IPO this year after its own CFO said it won't. We think one-in-twelve is closer to reality. ☺️ —— What would change our view: a public S-1 appearing on EDGAR before end of October, or the finance chief walking back the 2027 timeline.
2
1
3
1,196