THE PDF IS A CRIME AGAINST MACHINES
How Altitude built a four-pass AI pipeline to kill the most hated task in finance
By @deni_ersht , Co-founder & CPO at Squads
Bill pay should work like this: invoice arrives, machine reads it, payment queues. Done. The hard part is building the reading layer accurate enough to trust with a real payment.
Every finance team has that one person who is basically a human API between an email inbox and a payment form. For a company running 30-40 vendors - pretty normal for a Series A crypto startup paying for infrastructure, auditors, legal, devtools - that's 15-20 hours a month of pure data entry.
These are companies building on programmable money. They should have better tools. That's the seam where AI should sit - not in some abstract way, but a very concrete one: a document comes in, a machine reads it, and the system either pays it or asks you to confirm.
The Problem Isn't Reading. It's Knowing When You're Wrong.
Invoice parsing tools have existed for years. Finance teams still process invoices by hand not because the technology didn't exist, but because none of it was accurate enough to trust with a payment. It filled in fields confidently. Some were wrong. After correcting enough of them, it felt faster to just do it yourself.
That's the problem we designed around. The system isn't trying to be right 100% of the time. It's trying to know when it's wrong.
Why PDFs Are Hard
The PDF format was designed for human eyes, not machines. There's no semantic layer - no field called "amount due," no tag that says "this is a bank account number." OCR can tell you there's a "42,500" on the page. It can't tell you whether that's the subtotal, the tax, or the total due.
Invoices have no standard format. Bank details could be IBAN, SWIFT, ACH, or a Solana wallet address - sometimes all on the same invoice. For crypto-native companies specifically: a misread character in a wallet address doesn't produce a wrong street address. It produces money sent to the wrong wallet.
Off-the-shelf tools give you one pass, one model, one confidence score. Fine for the median invoice. The 30% that are over-styled, multi-currency, or non-standard - that's where they break. That's also where finance teams spend most of their time.
Why We Run Three Models
Most companies building AI products right now make a model commitment. They pick one, build around it, and ship. That's a reasonable bet if you're building a chatbot or a writing assistant. It's a bad bet if the output is a payment.
We run three models in priority order: Mistral primary, Gemini fallback, Claude second fallback. That's not a backup plan. It's a view that the model layer is still too unreliable and too fast-moving to commit to any single one.
Models are good, but not uniformly good. Performance varies by document type, language, layout, and the specific failure mode you're trying to solve. Mistral handles the majority of clean invoices well. Gemini catches things Mistral misses on messier documents. Claude catches things both miss. No single model dominates across the full distribution of invoices we see.
There's also a second reason. The AI space moves fast. A model that leads on document extraction today may not lead in six months. By owning the orchestration layer - rather than building around a single provider's API - we can swap models, adjust cascade priority, and run our own evals as the landscape shifts. We tested against 2,000+ real invoices to arrive at the current order. We'll keep collecting data and adjusting the cascade as the model landscape evolves.
Most companies want to say "we use [model]" because it sounds confident. Running a cascade isn't hedging. It's an honest acknowledgment that no single model should be the last word on a payment.
The 4-Pass Pipeline
The cascade runs inside a four-pass pipeline. Each pass exists because of a specific failure mode we observed in real invoices. If pass one returns 100% confidence, we stop. If not, we keep going. Anything still uncertain after three passes gets flagged for human review. Pass four always runs - pure logic, no AI.
PASS 1: One-shot extraction
Mistral primary, Gemini fallback, Claude second. All fields extracted in a single call. 100% confidence across the board - pipeline stops. That's 30-40% of invoices. For the rest, pass two runs.
PASS 2: Chain-of-thought re-extraction
The naive fix is to retry the same model on the same document. We don't do that - retrying the same model on the same input produces the same output. Instead, pass two runs a fresh extraction using format-aware chain-of-thought prompting, using what pass one learned about the document's number format, date style, and language. The fallback provider receives those observations alongside explicit markers for what came back uncertain, and reasons through the gaps step by step.
Classic OCR returns "$10,000." A language model reasons that $9,200 + $800 tax = $10,000 and identifies "Total" at the bottom of the column as the right field. That's semantic understanding of document structure.
PASS 3: Targeted field extraction
Still below threshold. The document has a specific problem - non-English text, handwritten notes, multi-currency with implicit FX, credits that make the correct total genuinely ambiguous. Re-reasoning the whole document won't fix a problem isolated to two fields. Focused prompts fire in parallel for unresolved fields only.
PASS 4: Cross-field validation
Runs on every invoice that makes it past pass one with less than 100% confidence. No AI. Do line items sum to the subtotal? Does subtotal + tax equal the total? Is invoice date before due date?
Banking details get a separate check against the user's vendor contact book. Vendor name matches but banking details don't - flagged immediately, regardless of confidence score. That's the most common invoice fraud pattern: a familiar vendor name with substituted payment details. We flag every mismatch. Friction is what keeps intelligence from outrunning accuracy.
The Tool Should Disappear
@ivanhzhao put it well: the best tools disappear into the workflow. The tool should amplify your intent, not announce its presence.
For bill pay that means: when you upload an invoice and everything is correct, you shouldn't think "the AI is really good." You should think "great, next invoice."
The shift we're after is from data entry to data review. Most fields filled, marked confident. A couple need a glance. The user's job becomes judgment, not typing. The most valuable thing AI does in finance isn't speed - it's removing the gap between information and action. Today there's a PDF sitting in someone's inbox for three days before it becomes a payment. That latency is pure waste.
Agentic Payables
@btaylor recently said the atomic unit of AI productivity is a process, not a person.
That's exactly how we think about bill pay. The old world had three people touching an invoice - someone receives it, someone enters the data, someone approves it. We didn't build this to make each of those people 20% faster. We automated the process. The productivity gain isn't incremental. It's structural.
Eliminating data entry is just the first layer.
The next question follows naturally: why does a human need to approve a routine invoice? Known vendor, recurring amount, approved budget line - those conditions are fully determinable from structured data. A policy engine evaluates them. An agent executes the payment. The CFO's job shifts from processing to judgment.
This is where we're building. The extraction pipeline is live. The policy engine is in development. The thesis: once you have reliable document intelligence feeding into onchain policy enforcement, routine payments don't need a human in the loop.
What makes that safe is that the policy layer lives onchain. The agent operates under spending limits enforced by a smart contract - not application logic, not a database flag. Even worst case - the AI hallucinates an amount, there's a prompt injection buried in a malicious invoice - the smart contract says "this exceeds the per-transaction limit" and the transaction simply doesn't happen.
A bank's spending limit is a database flag an admin can override. Ours is a smart contract that even we can't override.
Everyone's trying to solve AI safety at the model layer. We solve it at the infrastructure layer. The model can be as wrong as it wants. The smart contract doesn't care about confidence scores - it cares about spending limits.
Policies set by humans. Executed by agents. Enforced by code.




