IBM built an AI search that finds the right section of a book 82.6% of the time on the first try.
The paper is called STAIR, out of IBM Research.
When you ask an AI about a long document today, it usually runs RAG. The system cuts the document into equal-sized chunks, turns each chunk into numbers, and hands back the chunks that look closest to your question. It throws away the chapters, sections and headings on the way. Your 500-page manual becomes a pile of loose paragraphs.
STAIR keeps the structure. The model sees the book's table of contents and learns to answer each question with the right section title. The book's own outline becomes the index, so the system skips embeddings and vector databases.
To test it, the team built a benchmark from 18 books across law, medicine, finance, education and the sciences, with tens of thousands of questions.
The best comparable method got 76.9%, standard dense retrieval got 68.7%, and keyword search got 59.5%. An off-the-shelf Mistral model got 13.8%.
The viral posts lead with the hallucination rate. STAIR pointed to a section that doesn't exist in 0.05% of answers. The closest method did that in 3.25%, which is where the "65x" comes from. That number is about finding the right section. The paper doesn't test whether the answer built from that section is correct.
The viral posts also skip the limits. STAIR needs a document with a table of contents. The team fine-tuned a 7B model on each book separately, 200 passes per book. They haven't tested it on company-sized collections.
The idea is the part worth keeping. Authors spend months organising a book into chapters and sections, and STAIR puts that work to use.
Most AI search ignores the table of contents.
This paper shows it's the best index you already have.