Trying to figure out how search works

Doug Turnbull retweeted
Personally I believe our “late chunking” arxiv.org/abs/2409.04701 is elegant enough and practical enough, it’s probably still undervalued but whatever I’m fine with that. btw the best person to answer contextual embedding is probably @jxmnop
2
1
5
376
Doug Turnbull retweeted
I feel this pain. The corollary here that I’ve noticed: if you send cold emails or messages to me that *aren’t* written or inspired by AI, you are *way* more likely to get a response than in the pre-AI days. In some ways, the bar has never been lower
"Sorry for the delayed response! Middling generative AI has fundamentally altered the strategic balance between offense & defense in the war for our attention, so I have chosen to ignore this messaging channel. Out beyond kino & slop, there is a field. I'll meet you there."
7
6
50
5,359
The emphasis on chunking as a *retrieval* strategy for single-vector search has long been misguided. You can't a-priori anticipate every query some text might answer. Then try to focus in on that bit of text (cutting off information). And a single vector obviously can't capture pages and pages of information
5
2
32
2,734
And why I'm talking about why Embeddings Don't Solve RAG on Monday
1
2
11
1,237
Follow William for these tidbits - I'm pretty sure he knows much more about embeddings than me :)
Replying to @softwaredoug
contextual embeddings train "single vector" embedding models to resolve this issue by modeling the problem end to end during training: blog.voyageai.com/2025/07/23… today contextual embeddings are only available from perplexity (open and mit licensed) and voyage (latest being voyage-context-4).
1
5
760
An important criteria for what vector database I use, seems to be wherever @benwtrent works
12
1,816
RT @ArkidMitra: They don't about this enough. This is the highest ROI for any company building in house company brains/commerce search.
1
21
Kudos to @LightOnIO for being ChatGPT's preferred late interaction model ;)
1
7
21
5,776
Almost nobody actually needs a knowledge graph for search. What they usually want is to organize entities in a sane way. Which probably means a hierarchical taxonomy. softwaredoug.com/blog/2026/0…
7
2
36
2,533
I've changed the mission for cheat at search
1
1
10
683
Previously my positioning was for the deep search nerds Twoish years ago, it was about convincing them to use LLMs for their work. That's not particularly hard proposition these days :)
1
1
150
Old and busted: search engines based on grep New hotness: search engines using jev
3
310