I profess CS-ly at @UWaterloo. Previously, I monkeyed code for @Twitter, slides for @Cloudera, and scienced for @yupp_ai.

Nearby data lake
I'm in a good mood because "grep is all you need" and "BM25 is all you need" both recently drew from (separate) Bernoullis and came back positive. Now I can claim that both have been +1ed by LLMs (wait, I meant peers). #nerdref
5
59
8,996
Summer passed in a blur... ๐Ÿ–๏ธ๐Ÿฆฉ๐Ÿน and I forgot to congratulate @HosnaOH for finishing her Master's! Her work with @amirhkarimi_ "A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation" arxiv.org/abs/2605.15761 just got a +1 from NeurIPS, so yay! ๐ŸŽ‰
1
2
51
3,535
1/ ๐Ÿงต BrowseComp-Plus (ACL 2026) set the standard for disentangled agentic search eval... ๐Ÿ™Œ ๐๐ซ๐จ๐ฐ๐ฌ๐ž๐‚๐จ๐ฆ๐ฉ-๐๐ฅ๐ฎ๐ฌ_๐‚๐Œ makes it even better! ๐Ÿš€ Same BCP questions grounded on ๐‚๐ฅ๐ข๐ฆ๐›๐Œ๐ข๐ฑ-๐Ÿ’๐ŸŽ๐ŸŽ๐ with ๐Ÿ“๐Ÿ“๐Ÿ‘๐Œ ๐๐จ๐œ๐ฌ. ๐Ÿ’ป github.com/castorini/cmass
๐Ÿš€ Introducing BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent. It is a new Deep-Research evaluation benchmark built on top of BrowseComp. It features - ๐Ÿ“š a fixed, carefully curated corpus of web documents - โœ… human-verified positive documents - โš”๏ธ web-mined challenging hard negatives. With BrowseComp-Plus, you can thoroughly evaluate and compare the performance of different components in a deep-research system. e.g. GPT-5 + Qwen3-Embedding. Code, dataset, and leaderboard links are provided at the end of this thread.
1
9
34
4,469
Jimmy Lin retweeted
In the past year, we've seen rapid progress in search agents! However, how well do these search agents work outside English? Introducingโ›ต๏ธMAST @๐Ÿ”ฅFIRE 2026, a new shared task focused on multilingual agentic search.
1
8
31
6,175
Agentic reproducibility is the unlock, I claim. To what? ๐Ÿ™‹โ€โ™‚๏ธ Find out by reading my piece "Reproducibility in the Era of AI Agents" ๐Ÿค– in the June 2026 issue of SIGIR Forum. sigir.org/forum/issues/june-โ€ฆ
What I'm cooking up... ๐Ÿ‘จโ€๐Ÿณ
1
6
30
3,819
Side note: I wrote this piece back in May, ~3 months ago. It's not aged too badly, I think? ๐Ÿฅ› Such is the speed of academic publishing...
600
Jimmy Lin retweeted
And with that, all TREC RAG 2026 submissions are in ๐ŸŽ‰ We had an incredible turnout this year, with 30 teams submitting more than 80 RAG runs and 110 retrieval runs! A huge thank you to everyone who participated! We canโ€™t wait to start evaluating your submissions ๐Ÿš€
5
11
4,018
I always tell this joke in my classes about IR, and it goes something like: well, short of inventing telepathic retrieval, we remain stuck with some form of query formulation, be it a few keywords (before) or a prompt (now)... little did I know... ๐Ÿ˜œ
Two weeks ago, I resigned from OpenAI to join Conduit as a founding researcher, where we're training models to non-invasively read the human mind. I've written some thoughts about what telepathy could look like by 2035 and how to get there:
2
9
1,930
๐Ÿช  Belatedly plugging a paper that attempts to bridge information nuggets and side-by-side comparisons with @arena data for complex LLM responses (tl;dr - pointwise vs. pairwise evaluations) dl.acm.org/doi/10.1145/38057โ€ฆ Paper led by @Sahel_Sharify @Ushivani3 @beirmug!
2
3
11
1,420
This work serves as an inspiration for the evaluation methodology to be deployed at @TREC_RAG 2026 trec-rag.github.io/ - submissions due in a few days!
433
As they say, now for something different! "Rethinking Deep Learning-Based Rainfall-runoff Modelling for Data-Rich Hydrology" ๐ŸŒง๏ธ๐Ÿ’งโ„๏ธ defended by DR. Qiutong (Glen) Yu, co-advised with Bryan Tolson.
Made with AI
1
1
19
1,739
The saga continues ๐Ÿš€ All you need... grep? โœ… BM25? โœ… psssh... ๐ŸŸข๐Ÿ”ด Boolean! โ‰๏ธ
This was fun. Humans need powerful search tools, but even exact-match Boolean queries can work for agentic search. Our Boolean ranking uses only match proximity and density. No learning, vectors, or even term statistics required. arxiv.org/pdf/2607.11362
1
4
51
8,481
It's now officially ๐Ÿ‘จโ€๐ŸŽ“ DR @xueguang_ma - who yesterday successfully defended his Ph.D. ๐ŸŽ‰ Congrats!
7
9
64
15,853
Jimmy Lin retweeted
๐Ÿ’กLLM agents should do more than rewrite promptsโ€”they should actively search for and organize useful multimodal context to support the generator. Introducing SearchGen, which co-trains an LLM agent and a generator to expand the knowledge boundary of image generation. Try our Live Demo: tinyurl.com/searchgendemo Search a frozen 1M Web corpus; no live API needed. Read our project blog: haozheh3.github.io/SearchGenโ€ฆ (1/N)
1
6
17
2,092
Jimmy Lin retweeted
Our paper got an oral at the CTB workshop @ ICML๐Ÿ‡ฐ๐Ÿ‡ท How stable are LLM leaderboards? ๐Ÿค” Less than you'd hope. We built a unified framework for auditing how small changes to pairwise votes ripple through leaderboard rankings ๐Ÿ“„ arxiv.org/pdf/2605.15761 ๐ŸŒ hosnahoseini.github.io/leadeโ€ฆ
2
5
20
7,103
Jimmy Lin retweeted
๐Ÿ† Thrilled that our ACL 2026 paper โ€œTo Lie or Not to Lie?โ€ was selected as an ACL 2026 SAC Highlight! ๐ŸŽ‰ The work asks a simple but urgent question: are LLMs equally safe around the world? ๐Ÿงต๐Ÿ‘‡ @aclmeeting #ACL2026NLP #NLProc
6
16
65
4,102
[1/3] Our agentic search work is at #ACL2026! ๐ŸŽ‰ I can't make it to San Diego in person, but luckily @zijian42chen is presenting the poster โ€” go say hi! The one-line takeaway: ๐‘น๐’†๐’“๐’‚๐’๐’Œ ๐’ƒ๐’†๐’‡๐’๐’“๐’† ๐’š๐’๐’– ๐’“๐’†๐’‚๐’”๐’๐’. ๐Ÿงต
1
4
16
962
Jimmy Lin retweeted
What if a website helper were a program tree, not a single chatbot? I built one for my course from 30 ProgramAsWeights programs. Each question falls through its own path of routing, search, answering, validation. Now anyone can ask a coding agent to build one with one prompt.
1
9
18
5,271