// building a parallel web for AIs // cofounder @p0 //

San Francisco, CA
Pinned Tweet
turbo mode is web search that's super fast (~200ms), super affordable ($1 per 1,000 searches), and high quality, opening up a whole new frontier of what agents can do with the web it's built for the scale and shape of demand the agentic web requires
Introducing Parallel Search Turbo: the fastest and most affordable way to bring the web to your AI.
2
9
75
12,129
“Ads do not work with agents in their current form. Agents show up, no one sees ads, and you make no money. We are effectively building an AdSense for agents showing up to read your content. We like to pay content owners a variable amount of money every time an agent derives benefit from reading their information.” @paraga
14
7
113
21,378
Travers retweeted
Free ad space at DevDay
7
6
107
6,640
Travers retweeted
Parallel is now live in @Lovable. Bring fresh, relevant, reliable web data into any app.
4
9
49
4,235
“Agents will use the web 1,000x more than humans. Hence, new tech is needed and new business models are needed. No tech built for a certain scale survives three orders of magnitude. When you need new business models alongside new technology, a problem becomes really interesting.” @paraga
25
25
151
25,783
it's true - in hypergrowth mode the biggest shortage is toilets, not GPUs we had to airdrop more in our parking lot to address the bottleneck
when a 50-person startup shares 2 bathrooms we call it pacing the frontier
5
50
3,860
better data = better agents @p0 can pull from crunchbase, harmonic, polymarket, pubmed and a growing list of trusted sources
Today we're launching Data Connectors in Parallel. Your agents can now work with specialized third-party data through Parallel's best-in-class agentic web research APIs. Our first partners are @AlliumLabs, @useapolloio, @Baselayerhq, @CarbonArcAI, @crunchbase, Faraday AI, @harmonic_ai, @MiddeskHQ, @Similarweb , @particle_news, @pensa_systems, and @Polymarket. We're also launching free, opt-in access to biomedical sources including PubMed, ClinicalTrials.gov, and ChEMBL.
1
4
32
5,200
Travers retweeted
It is becoming table stakes to offer private data alongside search. To me, this marks the beginning of a new era of the web, where the most valuable information no longer sits siloed in a database, but rather sits on equal footing with web data, both discoverable and paid.
Today we're launching Data Connectors in Parallel. Your agents can now work with specialized third-party data through Parallel's best-in-class agentic web research APIs. Our first partners are @AlliumLabs, @useapolloio, @Baselayerhq, @CarbonArcAI, @crunchbase, Faraday AI, @harmonic_ai, @MiddeskHQ, @Similarweb , @particle_news, @pensa_systems, and @Polymarket. We're also launching free, opt-in access to biomedical sources including PubMed, ClinicalTrials.gov, and ChEMBL.
2
7
67
12,330
Excited to partner with @p0 on Managed Deep Agents! You now have parallel built into your agents so your agents search the web fast and efficiently! @travers00 @hwchase17 @VictorMoreira16
6
18
43
19,970
opus 5.5 leading the charts, a very good model broadly and for web search
Claude Opus 5.5 takes the #1 spot on Parallel's Search Capability Leaderboard, with a +4.3 gain over the second-highest score from Opus 5.
2
22
1,429
no mogging
Parallel Search MCP is now on @salesforce AgentExchange! It gives Agentforce agents real-time, citation-backed access to the open web: web search and page-content extraction, grounded in fresh facts and data.
12
11
352
62,900
from our model leaderboard: - Astra tops the charts for overall capability, then Sol - DeepSeek v4.1 Flash gets Sol quality at 1/5th the cost - DeepSeek v4.1 Flash also shows the biggest intelligence lift from search vs. base model
today we launched a leaderboard for how frontier models perform on web search tasks "which model should I use for search-heavy work" is a question we get constantly and there weren't good public evals for it. nothing that tests search capability end to end, with cost and latency next to accuracy so we made the methodology we use internally public. 24 models w/ same search, with and without web access
2
4
26
2,319
today we launched a leaderboard for how frontier models perform on web search tasks "which model should I use for search-heavy work" is a question we get constantly and there weren't good public evals for it. nothing that tests search capability end to end, with cost and latency next to accuracy so we made the methodology we use internally public. 24 models w/ same search, with and without web access
Introducing the Search Capability Leaderboard. Compare leading models for agentic search: answer quality, cost, and how much each gains from web access. Built to help you choose a model for search-heavy workflows. parallel.ai/leaderboard?utm_…
1
6
51
10,596
Travers retweeted
Parallel is expanding across the pond 🇬🇧 Keep an eye on our careers page for open roles: parallel.ai/careers
1
5
36
6,078
most teams pick web search for their agent without properly evaluating it makes sense, it looks fungible on the surface. it isn't. quality, latency, token efficiency all diverge, and the gap shows up in prod. we wrote a short guide - give it to your agent.
2
2
41
3,953
Travers retweeted
One thing about running local models: they're lacking great web search. There are a ton of options, but useful benchmarks didn't exist. So I ran a bake-off this weekend... I tested 4 leading open-weight models: @deepseek_ai (V4 Flash), @Zai_org (GLM 5.3 Flash), @Alibaba_Qwen (3.8 27B & Flash Next) each with 10 search provider tiers from @brave, @ExaAILabs, @p0, @serperapi, @tavilyai. 2,000 graded answers, real money paid. Clear winner: Qwen 3.8 27B + Parallel Turbo. 50/50 correct answers, 5.7s per answer, $1.26 per 1k questions. Full data: local-search-bench.vercel.ap… All tests were run on @NVIDIAAI DGX Sparks via @MiaAI_lab's recipes, orchestrated and judged by me using Fable 5 in @AnthropicAI Claude Code.
21
24
194
17,218
scaling has multiple axes beyond parametric memory and they are rapidly becoming more legible useful intelligence = model capability × external information × inference-time allocation the model doesn't need to memorize the world. retrieval gives them real world state to reason on
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
1
2
19
1,353
better search reduces inference per task and increases total inference by making each task cheaper, better retrieval unlocks more useful work search becomes a compounding part of the inference loop not a one time tool call
2
6
52
4,063
GLM 5.3 Flash lives up to the hype. similar performance to GPT-5.6 Luna, half the per-task cost our team benchmarked GLM-5.3-Flash against GPT-5.6 Luna with a variety of different web search options to see where this new model stacks up. no surprises: using GLM 5.3 with Parallel Search tops the leaderboard and is up to 14x cheaper on end-to-end tasks.
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
6
10
67
9,674
topping the charts @Zai_org
3
324