Organizing human intelligence to power the AI economy.

San Francisco
Pinned Tweet
Join @brendanfoody and Harvey's @gabepereyra for the debut episode of The Frontier podcast. Learn about: - Frontier and open source AI - Post-training models for professional work - Future of AI and legal work Check out the episode below.
Replying to @gabepereyra
@gabepereyra is the President & Co-Founder of @harvey, the most prominent legal AI startup and an industry leader in owning their own intelligence. I'm excited to share our conversation, which spans closed & open source, post-training, and how the law profession will evolve.
4
1
25
7,329
Claude Opus 5.5 is the new #1 on APEX-SWE. Overall, Opus 5.5 scores 67.6% Pass@1, a +3.9 point gain over the previous leader, Opus 5, at 63.7%. APEX-SWE measures observability, debugging from production telemetry, and integration, if a model can build an end-to-end system. Compared with Opus 5: Observability Opus 5.5: 69.8% (#1) Opus 5: 63.5% (#2) It gains +6.3 points. Opus 5.5 leads the board by 6.3 over Opus 5 and 10.8 over Fable 5.1. Integration Opus 5.5: 65.5% (#4) Opus 5: 64.0% (#9) Opus 5.5 improved at reading telemetry and finding faults, but building systems from scratch moved less. Congrats to @AnthropicAI. APEX-SWE leaderboard: mercor.com/apex/apex-swe-lea…
10
6
36
2,609
Mercor is committing $5M to a new AI Capabilities Fund. In conversations with researchers and enterprises, one challenge that comes up repeatedly is that many AI evals do not reflect how models are used. But to provide useful signal for eval and training, benchmarks need to capture actual workflows in representative environments. This will only become more difficult as models are used for more high-value, complex tasks. The Capabilities Fund supports researchers working on frontier eval challenges across a range of domains and data shapes, including: - Reward calibration and aligning grades with expert preferences - Creating high-quality realistic environments - Long-horizon workflows - Ambiguous requests and under-specified outcomes - Navigating social, temporal, and business context We will fund researcher time, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts, as well as free use of our evals and analysis platform. Grants are available to independent researchers, non-profits, and academics. To submit an Expression of Interest, go to: mercor.com/careers/?ashby_ji…
6
5
34
3,080
OpenAI GPT‑6 Sol and Luna are on the APEX leaderboards. GPT-6 Sol scores 54.3% Pass@1 (#18) on APEX-Agents. That is +2.9 pts over GPT-5.6 Sol. GPT-6 Luna scores 44.3% Pass@1 (#29), +1.3 pts over GPT-5.6 Luna. On APEX-SWE, Luna scores 38.8% Pass@1, +6.7 pts. OpenAI reduced API prices for Sol and Luna by 50% compared with their GPT‑5.6 promotional pricing. GPT-6 Sol improves in finance and consulting domains: Investment banking: 55.2% Pass@1 (#13), +5.5 pts over GPT-5.6 Sol Management consulting: 51.6% (#16), +9.0 pts GPT-6 Sol is only 0.7 pts behind GPT-6 Astra in investment banking. These gains are not consistent across all domains. Compared with GPT-5.6, Sol’s corporate law score drops 5.7 pts to 56.2% (#27). Luna drops 2.8 pts to 55.6%. GPT-6 Luna improves most on coding work: On APEX-SWE, GPT-6 Luna scores 38.8% Pass@1, gaining +6.7 pts over GPT-5.6. Integration: 59.3%, +13.3 pts Observability: 18.3%, no change Sol scores 45.0% Pass@1 (#18), about level with GPT-5.6 Sol (45.8%). Integration: 62.3% (#13), level with GPT-6 Astra (62.0%) Observability: 27.8% (#19), down 3.7 pts For both models, debugging with production telemetry is still the hardest part. Congrats to the @OpenAI team on the launch. See full leaderboard: mercor.com/apex/apex-agents-…
5
1
30
1,965
Mercor has now partnered with > 300 enterprises to help them monetize their data, paying companies up to $10M per enterprise. Across these companies, we’ve encountered 90 different systems: -55% contained Slack -53% GitHub -53% Google Drive -50% Gmail -40% Figma -35% Jira A company with 200 employees runs on many of the same systems as an enterprise with 5,000+. For AI to do economically valuable work, it needs to understand how work actually happens across these systems — not just within any one of them.
23
17
277
418,548
APEX is now on the Specialized Intelligence Index, @FireworksAI_HQ's new leaderboard for real-work benchmarks by industry. We built APEX to test frontier models’ ability to do professional work. We’re proud to partner with Fireworks to make these benchmarks more widely available. APEX General Practitioner (MD) APEX-Agents Corporate Law APEX-SWE Find out which models lead each category. See the full results on SII: fireworks.ai/specialized-int…
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day. Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
3
25
3,032
Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1) APEX-Accounting: 15.4% Pass@1 (#1) 62.0% mean score (#1) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1) Corporate Law: 71.2% Pass@1 (#3) Investment Banking: 69.3% Pass@1 (#2) Accounting: 62.0% mean score (#1) Token use can indicate domains where the model is most effective. Consulting leads at 80% and uses the fewest tokens at 2.0M per attempt. Law has the highest partial credit (85.7% mean) but costs the most at 4.9M tokens. In Accounting, Opus 5.5 gets partial credit on most tasks but fully passes only 1 in 6. More tokens can increase capability on Opus 5.5. Max effort consumes 3.50M tokens per attempt. That’s 2.1x more than Opus 5 and 1.2x Fable 5.1. But list price fell to $4/$20 per M (Opus 5 was $5/$25), and cache reads are $0.20, so 2.1x the tokens is only about 1.7x the dollars. Effort level has a significant impact on benchmark scores and token usage. Medium effort: 52.3% on 756k tokens Max effort: 73.4% on 3.50M tokens Increasing effort to max gains 21 points for 4.6x the tokens on agentic work. On APEX-Accounting, the same jump buys 2.8 points for 3.4x the tokens. When Opus 5.5 solves a task, it solves it consistently. 142 of 239 tasks passed on all 4 runs. Failing runs used 3.1M tokens vs 1.9M for passing runs, and took twice as long. 13 runs hit a hard failure, all in Corporate Law. 8 of them burned 25M to 42M tokens before dying, showing that long-horizon legal work can still send the model into a loop. Congratulations to the @claudeai team. See full leaderboard: mercor.com/apex/apex-agents-…
9
4
62
20,952
Mercor retweeted
Don’t miss newly announced Fireworks Forge speaker: @BrendanFoody, CEO and Co-founder of @Mercor. Join the people, teams, and companies building their own frontier on open models. Nov 3, San Francisco. Apply to attend: fireworks.ai/forge#apply
1
4
15
3,739
Grok 4.7 is now live on the APEX leaderboards. APEX-SWE: 53.6% Pass@1 (#5) APEX-Agents: 54.6% Pass@1 (#13) Compared to other models with similar cost and latency profiles, it’s a strong model for agentic coding tasks. Congrats to the @SpaceXAI team.
2
2
33
14,340
Grok 4.7 is very token efficient. On APEX-SWE, it uses 46% fewer completion tokens per task than Grok 4.6 (23.8k vs 44.1k). Full leaderboard: mercor.com/apex/apex-swe-lea…
1
11
2,081
For long-horizon tasks on APEX-Agents, Grok 4.7 is much faster than its full-size sibling, Grok 4.6. Built for speed, Grok 4.7 runs each APEX-Agents task with 60% fewer completion tokens and finishes 2.5x faster than its full-size sibling, Grok 4.6. Median time per task falls to 3.1 minutes from 7.7.
1
2
280
Four different Flash models now rank within the top 15 on APEX-Agents. Gemini 3.7 Flash: 67.8% (#2) Gemini 3.8 Flash: 64.3% (#6) Grok 4.7: 54.6% (#13) GLM-5.3-Flash: 52.8% (#15) Grok 4.7 leads GLM-5.3-Flash by 1.8 points on 240 tasks. Full leaderboard: mercor.com/apex/apex-agents-…
1
2
222
Mercor retweeted
I’ve been getting lots of DMs asking me how to pass Mercor.com assessments. 👀 Well… Mercor just made things a lot easier. They’ve released Mercor Academy — a learning platform where you can build and certify skills that are relevant to the work available on Mercor. You can now take courses, learn the skills you need, and earn certifications that appear on your Mercor profile. 🎓 One of the courses currently available is “Introduction to Agents”, where you learn what AI agents are, the components behind them, and how they work. If you’re trying to get into AI training, data work, or other opportunities on Mercor, this is definitely something worth checking out. Instead of just applying and hoping you pass the assessment, you can actually build the skills first. Mercor is basically giving you the roadmap. 🗺️ Go to your Mercor account → Learn → Academy and see what courses are available.
18
18
158
9,133
Mercor retweeted
We’ve just added three new speakers to our incredible Fireworks Forge lineup! We are proud to welcome @BrendanFoody @pirroh and @jefftangney to the stage as we bring together the people, teams, and companies building their own frontier on open models.
7
8
45
9,147
Replying to @gabepereyra
@gabepereyra is the President & Co-Founder of @harvey, the most prominent legal AI startup and an industry leader in owning their own intelligence. I'm excited to share our conversation, which spans closed & open source, post-training, and how the law profession will evolve.
8
12
146
595,953
People often ask me what books and podcasts I recommend for learning about AI. The truth is, I learn most about the future of AI and the data market through conversations with leading AI researchers. I want to make more of those conversations available to everyone, so I’m starting a podcast interviewing the people shaping the frontier of AI. More soon 👀
13
16
311
315,760
We’ve now encountered 90 different systems across 300+ company datasets. Company data isn’t a single corpus sitting in a database waiting to be exported. It’s fragmented across communications, code, documents, CRM, finance, design, project management, and dozens of specialized systems. So licensing real-world data isn’t really a “data marketplace” problem. It’s an infrastructure problem. Connectors, permissions, extraction, identity resolution, anonymization, validation, and packaging all have to work before the data becomes useful for research. The companies that solve that infrastructure layer will unlock a very different class of training data.
6
5
75
8,403
We’re bringing together founders, investors, and operators to talk about what we’re learning as this new market takes shape. We’ll discuss why real-world company data matters for frontier AI, what makes these datasets valuable, and how companies can participate in data partnerships with Mercor. Join us October 8 in San Francisco. RSVP: partiful.com/e/i28rxPtnkHXox… #SFTechWeek
2
4
1,434
Harvey is leading the wave of app-layer startups post-training models for their domain, in the same way Cursor did in code. @asadovsky is one of the best post-training leaders I've worked with across MAI and GDM. I couldn't be more excited for what this team accomplishes.
Excited to welcome @asadovsky as Harvey’s Chief Research Officer. Before Harvey, Adam co-led post-training at Microsoft AI and Google DeepMind. As a CVP at Microsoft AI, he helped build MAI-Thinking-1, Microsoft’s reasoning model. As part of Gemini’s leadership team he helped train Gemini 1.0 through 2.5, including fine-tuning, RL, data, and evals. His prior work as a Distinguished Engineer at Google spanned Assistant, Search Quality, and Search Infrastructure. I met Adam three years ago when I sent him a cold LinkedIn DM and was surprised he responded. At a time when most dismissed the application layer and legal, Adam was curious and generous with his time. He quickly became someone I regularly turned to for advice on AI as we scaled Harvey over the past three years. When we first met, we were too early to hire someone of his caliber and scale, but I always hoped we’d eventually work together. As Winston and I got to know him better, what stood out even beyond his technical achievements was his character. Despite his incredible technical career, he remains curious, humble, practical, and cares deeply about the teams he builds. We couldn’t think of a better leader to help us build frontier intelligence for the professionals and institutions we serve.
3
12
223
270,400
At this important juncture, building better evaluations across a diverse set of domains is the most impactful way to align models. We need subject matter experts across cyber, biology, chemistry, and many other industries to build benchmarks that help pace the frontier.
Mercor is committing $5M to a new AI Safety Fund. One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. This is why we believe investing today in safety research, evals, and verification is critical. We want to work with AI researchers who are focused on critical safety risks, including: - Misalignment: deceptive alignment, reward hacking, scheming - Sandbox escape and agent containment failures - Evaluation awareness: models that behave differently when tested - Interpretability and scalable oversight - Red-teaming methodology and safety eval design We will fund the researcher hours, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts red-teaming, grading and annotation, and free use of our evals and analysis platform. Grants are open to independent researchers, non-profits, and academics.
10
10
97
508,027