We're hiring: Research Engineer / Research Scientist, Coding Capability 🧑‍💻 Build the data, post-training, and evals that push frontier agentic coding forward. What you'll do: → SFT and RL data for code generation, debugging, tool use, and long-horizon tasks → Synthetic coding data at scale, including execution-verified solutions and RL environments → Own the post-training recipe end to end, from data to benchmark → Keep evals honest by tracking contamination and leakage Who we want: a strong track record in competitions, research, trading, or open source. Any degree level. Why Apodex: real ownership, serious compute, and research published in the open. 📍 Full-time · US / Singapore 📩 simon@apodex.com @SimonShaoleiDu #Hiring #AIResearch #PostTraining #CodingAgents #ReinforcementLearning #AIforDiscovery
Made with AI
6
3
24
1,963
Scientific discovery poses significant challenges for measuring an AI system, as effective signals on discovery outcome are slow, expensive and scarce. We present TRACES, a benchmark and a methodology to evaluate discoverative AI systems based on their process, along six fundamental capabilities: T – Tool usage R – Recovery from errors A – Alternatives management C – Coherence over many steps E – Evidence loyalty S – Scope descriptions that are clear We evaluated 4 frontier open-sourced models: @Kimi_Moonshot kimi-k3, @Alibaba_Qwen qwen3.8-max, @deepseek_ai deepseek-v4.1-flash, and @Zai_org GLM-5.3. The attached figure tells a nuanced story. While qwen3.8-max has comparatively the highest TRACES process scores, kimi-k3 might be the sweet spot of balancing discoverative effort with process quality. Jev @typesafeai , launched last week, works fantastically well on the other end of the problem space: well-defined objectives, calibrated confidence, answers in milliseconds at negligible cost, so verification can run continuously across millions of decisions. In contrast, TRACES (discovertraces.ai/) targets long-horizon open-ended AI discoverative tasks with emphasis on complexity of the process and richness of process diagnostic measurements. hashtag#AIforScience hashtag#ScientificDiscovery hashtag#Benchmarking hashtag#AIResearch
2
2
14
752
Big congrats to @profshengwang , just named to @techreview 's Innovators Under 35 Asia Pacific list. 🎉 He joins fellow Apodex Lead Scientist @SimonShaoleiDu, who made the list last year. Two years, two honorees. We're pretty proud. @profshengwang builds biomedical AI that brings pathology slides, scans, and molecular data into one model. GigaPath, which he and his team built, has been downloaded more than 5 million times and is used worldwide. @SimonShaoleiDu works on the theory behind modern AI. In 2018-2019, he gave the first rigorous proof of why the standard way we train deep networks (gradient descent) can find the optimal solution. Sheng and Simon tackle very different problems, but share one conviction: AI can do more than summarize what's already known. It can help scientists form hypotheses, test ideas, and uncover genuinely new insights. AI as a partner in discovery: that's the vision our founder @tianqiao_chen set for Apodex, and it shapes everything we build. Congratulations, Sheng and Simon! Sheng's profile: coming soon Simon's profile: innovatorsunder35.com/the-li… Learn more about Apodex: apodex.com
5
1
26
1,827
Most AI benchmarks test retrieval: can a model find an answer we already know? But science’s hardest problems demand discovery: can a system earn an answer no one knows yet? Meet TRACES 🧭—the world’s first benchmark for discoverative AI: systems that can weigh evidence, test hypotheses, and reach verifiable conclusions on problems with no answer key. Proposed by our founder, @tianqiao_chen, TRACES measures six defining capabilities of discoverative intelligence. Recently, we published: · A rigorous definition of “discoverative intelligence” · A rubric that distinguishes sound investigation from lucky guesses · An open call for both problems and solvers · Website:discovertraces.ai/
5
10
111
1,069,910
We just shipped an update to Apodex: Library and Environment · Connectors. 🫱Try it now: apodex.ai Keep the files you rely on in one place, ready at the start of every chat. Connect the sources you use GitHub, research databases, market data and Apodex can find it.
10
2
35
180,700
📚 Library Ever re-uploaded the same paper or spreadsheet before every question? Not anymore. Save the files you rely on once, they're already there at the start of every new chat. No more digging through old chats for an attachment. Type @ to point Apodex at one specific file for this message. Citations in the answer link straight back to the source, click through anytime. Personal and company libraries stay strictly separate: your own files never enter company chats, while anything saved in a company chat is shared with the whole org. No folders, no share links, no auto-save. It's not a drive, just the reference shelf you actually use, so you never upload the same file twice.
6
1
9
531
🔌 Environment · Connectors What Apodex can do is bounded by the world it can reach,now you decide where that boundary is.13 built-in sources across three areas: collaboration & knowledge (GitHub), research & public data (arXiv, PubMed, Europe PMC, Crossref, ClinicalTrials.gov, bioRxiv/medRxiv, DataCite, openFDA), and market data (EODHD, Fiscal.ai — already included in your plan, bring your own key anytime). Type @ to search your Library and connected sources together, grouped by source. A per-turn selector lets you choose exactly what this inquiry can read. Every connection is read-only. Credentials are encrypted, and disconnecting revokes access instantly. For teams: admins control what's available org-wide, but never see what anyone connects or reads.
1
4
293
A generational leap, but not on every dimension. Yesterday we put the TRACES board up. Here is the first comparison worth pulling out of it. GPT-6-astra @OpenAI tops the board but that is not the interesting part. Against GPT-5.6-sol, same harness, 140 agent episodes, it gains on all six capabilities. The gains are not even, and the uneven part is the story: ➡️ Tools, selecting, calling and correctly interpreting external tools: 2.34 → 2.80 ➡️ Repair, locating and correcting its own errors once feedback arrives: 1.81 → 2.01 ➡️ Alternatives, laying out competing hypotheses and keeping or discarding them as evidence accumulates: 2.10 → 2.80 ➡️ Coherence, holding state, constraints and logic intact across a long chain of work: 2.63 → 3.15 ➡️ Evidence, grounding every conclusion in observation, data, experiment or citation: 2.26 → 2.76 ➡️ Scope, stating where a conclusion holds and where it does not: 2.20 → 2.75 Five of the six moved into the top three on the board. Repair did not. It gained 0.20 points where every other capability gained at least twice as much. Now put Repair next to Alternatives, the largest gain of the six. This generation got far better at holding several possible paths open and pruning them as evidence arrives. That matters for discovery. Finding something genuinely unknown is not just reasoning further. It is knowing how to maintain, investigate and select alternative hypothesis in a rigorously, systematic way. A single score tells you who won. TRACES tells you why, and where the frontier still breaks. Live leaderboard: discovertraces.ai/liveleader…
76
111
390
538,504
AI Capability is accelerating. Is our ability to evaluate it keeping pace? That question runs underneath a lot of this week's conversation, including @hilbertspaess 's departure from @AnthropicAI after three years of pretraining research there and at @OpenAI . We are not going to adjudicate his read on either lab. But the question his exit puts in front of everyone building these systems is a fair one, and it is ours too: the systems are getting more capable faster than the tools we use to characterize them are getting better. Most evaluation still scores the answer. Right or wrong, how often, how fast. That tells you what a system can do. It tells you almost nothing about how it got there, which is where the failures actually live. That is what we built TRACES for. It scores the work itself, across six capabilities: how a system uses tools, repairs its own errors, holds competing hypotheses open, keeps a long chain of work coherent, grounds its conclusions in evidence, and states the limits of what it found. Same harness, same tasks, scored on process rather than outcome alone. Knowing how capable a system is was never the hard part. Knowing how it gets there, and where it breaks, is. discovertraces.ai/
1
2
8
13,387
Apodex retweeted
GPT6 Astra achieved the best performance on our discovery benchmark TRACES, which not only measures outcomes, but also the problem solving process. @gdb @sama
GPT-6 Astra @OpenAI is here. Everyone's asking how smart it is. We asked a different question: can it discover the unknown? We ran it through TRACES — our benchmark for scientific discovery capability. Six process dimensions, not one score. Result: it tops our leaderboard. Real generational jump — Evidence and Coherence both up ~20-25% over GPT-5.6 Sol and GPT-5.5. Highest Coherence we've ever recorded. But the six dimensions don't move together, and that's exactly why one "how smart" number isn't enough. Some capabilities leap forward, others hold steady, in the same model. That unevenness is what a process score is built to catch. Full breakdown, live: discovertraces.ai/liveleader… #GPT6 #TRACES #AIforScience #LLM #AIBenchmark #Apodex #AGI #AIResearch #ScientificDiscovery #FoundationModels
2
5
394
Apodex 1.1 mini is now available in GGUF 🚀 Built for easier local deployment and community experimentation. GGUF → huggingface.co/apodex/Apodex… Explore all Apodex 1.1 model formats → huggingface.co/collections/a…
5
18
125
7,312
How do you test an AI on a problem nobody has solved yet? Our AI Research Scientist Brian Wang sat down with iTnews to explain TRACES, the benchmark we built for exactly that. Full conversation 👇 itnews.asia/feature/traces-p…
4
5
20
39,077
So TRACES scores the working record, not the final answer. Six capabilities: Tools, Repair, Alternatives, Coherence, Evidence, Scope. Brian's framing: less an exam, more one researcher watching another work.
2
4
214
Also in the interview: where the 20 environments came from (a 2-month scan of 561 industries), who TRACES is for, and when it becomes essential rather than nice-to-have. Full conversation 👇 itnews.asia/feature/traces-p…
2
97
GPT-6 Astra @OpenAI is here. Everyone's asking how smart it is. We asked a different question: can it discover the unknown? We ran it through TRACES — our benchmark for scientific discovery capability. Six process dimensions, not one score. Result: it tops our leaderboard. Real generational jump — Evidence and Coherence both up ~20-25% over GPT-5.6 Sol and GPT-5.5. Highest Coherence we've ever recorded. But the six dimensions don't move together, and that's exactly why one "how smart" number isn't enough. Some capabilities leap forward, others hold steady, in the same model. That unevenness is what a process score is built to catch. Full breakdown, live: discovertraces.ai/liveleader… #GPT6 #TRACES #AIforScience #LLM #AIBenchmark #Apodex #AGI #AIResearch #ScientificDiscovery #FoundationModels
62
96
306
425,437