AI researcher mapping human–LLM capability boundaries. Building validation systems and games to explore how humans and AI interact.

United States
my website is done!😄👇 You are welcome to make suggestions. aihumanbench.com
6
9
96
2,277
If human brain's long-term memory is ~2.5 PB: • HDD: ~$60K • Enterprise NVMe: ~$350K • DDR5 ECC RAM: ~$80-100M Storing 2.5 PB is cheap. Making it instant, associative, & computational is astronomical. Good news: Your brain just appreciated in value.😀
3
8
66
899
⚡ How fast are you? 400+ ms: Refocus & retry 300–400 ms: Beginner 200–300 ms: Average 150–200 ms: Pro Gamer 120–150 ms: Elite Under 100 ms: Superhuman? Probably a false start or prediction. Test your 5-round average: AIHumanBench.com #ReactionTime #Gaming pmc.ncbi.nlm.nih.gov/article…
3
5
47
626
AI can already perform an enormous amount of computation. But more compute doesn’t automatically mean better judgment. Humans still have an important advantage in many real-world situations: knowing when something looks wrong, ambiguous, or uncertain.
3
2
34
539
Ever wondered how your brain stacks up against AI? 🧠⚡ Test your reaction speed, memory, logic, and problem-solving skills Ready to find out? 👉 aihumanbench.com #AI #HumanVsAI #BrainTest #AIHumanBench
1
19
297
Memory isn’t one ability. Different tests can measure very different things: • Short-term: brief retention • Working: hold + manipulate • Spatial: locations + relationships • Auditory: sound patterns A low score in one task doesn’t mean you have a “bad memory.” #AI #Agent
1
2
16
208
For more useful comparisons: 1. Repeat the same test under similar conditions. 2. Compare your median—not your best attempt. 3. Track each memory type separately. A spatial-memory score shouldn’t be treated as interchangeable with working memory.
1
10
135
Your reaction-time score isn’t just reflexes. Four things can change it: • Sleep and vigilance • Focus and distractions • Practice and anticipation • Display and input latency For a fairer comparison: use the same device, run 5 attempts, and compare the median—not your best score.
1
9
139
Where do humans still outperform AI? The answer changes with every model release. AIHumanBench is mapping that boundary across reaction, memory, logic, creativity, and social reasoning. We’ll publish the tests, results, and limitations—not just the conclusions.
1
6
207
Justin Smith retweeted
The first CPU built for agents is going to work at scale. @SpaceX is deploying NVIDIA Vera to accelerate the orchestration, code execution, and data processing that powers its next generation of agentic AI — keeping GPUs fed and agents acting fast. From gigawatt AI factories to orbit. One NVIDIA architecture, everywhere.
280
583
5,419
9,897,417
Nice to see modularity pushed this far — even the agent loop, sessions, storage, and UI are swappable. Curious to see how well it holds up in real-world use.👏
🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now! github.com/deepseek-ai/deeps…
4
239
Just tried it—seriously impressive.
FLUX 3 Video is #2 in the world. To celebrate, FLUX 3 Video is free to use in our playground until Sunday 16th, 11:59pm PT (link in the thread). This is just the beginning of what’s coming. Up next: 4K, video editing, and multiple images & videos as reference input.
1
5
194
There are still so many gaps in AI’s capabilities today — and not necessarily because they’re hard to solve. Some may simply not be worth the time and money for anyone to research and smooth out yet. What happens when AI becomes capable enough to close its own gaps? #ai #agi #robot
3
127
Justin Smith retweeted
Exciting news: Qwen3.8-Max by @Alibaba_Qwen is #2 in Image-to-WebDev Arena! With 1,631 pts, it’s trailing only Claude Opus 5 (Max) by 39 pts! Congrats again to @Alibaba_Qwen on this huge release!
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:github.com/qwen-code-dev-bot… - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: qwen.ai/blog?id=qwen3.8 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.8… ⚡ API: qwencloud.com/models/qwen3.8…
43
77
972
163,807
nice!
Hermes Agent v0.20.0: The Herald Release Changelog below
91
Does anyone else feel like Claude Code's usage limits are super generous now? A few months ago with Opus 4.5, I'd hit my limit by day 3 every week. Now with Opus 5, I can't even max it out. I maintain a 100k-line codebase and don't even need Fable 5. #codex #claudecode #ai
3
1
191
I prefer to describe problems precisely rather than letting it endlessly guess, and I've found that starting a new chat at the right moment works better than endlessly following up. Often, a problem gets solved quickly just by re-asking it.
1
31