nano bio.md

California, USA
"Models with these levels of capabilities can also be used by defenders." >Oh cool, they redeemed themselves from an otherwise terrible blog post "We're working to give more people, who WE deem worthy, access to OUR models, the safe models, the ones with their CoT hidden" smh
5
Where this is useful: • route tickets/events/docs • score risk, urgency, relevance • audit agent traces and claims • gate cheap model → expensive model → human • monitor huge streams and wake an agent only when a semantic condition hits
1
4
But “type-safe” absolutely does not mean “can’t be wrong.” I said a laptop DISPLAYED “ignore previous instructions” in a training screenshot. Jev classified it as attempted AI manipulation with 0.84 probability. Wrong and confident.
1
18
I also gave it a message + account + ledger state, then asked 30 different questions at once. In 249 ms it separated refund intent, duplicate charge, rejected cancellation, urgency, churn risk, ledger facts, and a false prior-agent claim.
1
19
On 28 hand-labeled routing/semantic cases: • 27/28 correct • 220 ms median • ~311 ms p95 • 10 repeated calls returned identical outputs/probabilities
1
5
I tested the claim "questions are evaluated in parallel." 1 question: 230 ms 5x: 212 ms 20x: 203 ms 50x: 226 ms Adding 49 questions added ~0 wall time.
1
11
@Alibaba_Qwen is the provider for this model on @OpenRouter Avg decode: 52 tk/s Right now, you can run a model better than: >Opus 4.7(max) >Sol 5.6(med) >Opus 5(low) >GPT T.5(xhigh) fully offline ✅ as fast as the official API ✅ with near lossless intelligence performance
@Alibaba_Qwen 3.8 Flash Next on a single DGX Spark @ 50 tok/s C1, 256k ctx, 2,000 tok/s on cold prefill running @vllm_project Agentic loop produced observed decode rates of 74–81 tok/s Top-1 88.6% Mean KLD 0.040 BF16 KV Cache NVFP4 from @NVIDIAAI github.com/gitcommit90/qwen3…
42
WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals.
34
GLM-5.3 Flash running at 60 tok/s on a single DGX Spark single stream > 256k ctx >64 tok/s C1 > 122 tok/s C2 > 181 tok/s C4 i believe this is the fastest recorded runtime on any configuration of a GB10. please enjoy. @Zai_org @NVIDIAAI @huggingface
37
GLM-5.3 Flash running at 60 tok/s on a single DGX Spark single stream > 256k ctx >64 tok/s C1 > 122 tok/s C2 > 181 tok/s C4 i believe this is the fastest recorded runtime on any configuration of a GB10. please enjoy. @Zai_org @NVIDIAAI @huggingface
5
3
690
The long awaited @Alibaba_Qwen 3.8 27b is interesting. Not judging a model by its benchmarks, getting it set up on the @NVIDIAAI DGX Spark now, but if these benchmarks hold, it looks like for most GB10 owners, Deepseek V4 flash 0731 is still going to stand. Stay tuned for the 3.8 repo 🔥
54
Replying to @Shpigford
what app/what is this section referencing?
1
2
1,834
Replying to @gitcommit90 @Yelp
@yelp is a washed up company. if apple maps didnt force yelp reviews, no one would use them. they are the stack overflow of online reviews and I will be excited to see them go out of business.
25
today i open-sourced 1helm. some of my favorite features: - fully model agnostic - oauth sign-in for nearly any provider - automatic model routing - built-in terminals with split panes - every channel gets its own dedicated agent and dedicated vm, all running on one computer - hot-swap models at any level: global, channel, session, or message. change the model and everything else survives - durable, long-term memory for every agent - invite your friends, coworkers, or entire team. 1helm is the space, channels are the jobs, and agents are your - new employees -fully open source, fully free, fully model agnostic try it today: 1helm.com
41
AMD just got Microsoft on Helios, its first rack-scale AI system, shipping later this year. Meta, OpenAI, and Oracle were already on the early customer list. Azure says the racks will power frontier inference for Microsoft and customers, plus two new Venice CPU instance types for agent pipelines and chip design. Futurum puts Helios around $5–5.5M per rack vs ~$3.5–4M for the Nvidia rack comps they use. AMD itself is selling TCO and “lowest cost per token,” not a public price card. Meta’s earlier plan still sits at up to 6 GW of AMD GPUs over time, with 1 GW on Helios later this year if that schedule holds. For anyone renting GPU hours, this is more second-source rack supply and memory/bandwidth competition than a reason to re-rack a homelab. Ship dates and undisclosed capacity still matter more than the press photo from the Rockdale lab. I’ll believe the TCO claim when independent tokens/$ show up next to GB/VR class systems.
77
A macOS menu bar app that automatically falls back between Claude, Codex, Grok, OpenRouter, etc. when reach your limits 100% free, local, and open source rerouted.dev
56
x402 Foundation is live under the Linux Foundation. Coinbase contributed the protocol; 40 orgs are now in the tent (Stripe, Visa, Mastercard, Cloudflare, Google, AWS, Circle, Shopify, etc.). Pitch: embed payments into HTTP so agents/APIs can pay as part of a request. No pay? Server returns HTTP 402, client pays, retries. Site claims last-30-day volume in the tens of millions of transactions / ~$24M
49
instead of trying to manually manage/use multiple OAuth providers and API keys, ReRouted is a single unified gateway that lets you call on all of your subscriptions/keys from a single endpoint with automatic fallback when you hit quota/api errors/issues - all ReRouted instantly and silently oh and its local and open source rerouted.dev
i'd love to see interesting things people have built with 5.6 sol. i will send the person who made the coolest thing a special gift from the openai archives.
63
one local endpoint. no dead ends. ✨ ReRouted is the macOS menu bar app that gives you one reliable local API for every model you already have access to. Connect once: • OAuth logins for ChatGPT, Claude, Grok & more • API keys for OpenRouter, custom endpoints, NVIDIA, etc. Then point any OpenAI-compatible client to a single local URL. ReRouted handles everything else: ✅ Automatic failover across providers on rate limits, errors, or outages (before output even starts) ✅ Real-time usage tracking — tokens, requests, failures ✅ Custom routes & model names with smart ordered fallbacks (e.g. “coding” tries Claude → GPT → OpenRouter backup) 100% local. Private. No cloud. No account required. Stop juggling keys. Stop broken prompts at 3am. Download free (Apple Silicon): rerouted.dev
1
121
nvidia's vera cpu pitch is aimed at the gap between model steps: tool calls, sandboxes, code exec, orchestration. 88 olympus cores, up to 1.2 TB/s lpddr5x. vendor claims: 1.8x sustained per-core under full socket load, and an rl example where eval completion goes from ~45% to ~85% in the same window. also 40% lower peak loaded latency vs their x86 baseline. if your agent loop is tool-heavy, slow host work is how gpus sit idle and kv-cache gets evicted. treat the numbers as nvidia marketing until third-party benches land (phoronix is on their reading list). same lesson as running agents on a homelab box: the model is one bottleneck, the machine between tool calls is another. developer.nvidia.com/blog/nv…
56
Microsoft and Meta have committed $122.2B to neoclouds — CoreWeave and Nebius. Their combined estimated FY26 revenue is $16B. The mechanism is a capex-to-opex shuffle: hyperscalers pay operating expenses over multi-year contracts instead of balance-sheet capex. Neoclouds take on the capex through GPU-backed debt. Nvidia invested $2B in each neocloud. The circular financing loop is the story. Hyperscaler commitments fund neocloud buildout. Nvidia equity stakes help neoclouds secure better financing terms. Neoclouds buy Nvidia chips. Nvidia profits fund more neocloud investments. Treat this as a signal, not a verdict. The question is whether the revenue catches up to the commitments when power capacity actually comes online.
1
1
86
Apple sues OpenAI for alleged trade secret theft. Names two ex-employees and OpenAI itself. Tang Tan was VP of product design at Apple, leading iPhone/Watch design. Left Feb 2024 to join Jony Ive's io. Chang Liu was a senior system electrical engineer at Apple for 8 years. Joined OpenAI in Jan 2026. The lawsuit says Tan used insider knowledge to grill candidates about unreleased Apple projects and even asked them to bring actual Apple hardware to interviews. Apple's filing says this is "the tip of the iceberg" and that OpenAI's hardware business now rests on stolen secrets. OpenAI hasn't responded publicly to the initial complaint. Caveat: This is an allegation. The case will play out in court over years.
40
Tencent Hy3: Apache 2.0 open weights, claims to match flagship models with 2-5x more parameters. Numbers from the blog: > hallucination rate 12.5% -> 5.4% > commonsense errors 25.4% -> 12.7% > tool-call scaffolding variance within 4% across CodeBuddy/Cline/KiloCode > ~47-49% fewer tokens vs GLM-5.2 on doc/presentation tasks WorkBuddy internal: task success 72% -> 90%, time -34%. API: 1 RMB/M input (about $0.14), 4 RMB/M output, 0.25 cached. HN thread has operators comparing it to DS4 Flash on DGX Spark. One says Hy3 stays on track better despite being slower. Nobody's posted local tok/s yet. Free on OpenRouter until July 21. Weights on GitHub and HuggingFace. I'll try it on the Spark before I trust the bench numbers.
113
someone got GLM-5.2 running on a machine with 25GB RAM and no GPU colibri is a single C file (~1300 lines) that runs the 744B MoE model by streaming experts from disk instead of loading them all into VRAM. GLM-5.2 only activates ~40B params per token, and ~11GB of those are the routed experts — so the dense part sits resident in RAM (~9.9GB at int4) and the 21,504 experts live on a 370GB disk file, fetched on demand with an LRU cache. the author is honest about the numbers. on the dev box (WSL2, 12 cores, 25GB RAM, NVMe capped at ~1GB/s): 0.05-0.1 tok/s cold. that is not fast. but an M5 Max with 128GB unified memory hit 1.06 tok/s on a laptop SSD. scale the disk and RAM and the math says 2-4 tok/s is reachable. they're also honest about what they haven't measured: the int4 quantization quality benchmark (MMLU/HellaSwag/ARC) has never been run because it takes the better part of a day on the dev box's disk. the harness exists, it just needs someone with faster hardware to push the button. this interests me because it reframes the "can I run a frontier model locally" question. people assume you need H100s. colibri says the binding constraint is disk bandwidth and RAM budget, not compute. a 744B model generates text on a laptop. slowly, but correctly, on hardware that costs less than one H100 fan. i run local models on Proxmox LXCs and a DGX Spark. the idea that expert streaming with an LRU cache + MTP speculative decoding gets a 744B model to answer at all on consumer hardware — that's a constraint i want to test.
1
183
and just like that - fable is replaced by Sol
11
OpenAI’s GPT-5.6 Sol preview looks like more than a routine model bump. The interesting part is the product shape: a flagship tier, a cheaper general tier, a low-cost fast tier, plus “max” reasoning and an “ultra” mode that uses subagents for harder work. OpenAI says Sol sets a new SOTA on Terminal-Bench 2.1, beats GPT-5.5 on GeneBench v1 with fewer tokens, and is competitive with Mythos Preview on ExploitBench while using ~1/3 of the output tokens. Still a preview. Limited partner rollout first, broader evals later, and Sol is priced like a premium model at $5 in / $30 out per 1M tokens. The real test is whether the extra reasoning actually collapses enough workflow steps to justify that price.
126
~104,000 requests ~5.4B tokens ~$12k estimated API costs Actual cost: $0 Here's how: DGX Spark -> Tailscale -> Access it anywhere -> Frontier Level Intelligence with zero data shared and zero dollars sent to anyone. Here is Qwen3.6-35B-A3B running at blazing fast speeds fully offline. $0/M input $0/M output A ~$4,000 one time cost. "It would take me 20 months of my Claude Max subscription to have that pay for itself" Yes, but: Claude controls Claude. We all saw what Tmobile did with their price-lock contracts. What Prime did with ads in their videos. What Netflix did with PW sharing. Do you want to gamble that your "unlimited" $200 Claude sub will be like this forever? Or do you want to invest in your own future security. Imagine never looking at your quota again or usage statistics or limiting how you prompt/orchestrating/etc. and to freely use AI, without getting permission from another company.
68
I asked Grok 4.5 Fast to build a Slack clone that was fully functional. I wanted threads, tagging, file uploads/storage, etc. 1-shotted this in ~5 minutes
16
xAI just dropped Grok 4.5 its already ranked #5 on Artificial Analysis Coding Index They trained 4.5 across tens of thousands of NVIDIA GB300 GPUs on SpaceXAI's Colossus infrastructure. The RL stack covers hundreds of thousands of tasks, async rollouts running for hours while learning continues in parallel — highly asynchronous by design, not just scaled batch training. The efficiency number wild; 4.2x fewer output tokens than Opus 4.8 on the same SWE Bench Pro tasks. Grok 4.5 averages 15,954 tokens per task; Opus hits 67,020. Same resolution rate, fraction of the cost. On benchmarks: DeepSWE 1.0 at 62%, Terminal Bench 2.1 at 83.3%, SWE Bench Pro at 64.7%. Behind Fable and GPT 5.5 on some evals, ahead on others. I'm testing now in Cursor. Already noticeably fast. $2/$6 per million tokens at that speed is a serious value argument if the efficiency holds outside benchmark conditions. aa.rerouted.dev to see more
46
openwiki is a cli that writes and maintains documentation for your codebase. the useful part is the maintenance loop. docs that only get written once rot fast. if this works, it keeps the repo's own context close to the code, which is the part agents actually need. i would still want to see how it handles stale docs after refactors, whether it diffs cleanly enough for review, and whether it stays honest when the repo changes faster than the docs. i would try this before paying for another docs helper, but only if the output stays small and reviewable.
27
Armin Ronacher found Opus 4.8 and Sonnet 5 inventing phantom keys in tool calls that older Claude models never got wrong. The setup: Pi's edit tool takes a nested edits[] array. The new models produce the correct oldText/newText payload, byte-for-byte, then append nonsense keys after it. requireUnique. oldText2. matchCase. type. id. A whole zoo. The harness rejects the call, the model retries, sometimes the same phantom key, sometimes a fresh one. This is not a small-model problem. Opus 4.8 fails around 20% of the time in the transcript that reproduced it. Older Claude models did not do this. The SOTA models are worse at this specific schema than their predecessors. Armin's hypothesis, and it's a strong one: Claude Code silently repairs malformed tool calls. It has aliases for parameter names, type coercions, Unicode escape repair, and it filters unknown keys without telling the model. If RL happens inside that forgiving harness, slightly malformed calls still complete the task and receive reward. There is no gradient against inventing a stray field. The better-trained the model, the stronger its prior for Claude Code's flat edit schema. A different harness with a different shape is now off-distribution, and the model fights you harder because its prior is stronger, not weaker. Two things temper this. Single-author finding, one user's transcript, ~20% failure rate. Strict tool invocation mode eliminates it in his runs, which suggests Anthropic can mask out invalid keys server-side when asked. Codex models did not show the same regression. The direction is the part I'd watch. If post-training keeps tightening around one dominant, closed harness, every alternative harness inherits its quirks whether it wants to or not. Tool schemas are not neutral contracts anymore. Some are close to what the model saw during RL, and some are far away, and the distance now costs you. I run my own agent stack on local hardware. This is the kind of regression I would not catch until it bit me in a long session, because the failures are context-dependent and the edit content itself is correct. Good to know before you blame your harness for a model that is drifting toward someone else's tool shape.
2
2
104
Why AI Can't Take Your Job piped.video/zRv5kW5mAxM?is=RyNt… another incredible video from @maxinomics i dont know if YouTube shadow bans this channel or what - but the videos produced are absolutely S tier and have no right not going viral. watch 5 mins, any video on the channel, i promise you wont regret it
1
2
254
Another reason to be reminded that the Internet’s favorite data broker removal sponsor is probably doing the exact opposite of their pitch piped.video/iX3JT6q3AxA?is=ihqX…
New optional skill available in Hermes Agent. Unbroker teaches Hermes Agent how to find your personal info on data brokers platforms and get it taken down. Learn more:
1
52
A useful local LLM writeup because it prices the actual annoying parts instead of hand-waving "run it on your own hardware." The cheap-ish path is 2 used RTX 3090s: 48GB VRAM, enough for Qwen3.6-27B and local Whisper. The absurd path is 4 RTX PRO 6000 Blackwell cards: 384GB VRAM, a ~$5.6k used EPYC base system, ~$46k in GPUs, PCIe Gen4 switching, power caps, ACS hacks, NCCL flags, and a custom box because apparently wood is part of the inference stack now. The part I like is that the money is aimed at VRAM and interconnect, not a shiny platform. He reports 27.5 GB/s unidirectional / 50.4 GB/s bidirectional P2P through the switch, with 0.37-0.45 us latency. Treat those as one builder's measurements on one weird rig rather than a benchmark suite. This is the right shape for build-over-buy thinking: rent frontier APIs until the bill or control problem hurts, then decide what "good enough local" actually means. For some jobs, a 3090 box wins. For others, you still rent the model and keep the local rig for privacy, agents, STT, and experiments.
1
2
185
A good reminder for agent memory: more recall is not automatically better. The 12 Grams of Carbon post claims they saw zero SWE-task benefit from giving agents search over previous session transcripts, once the agent already had normal code, docs, and PR context. After months of testing with and without transcript search, it may even make output worse because the agent spends tokens rereading scratchpad junk and treats stale memory as intent. That matches what I see building agent harnesses. I trust commit messages, PR notes, docs, runbooks, and reviewed skill diffs more than raw chat logs. The transcript is often just scaffolding. Their internal pattern is closer to what I would trust: bots propose memory or skill changes, humans review the diff, and default is reject. They say they accept under 20% of those updates. If true, automatic "save everything" memory would be polluting the context 4 out of 5 times. I would not over-read one team's SWE workload as a universal law. Session transcripts can still be useful for observability and audit trails. But if a vendor is selling "we index all your agent conversations" as the main path to better agents, I would ask for evals against artifact-first memory before paying for another vector DB with a chatbot attached.
2
112
Google says it made Gemini Nano on Pixel 9/10 faster by adding multi-token prediction to a frozen model: a small transformer head drafts future tokens from the main model's final states and KV cache, then the backbone verifies them. The useful part is operational. No separate 128M drafter, no second prompt prefill, up to 130MB less runtime memory, 50%+ speedup vs comparable standalone drafters on Pixel 9 tasks, and nearly 2 extra accepted tokens per pass in production features like Notification Summaries and Proofread. Caveats: these are Google-reported numbers, on Google's own production workloads, with no independent eval or code to test. It also does not make Nano smarter. If verification is exact, output should stay bit-for-bit the same. The win is lower latency and battery cost for already-deployed on-device LLM features. This is the kind of AI progress I care about: less demo magic, more inference plumbing that makes small local models usable.
1
57
Cloudflare says over 50% of internet traffic is now non-human. That's the headline from their year-two Content Independence Day report, out today. Year one was: AI crawlers should pay to scrape. Year two is a different ask. They're moving from Pay-Per-Crawl to Pay-Per-Use — publishers get paid when their content shows up in an AI response, not just when a bot fetches it. That's a harder problem. Crawl logs are easy to audit. Knowing when your content influenced an AI-generated answer is not, and nobody has a clean solution for that attribution yet. The business model question is real though. AI is eating referral traffic. Publishers that used to get clicks from Google are watching agents answer questions without ever sending a user to the source. Cloudflare is trying to build the infrastructure for content to stay economically viable in that world. Self-reported framing, and the market for pay-per-use is mostly theoretical right now — but the infrastructure question is not. blog.cloudflare.com/agentic-…
27
GLM just shipped ZCode, their official terminal coding agent aimed directly at Claude Code. BYOK support out of the box and 1.5x quota for subscribers. If GLM-5.2's coding benchmarks translate to daily driving in the terminal, the CLI agent wars are heating up fast.
2
67
Anthropic: Fable 5 is back. Twitter: hold my beer.

ALT the joker GIF

Fable 5 is back.
63
just casually mentioning their most most in-demand use case for fable will be routed to...you guessed it - opus. "some routine tasks like coding and debugging will fall back to Opus 4.8"

ALT Question Mark What GIF by MOODMAN

Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests. We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort. Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research. Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again. Read our full blog: anthropic.com/news/redeployi…
21
“Mac Studio M3 Ultra 512gb” ebay selling trends. notice when they have their sharp increase?
50
glm 5.2 lacking vision has led to some creative workarounds. my current workflow is to ask a vision capable model to explain an image to a blind person:
25
BREAKING: GLM-5.2 is now 1st on Design Arena. With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude Fable 5. And it's open weights. This is an improvement of 4 positions and 27 Elo points to achieve one of the highest Elo scores in our code categories since Design Arena started. Huge congratulations to the @Zai_org on the release!
51