Building Transformers.js at 🤗 @huggingface // @GoogleDevExpert for #WebTechnologies and #AI

Thun, Switzerland
🖐️ No clicks, no scripts — just an AI agent driving the browser itself. @nicodotdev, Machine Learning Engineer at @huggingface, built a multimodal AI agent that runs entirely client-side and holds context across conversations. Catch the talk 👉 jsnation.us/
1
1
2
593
I'll be back in SF for Web AI Summit! A year after the v4 preview, Transformers.js went from 1.7M to 10M monthly downloads. Come see what shipped. Also, @xenovacom will be talking about agents optimizing GPU kernels at web scale. The whole WebAI world in one room: rsvp.withgoogle.com/events/w…
2
14
689
🤷 Nico Martin retweeted
GLiNER2.5-Decide now runs entirely in your browser 🎉 ONNX weights + a WebGPU demo in @nicodotdev's open-jev: 4 typed decisions in ~100 ms on a MacBook, and your text never leaves the tab. Try it: huggingface.co/spaces/shreya…
Introducing GLiNER2.5-Decide, our new 340M parameter open weight, encoder-based decision model. GLiNER2.5-Decide is built for fast, deterministic classification. The model evaluates a set of user-defined typed questions and rules, and jointly decodes their answers, returning structured decisions with probability distributions and confidence scores. We evaluated the model’s performance on Fast Decisions, an unseen, internally generated classification suite based on 17 datasets testing real-world use cases across routing, triage, classification, sentiment, and content understanding. Measured against similar decision models, GLiNER2.5-Decide leads in 9 of the 17 datasets, achieving the highest average score: - GLiNER2.5-Decide: 60.1% - SemIf: 56.4% - JevK5: 57.5% - Laya: 46.6% This performance makes the model a strong fit for use cases like tool calling, model routing, browser and computer use, and LLM-as-a-judge. GLiNER2.5-Decide’s lightweight encoder architecture makes it easy to fine-tune the model for specific tasks, while being efficient enough to run locally on consumer-grade CPUs or in air-gapped environments, giving users greater control over where their data goes and where the model runs. To make building and experimenting with GLiNER2.5-Decide as easy as possible, we're also offering hosted inference. You can now use our API to run inference and fine-tune GLiNER models on specialized tasks right inside your own coding agent: agent.fastino.ai As with previous models, we’re also releasing the model weights on @huggingface under the Apache 2.0 license: huggingface.co/fastino/GLiNE…
7
6
21
1,134
Jev is the most exciting AI release in months. So I built the open, in-browser version 🎉 open-jev: System One style decisions running 100% on your device. Text in, typed answers out, one forward pass. Nothing is generated, so it can't hallucinate.
17
21
153
19,410
Under the hood: the KEV models by @jaredpalmer , open Apache-2.0 decision models with the same contract as @typesafeai 's Jev. Not the same model, but the same shape. I converted them to ONNX and they run on your GPU via WebGPU 🔥 huggingface.co/onnx-communit…
1
6
951
A 27 BILLION parameter model in the browser! 🎉🤯
NEW: Ternary Bonsai 2 just landed on Hugging Face — a 27B reasoning model in ternary weights, built for the agentic era 🤯 9x smaller than FP16 while retaining 98.2% of the intelligence. At <6GB in size, it can even run 100% locally in your browser on WebGPU! Try it yourself 👇
1
3
15
1,244
Bonjour Paris👋 First time in Paris since I joined @huggingface 🤗
13
108
2,879
🤷 Nico Martin retweeted
New in Transformers.js v4.3.0: Structured Output, New Models, WebGPU upgrade, and a Documentation Overhaul. $ npm i @huggingface/transformers
4
4
36
2,522
Transformers.js v4.3.0 is out 🎉 And with it our first sub-package: huggingface/transformers-structured-output Force any model to return exact JSON Schema or regex output, directly in the browser. Full walkthrough 👇
1
7
48
2,210
Without constraints you never know if the model returns clean JSON, JSON in markdown fences, or XML because it felt like it. Structured output strips every token that would break your schema at decoding time. Super fast, pure JavaScript, no WASM, barely any overhead🚀. Plugs into the existing logits processor🤗.
1
2
172
🤷 Nico Martin retweeted
Should AI run in the browser or on the server? Honest answer: it depends. @nicodotdev gives the final, honest take at React Day Berlin. Reserve your spot → reactday.berlin
1
5
9,198
Are these custom @zurichjs sneakers, @farisaziz12? 🥰
4
2
21
1,230
Did someone say Whisper + WebGPU in the browser? 👀 Super happy to see Transformers.js powering on-device transcription in @remotion ❤️ timestamped words straight to captions, no server involved🚀 Nice work @JNYBGR & team 🙌
Some cool new packages we've made! @⁠remotion/whisper-webgpu: Transcribe audio in the browser @⁠remotion/gsap: Use GSAP to animate in Remotion @⁠remotion/mac-cursors: Every Mac cursor as vector graphic Happy hacking 🥳 remotion.dev/api
1
12
1,424
Had a great time talking browser AI with @grabbou on React Universe On Air. We went pretty deep: why Transformers.js sits on ONNX Runtime, what WebGPU changes, caching and download UX, agents in the browser, and a first look at the WebGPU inference engine we are building at @huggingface. Chapters below if you want to skip around🤗.
How much AI can you run in the browser today? @nicodotdev from @huggingface joins @grabbou to explain how Transformers.js brings pretrained models into JavaScript for text, audio, image, and agent workflows. The conversation covers ONNX Runtime, WebGPU, model downloads and browser caching, CPU fallbacks, performance across devices, local speech recognition, and background removal. They also discuss browser agents and Hugging Face's early WebGPU inference engine, where current experiments show 5x to 10x speedups. Chapters: 0:43 Welcome to React Universe On Air 1:04 Nico Martin's work at Hugging Face 2:04 What Transformers.js is 3:30 Browser AI beyond large language models 6:00 Bringing Python ideas into a JavaScript API 10:08 How the Transformers.js pipeline API works 14:23 Choosing compatible models and architectures 15:46 Why Transformers.js uses ONNX Runtime 19:01 CPU inference, WebGPU, and browser coverage 22:57 Model downloads, browser caching, and UX 28:28 Startup time and performance across devices 30:57 Why Transformers.js focuses on the browser 34:07 What developers misunderstand about browser AI 38:44 Building AI agents in the browser 43:19 Local fallbacks and hybrid AI 44:28 The browser AI roadmap 47:04 Hugging Face's custom WebGPU inference engine 50:57 Why speedups matter on constrained hardware 52:07 Nico's favorite browser AI use cases
2
2
9
931
Two days ago we shipped @huggingface/kernels. Now here's the deep dive. Why we publish Jinja templates instead of WGSL files, how the browser compiles the fastest kernel for your GPU, and what that unlocks: → attention in 20 lines of JS, running on the GPU → 1M+ pixels animated by a single matmul, 400 fps vs 6 fps in plain JS → Fleet: benchmark your GPU, help us tune the kernels for every device out there Full video ↓
5
4
27
1,620
I had a great chat with @grabbou from @callstackio. Going live next monday 2PM CEST. 😊 piped.video/watch?v=elAVSra9…
Background removal and transcription run in the browser. @nicodotdev joins @grabbou to cover Transformers.js today, device limits, offline use, and which parts of an agent workflow can already run locally. 📺 YouTube premiere: piped.video/elAVSra9KCY 📆 Sept 7 at 2PM CEST
1
1
4
990