Machine Learning Librarian @huggingface šŸ¤— I like big datasets and small models.

Scotland
Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub huggingface.co/datasets/bigl…
20
260
1,864
101,820
Daniel van Strien retweeted
The Hub can be a firehose of new releases. I used @fastinoAI's GLiNER2.5-Decide to build a Space that tags every new model, dataset and Daily Paper on @huggingface into 48 topics, usually a minute or two after it changes. Runs on one CPU, no LLM. huggingface.co/spaces/davans…
6
9
42
1,678
Daniel van Strien retweeted
Hugging Face is where much of the machine learning community publishes and finds its datasets, while DuckDB is the in-process analytical database that queries files like CSV and Parquet directly, with no server or warehouse to install or run. šŸ¤— šŸ¦† Did you know that, since #DuckDB v0.10.3 (~May 2024), you can point a SELECT at a dataset on the Hugging Face Hub⁠, using the DuckDB `hf://` protocol, and query it, without downloading it first? This blog post covers how that integration works and the use cases it fits:Ā duckdb.org/2026/09/25/huggin…
8
15
116
5,740
Can now fine-tune an even better starter model using this recipe!
Trained a Jev-style classifier on @huggingface Jobs for ~$1.50. It's a 194M GLiNER2 model that suggests task tags for any Hub dataset from its column names and first row, and returns a label with a probability. Zero-shot, GLiNER2's first suggestion matched an owner's tag 10% of the time. After 17 minutes of fine-tuning: 69%. The fine-tuned model runs on a free CPU in about a second. Owners' tags are noisy, so some "wrong" answers are tags the owner left out. The recipe is open: one hf jobs command trains the same kind of model on your own labels. The README example (book titles) runs in ~2 minutes for about $0.02. Demo: huggingface.co/spaces/davans… Recipe: huggingface.co/datasets/uv-s…
4
42
4,662
Daniel van Strien retweeted
Introducing GLiNER2.5-Decide, our new 340M parameter open weight, encoder-based decision model. GLiNER2.5-Decide is built for fast, deterministic classification. The model evaluates a set of user-defined typed questions and rules, and jointly decodes their answers, returning structured decisions with probability distributions and confidence scores. We evaluated the model’s performance on Fast Decisions, an unseen, internally generated classification suite based on 17 datasets testing real-world use cases across routing, triage, classification, sentiment, and content understanding. Measured against similar decision models, GLiNER2.5-Decide leads in 9 of the 17 datasets, achieving the highest average score: - GLiNER2.5-Decide: 60.1% - SemIf: 56.4% - JevK5: 57.5% - Laya: 46.6% This performance makes the model a strong fit for use cases like tool calling, model routing, browser and computer use, and LLM-as-a-judge. GLiNER2.5-Decide’s lightweight encoder architecture makes it easy to fine-tune the model for specific tasks, while being efficient enough to run locally on consumer-grade CPUs or in air-gapped environments, giving users greater control over where their data goes and where the model runs. To make building and experimenting with GLiNER2.5-Decide as easy as possible, we're also offering hosted inference. You can now use our API to run inference and fine-tune GLiNER models on specialized tasks right inside your own coding agent: agent.fastino.ai As with previous models, we’re also releasing the model weights on @huggingface under the Apache 2.0 license: huggingface.co/fastino/GLiNE…
84
158
1,388
154,061
Pages were read by people, then by machines. Now you can pick an open OCR model from my @huggingface uv-scripts collection and run it on HF Jobs with one command! Every maintained recipe carries its own GPU, image and token settings in the script header. Checking 17 of them end-to-end this week cost about $1.20. github.com/davanstrien/uv-sc…
2
2
9
608
Footage: public-domain films from the Prelinger Archives. A coding agent found the clips by searching second-by-second captions of 1,864 films. More on that soon.
2
146
huggignface of robotics is huggingface we've just added lerobot episodes player on all lerobot datasets
1
6
35
3,837
Daniel van Strien retweeted
huggingface_hub v1.33.0 is out šŸ¤— šŸ“¤ Large Xet uploads: new Validating progress bar 🧩 hf skills add now covers every agent (Claude Code, Codex, Cursor…) šŸ› +12 bug fixes pip install -U huggingface_hub šŸŽ¬ Video made from the release notes… and yes, it's Opus 5.5-generated
1
1
4
510
It's very cool to see all the open Jev replications, but imo people aren't focusing enough on the not-so-secret sauce (the data). I made a @huggingface collection of some datasets that can be useful for Jev training and eval. huggingface.co/collections/d…
4
8
80
3,136
I am so happy the top-trending model on the Hub is a text classifier!
1
21
996
Daniel van Strien retweeted
Kudos to @eidon_ai for open sourcing their hardware! We decided to support their finger tracker in LeRobot! github.com/huggingface/lerob… See it in action: huggingface.co/spaces/lerobo… And build your own! Find BOM (<30$), URDF & assembly guide here: github.com/nepyope/Project-H…
After 2+ years in the robotics data space, we are shutting @Eidon_AI down. The thesis was right. But the business is brutally hard. We close this chapter by open-sourcing everything we built and sharing lessons for anyone venturing into the space.
Article

Robotics Data is a Broken Business

A technical bottleneck does not necessarily imply a good business. We recently shut down @eidon_ai after 2+ years in the data space. We started in early 2024 with a gamified multimodal data collection

2
12
53
7,081
Trained a Jev-style classifier on @huggingface Jobs for ~$1.50. It's a 194M GLiNER2 model that suggests task tags for any Hub dataset from its column names and first row, and returns a label with a probability. Zero-shot, GLiNER2's first suggestion matched an owner's tag 10% of the time. After 17 minutes of fine-tuning: 69%. The fine-tuned model runs on a free CPU in about a second. Owners' tags are noisy, so some "wrong" answers are tags the owner left out. The recipe is open: one hf jobs command trains the same kind of model on your own labels. The README example (book titles) runs in ~2 minutes for about $0.02. Demo: huggingface.co/spaces/davans… Recipe: huggingface.co/datasets/uv-s…
10
19
138
21,291
Daniel van Strien retweeted
9TB of egocentric video open sourced on @huggingface @eidon_ai is winding down and decided to open-sourced their data. 1,274 hours of first-person housework video, paired with 7-point IMU arm tracking. 13,451 recordings from 27 homes. CC-BY-4.0. huggingface.co/datasets/eido…
8
52
577
47,584
Daniel van Strien retweeted
After 2+ years in the robotics data space, we are shutting @Eidon_AI down. The thesis was right. But the business is brutally hard. We close this chapter by open-sourcing everything we built and sharing lessons for anyone venturing into the space.
Article

Robotics Data is a Broken Business

A technical bottleneck does not necessarily imply a good business. We recently shut down @eidon_ai after 2+ years in the data space. We started in early 2024 with a gamified multimodal data collection

96
136
1,488
230,190
Daniel van Strien retweeted
Latest CC dumps are on HF with CDN pre-warming for warcmaxxing 🄳 commoncrawl.org/get-started
1
9
673
Daniel van Strien retweeted
We heard you like decision models…coming soon.
16
15
255
18,791
Daniel van Strien retweeted
Introducing Limite 1B - Violetto. A model for high-frequency mathematical intelligence.
64
225
1,816
282,969