There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗?
For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training.
Here’s the work between downloading those and training a model 🧵
ALT Flow diagram of the Marin data pipeline. Starting with 25.25 trillion tokens, deduplication removes 2.13 trillion and benchmark decontamination removes another 13.66 billion, leaving 23.11 trillion. Documents are organized into 40 topics and five quality bands, yielding 200 sampling groups. Results from 872 small-model experiments feed a regression used to choose sampling weights.
Information Stewardship Forum 2026: Creating Community and Purpose Around US Government Information
A report from the inaugural Information Stewardship Forum, a three-day convening hosted by the Internet Archive in San Francisco on March 18–20, 2026, that brought together 120 people working to preserve and provide access to United States government information.
September 2026 Crawl Archive Now Available
We are pleased to announce the release of the September 202 6 crawl archive, consisting of 2.17 billion pages (or 361.4 TiB of uncompressed content.)
📷
Getting Started with Common Crawl Data on Hugging Face
We made parts of our crawl archive available experimentally on a Hugging Face Storage Bucket. This blog post explains how to use the data on Hugging Face, with tools and examples on which you can build.
CommonsDB, ISCC, Spaghetti, and Meatballs
Notes from the CommonsDB final conference in Alicante: a registry of 6.5 million open works keyed on ISCC, and the EUIPO's push for federated copyright infrastructure.
Bodhium Labs: From web pages to a question corpus
FineQuestions draws on 110 snapshots of FineWeb, the filtered English-language corpus derived from Common Crawl, spanning 2013–2025. Extraction yielded roughly 13.4 billion raw questions filling 22 TB. Normalization and aggregation reduced these to approximately 1.5 billion unique questions.
Web Graph Embeddings: An Experimental Dataset Release
We are publishing an experimental dataset of graph embeddings built from the Common Crawl Web Graph: a 128-dimensional vector for each of 52.9 million web hosts, learned from hyperlinks alone, along with two interactive Hugging Face Spaces.
the modern web you see incl websites you browse will likely disappear or become irrelevant altogether in the next few years. they’ll feel like absolute ghost towns.
when the previous era existed (pre ai) google aggregated attention & then sprayed traffic to the web based on intent.
modern ai has no need to spread traffic like this, it just needs to do something through an agent to agent negotiation. this leaves websites & web interfaces completely irrelevant.. which renders the browser pretty obsolete over time, just go look at your browser history & see how top heavy it all is.
this is why ai is such a giant game to win because it is the ultimate aggregation of attention / time spent. you never even have to leave like you did google! whoever controls this basically controls a large portion of the gdp. e.g. there might be a measure of success metric that is something like gdp per ai agent.
A Content Analysis of llms.txt Files from the July 2026 Crawl Archive
We analyzed the contents of 584,107 llms.txt files from the July 2026 crawl. Two thirds are produced by a plugin, half follow the structure the specification defines, and a few files even contain prompt injections.