Common Crawl is a non-profit foundation dedicated to the Open Web.

San Francisco, CA
1
67
Common Crawl Foundation retweeted
There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗? For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training. Here’s the work between downloading those and training a model 🧵
22
94
946
111,215
Information Stewardship Forum 2026: Creating Community and Purpose Around US Government Information A report from the inaugural Information Stewardship Forum, a three-day convening hosted by the Internet Archive in San Francisco on March 18–20, 2026, that brought together 120 people working to preserve and provide access to United States government information.
1
5
348
Getting Started with Common Crawl Data on Hugging Face We made parts of our crawl archive available experimentally on a Hugging Face Storage Bucket. This blog post explains how to use the data on Hugging Face, with tools and examples on which you can build.
1
5
14
877
CommonsDB, ISCC, Spaghetti, and Meatballs Notes from the CommonsDB final conference in Alicante: a registry of 6.5 million open works keyed on ISCC, and the EUIPO's push for federated copyright infrastructure.
1
1
4
364
Bodhium Labs: From web pages to a question corpus FineQuestions draws on 110 snapshots of FineWeb, the filtered English-language corpus derived from Common Crawl, spanning 2013–2025. Extraction yielded roughly 13.4 billion raw questions filling 22 TB. Normalization and aggregation reduced these to approximately 1.5 billion unique questions.
2
3
411
Web Graph Embeddings: An Experimental Dataset Release We are publishing an experimental dataset of graph embeddings built from the Common Crawl Web Graph: a 128-dimensional vector for each of 52.9 million web hosts, learned from hyperlinks alone, along with two interactive Hugging Face Spaces.
3
2
15
624
Common Crawl Foundation retweeted
the modern web you see incl websites you browse will likely disappear or become irrelevant altogether in the next few years. they’ll feel like absolute ghost towns. when the previous era existed (pre ai) google aggregated attention & then sprayed traffic to the web based on intent. modern ai has no need to spread traffic like this, it just needs to do something through an agent to agent negotiation. this leaves websites & web interfaces completely irrelevant.. which renders the browser pretty obsolete over time, just go look at your browser history & see how top heavy it all is. this is why ai is such a giant game to win because it is the ultimate aggregation of attention / time spent. you never even have to leave like you did google! whoever controls this basically controls a large portion of the gdp. e.g. there might be a measure of success metric that is something like gdp per ai agent.
54
34
621
42,926
A Content Analysis of llms.txt Files from the July 2026 Crawl Archive We analyzed the contents of 584,107 llms.txt files from the July 2026 crawl. Two thirds are produced by a plugin, half follow the structure the specification defines, and a few files even contain prompt injections.
1
1
5
456