You can query the entire internet, 100+ billion (!) rows, with
@duckdb in under a minute. Yes, it's crazy, but you can basically select * the internet. You might be familiar with the Wayback Machine from
@internetarchive. The
@CommonCrawl is a similar project that also has all its data available in a S3 bucket.
@dumkydewilde was interested to see how the 'vibe-coded' web has grown over the last few years. While
@Lovable ,
@Vercel and
@Cloudflare Pages have taken off tremendously, they still pale compared to the traditional Wordpress-Blogspot hegemony. See for yourself in the interactive Dive visualization, or read the full blog to do it yourself 👇.
- Dive:
motherduck.com/dive-gallery/…
- Blog:
motherduck.com/blog/querying…