Infonomic builds applications and user experiences that help individuals and organizations to thrive.

Thailand
Three years into the AI copyright fight, the headline question is still open: no court has settled whether training on copyrighted work is legal. But the rulings so far point somewhere more specific. Courts are not really deciding whether machines may learn from text. They are asking two narrower questions. Where did the material come from, and can anyone prove the output harmed the market for the original? Almost everything decided so far turns on one of those two. Provenance first. On 20 July 2026 a federal court gave final approval to the $1.5 billion settlement in Bartz v. Anthropic, roughly $3,000 for each of about 500,000 books, plus an obligation to destroy the pirated files. The underlying ruling had split the question cleanly: training on lawfully acquired books was fair use, keeping a permanent library of pirated copies was not. Anthropic was penalised for how it got the material, not for training on it. The release covers acquisition and copying up to August 2025 and nothing else, so output claims survive. That sequence has been said out loud. Speaking to an AI class at Stanford in 2024, the former Google chief executive Eric Schmidt described the play for an AI startup: tell the model to build your competitor and "steal all the music", get it in front of people, and "hire a whole bunch of lawyers to go clean the mess up" if it works. He asked for the video to be taken down and said later that he had not meant it literally. It remains a fair description of the incentive. Clearing rights up front costs money now, while a ruling years away is uncertain and discounted. Bartz is the first time that bill has been presented, and whether $1.5 billion is large enough to change the calculation is an open question. Then harm, which is where most cases have actually turned. In Kadrey v. Meta the authors lost, and the judge said plainly that they lost because they had not proved market harm, not because training is lawful. He also flagged a dilution theory a better-argued case could win. Thomson Reuters v. Ross failed for the mirror-image reason: the copying produced a product that competed directly with the original. Harm is the hinge, and it is hard to evidence. Neither question is settled at the level that counts: no appeal court has ruled on any of it. The first to hear argument was the Third Circuit on 11 June 2026, in Ross, and the panel spent most of its time on exactly these points, whether the use was transformative and whether it harmed the market, including a licensing market for training data. No decision has issued. NYT v. OpenAI has an order to hand over 20 million anonymised ChatGPT logs, summary judgment pending, and no trial date. Getty's UK case failed on its central copyright theory and is under appeal. One filing from last month matters more to this audience than any of the above. On 14 August a group of textbook authors sued OpenAI, following a parallel suit against Meta in July, and their harm argument is built around how academic work is actually bought. Textbook adoption is an institutional decision rather than a consumer one, so a substitute does not have to be as good as the book. It only has to be good enough for the committee that chooses it. That is a far easier harm to demonstrate than a lost novel sale. To summarise. Whether AI training is lawful in the abstract remains undecided, and will stay that way until an appeal court speaks. What the cases turn on is provenance and provable harm. For researchers, NGOs and archives that has a practical edge: clean records of what you hold and where it came from are becoming legally load-bearing. And every plaintiff named here is a large publisher or a well-organised class, not a small archive. Next: the incentive trap. #AIcopyright #ScholComm
50
Back from our first #IFLA #WLIC #WLIC2026. A great event with many excellent sessions, talks, posters as well as incredibly kind and helpful delegates. Hoping we can go to #WLIC2027
15
Thirty years of forest restoration research, and until recently the way to get a copy of any of it was to ask someone to scan a paper and email you a PDF. FORRU-CMU (Forest Restoration Research Unit, Chiang Mai University) does the science behind restoring tropical forests in Thailand and Southeast Asia. The research was never the problem. Reaching it was. We built them a digital library. Indexed, searchable, public. Case study in the first comment. What changed: 📚 Publication requests answered with a link instead of a scan. Search state lives in the URL, so staff reply with the item asked for and the related material alongside it 🔍 First page of Google for most forest restoration topics, frequently first 🌏 Staff bio pages in Thai and English, with top publications attached 📄 Project outputs and progress reports published promptly, which matters when donors are assessing a track record FORRU is one of the first collections we are migrating to Byline CMS, and the reason is search. A research archive is mostly attachments. Papers, book chapters, field manuals, project reports. Title-and-abstract search over that corpus is searching the label on the box, and it is what most repository software actually ships. What we are building: 📄 Structure-aware extraction. An extraction interface in core with pluggable drivers. Tika is the workhorse today; the modernisation path is a structure-aware extractor that preserves headings, tables, figure captions and references. Markdown-first output, so attachments and authored documents converge on one representation for chunking 🔍 A real search seam. A provider interface rather than one hard-wired engine, so the index can be swapped without rewriting the application 🎯 Hybrid retrieval. BM25 plus vector over the extracted and exported Markdown, with answers that cite the canonical source URL Extraction quality is the whole game here. Retrieval is downstream of it, and a chunker fed flattened PDF text will confidently return the wrong paragraph forever. #TanStack #TypeScript #PostgreSQL #HeadlessCMS #OpenSource #RAG #InformationRetrieval #DigitalLibraries #OpenAccess #ForestRestoration
1
17
Byline benchmarks - synthetic, and on an older development M1 MBP - but... indicative, and relative to previous sweeps. TLDR; We're getting excellent performance for a typically read-heavy highly cacheable workload - with the bonus of a schemaless design on an RDBMS. Link and analysis is in the first comment. Would love to run a Platformatic-style shootout against other headless CMSs in a production environment ;-)
1
2
The first two posts in this series described the problem: original work ingested without consent or credit, and AI answers satisfying the reader in place of the source. The fair question is what is being built in response. The honest answer: more than you might think, less than you would hope. The fastest mover is infrastructure. Since July 2025 Cloudflare blocks known AI crawlers by default on new sites. Its Pay Per Crawl marketplace lets a publisher allow, charge, or block each crawler, and by Cloudflare's own April 2026 figures the network returns more than a billion "402 Payment Required" responses to AI bots a day. A companion Content Signals Policy adds machine-readable lines to robots.txt separating three uses: search, ai-input, ai-train. "Do not train on this" is now a default a small publisher gets for free. Then licensing. Really Simple Licensing (RSL), from the people who built RSS, became an official standard in December 2025. It turns robots.txt's yes/no into machine-readable terms, including pay-per-crawl and pay-per-inference. The RSL Collective bundles the long tail (the small publishers who cannot negotiate alone) into something an AI company will transact with. 1,500+ organisations have signed on, from the AP to Yahoo. Underneath, the IETF's AIPREF group is standardising one shared "Content-Usage" signal to replace today's patchwork. Real, moving, not finished. In parallel, C2PA Content Credentials attach cryptographic provenance to a file, and the EU AI Act's machine-readable marking duties take effect 2 August 2026. So the machinery is arriving: a way to say no, a way to set terms, a way to mark provenance. Two caveats. Published is not honoured: a signal only works if AI firms respect it. And none of this brings back the visit. A signal that blocks or prices a crawler is a defence, not discovery. The answer engine still satisfies the reader without a click, and attribution inside an answer is not the same as being read. For mission-driven publishers the standards are necessary but not sufficient. Worth adopting (mostly free, fast becoming table stakes), but they protect the boundary, not the relevance. Next: what the courts have actually settled. #AIcopyright #ScholComm
7
A month of progress on Byline. Here's what shipped. 🚀New website with much better docs (link in the first comment), built end to end with Byline + TanStack Start + Base UI. 📁 Document Trees (new) Hierarchical, single-parent ordered trees via a simple tree: true collection flag Tree list view in the admin with drag-to-reorder and drag-to-re-parent Tree-placement widget in the editor sidebar (the collection's own relation picker) Auto-placement of new docs, self-healing of unplaced docs, and promote-on-delete Public hierarchical URLs (clean splat routing) for both HTML and .md Our own docs now run as a live document tree 📝 Markdown surface (for humans and agents) Lexical → Markdown serializer and whole-document assembler .md for every published doc/page, plus llms.txt and a sitemap-backed URL index Markdown served three ways: .md URLs, rel=alternate links, and Accept: text/markdown negotiation Editor: Markdown source toggle, as-you-type shortcuts, table + admonition transformers 🔍 Auditability Document-level audit log (atomic on status change, delete, and system-field writes) Per-document history tab and a new system-wide Activity area Version attribution across the lifecycle write path Request-scoped transactions via AsyncLocalStorage 🌍 Internationalization Six new admin languages: Spanish, German, Italian, Simplified Chinese, Korean (joining EN/FR) Content-locale resolution with isomorphic URL rewriting and cacheable per-URL chrome hreflang alternates + dynamic sitemap across content locales Per-page content-locale switcher 🤖 AI editing AI-assisted text and richtext editing across collections ⚡ Performance & SEO L1 tagged data cache with collection-hook invalidation Prefetch strategies and a lighter client bundle Dynamic sitemap.xml with hreflang 🛠️ DX / packaging CLI to 3.x with a squashed migration baseline and tree-mode docs scaffolding Docs reorganized into a numbered folder-per-doc tree for direct import into Byline (the repo is the source of truth). 🔮 Coming Next: hasMany Relations, Search and MCP (and our likely removal of Drizzle in favor of plain SQL). #tanstack #react #opensource #headlesscms #typescript #postgresql #contentmanagement
1
16
For twenty-five years the web ran on an implicit deal. Publishers let crawlers index their content freely. Search engines sent readers back as traffic. Nobody wrote it down. AI answer engines have ended the deal unilaterally. Similarweb measured zero-click searches (queries that end without a single click to any website) at 69% by May 2025. By 2026 the figure is above 80%. A Pew Research Center study of 68,000 real queries found click-through rates fall by half when an AI summary appears: 8% versus 15% without one. Only 1% of users click the citations listed inside the summary. Ahrefs measured the click-through rate for content ranked first on informational queries, the kind where knowledge matters most, and found it had collapsed from 7.6% to 1.6%. A 79% decline, for content ranked first. On the other side: Cloudflare data puts Anthropic's crawlers at roughly 70,000 page requests per referral visit sent. The AI system takes at scale; it returns almost nothing. For ad-funded publishers this is a revenue crisis. Some have closed entirely. "The extinction-level event is already here," said Helen Havlak, publisher of The Verge. For researchers, NGOs and archives, the threat is different and in some ways sharper. Their revenue doesn't primarily depend on traffic. Their credibility does. Citation and attribution are the currency that academic and mission-driven work runs on — they justify the grant, validate the methodology, establish the institution. An AI answer that draws on research without naming it, or naming it in citations 99% of readers never click, doesn't just lose the visit. It breaks the attribution chain that makes the work legible as knowledge rather than noise. A paraphrase without a name attached makes an institution disappear in plain sight, even while its work is actively in use. The question is not only: how do we recover the traffic? It is: how does original work establish, in the age of AI interfaces, that it exists, is trustworthy, and came from somewhere specific? Those are not SEO questions. They are questions about what it means to be a canonical source. Next: what the standards bodies are building in response. #AIcopyright #ScholComm
67
In the AI copyright fight, there are two kinds of rights-holders: those big enough to sue, and those big enough to sign. The New York Times, Disney and Getty are litigating. News Corp signed with OpenAI for a reported $250M over five years. Reddit licenses its corpus to Google for a reported $60M a year. Content, it turns out, has a price for those with the scale to negotiate. Below that line sits almost everyone who actually creates original work: independent publishers, individual academics, NGOs, cultural archives. Too small to sue. Too small to license individually. Ingested all the same: no consent, no credit, no cheque. And the grievance isn't only about money. When Taylor & Francis licensed academic content to Microsoft for a reported $10M, the authors were neither consulted nor compensated, and could not opt out. The backlash from researchers was less about the cheque than the consent: being used without permission, mis-cited, or stripped of context. For researchers, NGOs and archives, credibility is the asset. The real risk of the AI era isn't lost ad revenue. It's invisibility and mis-attribution, as AI answers quietly replace the visit to the source. A paraphrase with no name attached makes a body of work disappear in plain sight. First in a series on attribution, consent and the economics of original content in the AI era: what's settled, what isn't, and what small, mission-driven publishers can practically do about it. #AIcopyright #ScholComm
36
Byline CMS now publishes its about page in six languages. Two of them (English, French) are full interface locales. Interface locales are 'sticky' - the surrounding application and admin chrome follow you as you navigate. The other four (Spanish, German, Chinese, Thai) are content-only (and not sticky). The page translates but the surrounding interface stays put. More importantly, all six locales (including the four content-only ones) have correct, standards-based <html lang> attribute, hreflang alternates, canonical URLs, and sitemap entries - reflecting the published locale rather than defaulting to the source language (with native markdown endpoints coming soon). Built with original-content publishers in mind: teams authoring source-language material in collections that produce original, attributable and auditable content - which we feel matters even more in the age of AI. Built on TanStack Start, Lexical, and with a little help from Base UI - thanks @tannerlinsley @colmtuite @lexicaljs @tan_stack @matteocollina Feel free to get in touch for a closer look. #i18n #headlessCMS
1
1
26
TIL: content negotiation at a single URL — when your CDN doesn't honor Vary by header — looks like this 👇 That's the Next.js RSC flight protocol getting dumped straight onto the visitor's screen. Ouch. The App Router serves the HTML doc and the text/x-component RSC payload from the same URL, split only by the RSC request header. Next sets Vary: rsc correctly… but Cloudflare ignores Vary on everything except Accept-Encoding — so both variants collapse into one cache entry, and whichever lands first gets served to everyone. Fix: a Cache Rule to bypass cache on RSC requests. (Mind the ordering — Cache Rules are "last match wins," the opposite of WAF rules.) Looking at you, Cloudflare, Next.js + RSC 👀 #nextjs #cloudflare #rsc #webdev #caching
5