RAG doesn’t require the model to memorize an enterprise’s knowledge; it retrieves the right knowledge at inference time, while embeddings and rerankers determine whether the model gets the right context in the first place.
Today we move up the stack and answer a question that appears constantly in enterprise AI.
Link to post in comments 👇
Vercel is generating $600 million in annualized revenue, up 148% from a year earlier, as coding agents become a major source of new business.
Agents now account for about half of new business, up from less than 3% at the start of the year.
Read more details: thein.fo/4eaW4bv
Born in Higgsfield. All over your feed.
Introducing Higgsfield AI Influencer.
Create your own AI influencer and bring them into any trend.
Try now with up to 5 FREE generations on Higgsfield and in the ChatGPT extension.
Powered by Genjutsu.
Credits: @JeanPhilMadame
As a JAY-Z fan all my life;
1. I caught the “keep 1 eye open like CBS” line immediately. I did NOT register that he also said “SEE BS”.
2. I didn’t know the song was about paranoia, from the 1st bar to the hook, to the last bar of the song. All about paranoia.
Funny, but this is the real lesson, in AI infra, the rep who can answer the KV cache question is the one who wins the deal.
I've been writing up the technical side of selling inference so GTM folks don't get caught at the booth
itsthedoom.substack.com/p/ba…
I’ve started thinking about quantization less as model compression and more as a capacity lever.
If you can reduce the memory footprint while preserving acceptable quality, that potentially changes how much expensive GPU infrastructure the workload requires.
link to post in comments 👇
One observations about inference economics is that some of the biggest gains aren’t necessarily about cheaper GPUs. They’re about eliminating unnecessary computation and extracting more useful work from the compute you’re already paying for.
Batching gets more useful work from each GPU; caching avoids making the GPU repeat work it has already done. Both turn serving architecture directly into inference economics.
What if you could make a model require substantially less expensive compute, but doing so might slightly change its quality?
full post in comments 👇