Embeddings, OCR, reranking and generation don't necessarily run best on the same inference runtime. SIE uses multiple compute engines, including PyTorch, Flash Attention and SGLang, and selects one per model. The application sees one inference layer. Interested?
2
4
66
Keep the gateway thin: inspect the request, add the metadata workers need, enqueue it. When requests contain images or video, repeated serialization and copying can become expensive. The shared entry point shouldn't turn into the next bottleneck.
53
Maximum throughput can be a trap. Past the knee of the throughput - latency curve, more traffic may add little throughput while latency climbs fast. The useful benchmark is where you'd actually operate, not how hard you can push before it breaks...
41
Stop manually optimizing inference. Let an agent bash its head against the benchmark to rewrite forward passes. Faster tokens, less compute. Grab the SIE repo and try it yourself: github.com/superlinked/sie #AI #MachineLearning #LLM
1
1
85
Stop burning money on massive model training. You can fine-tune 235B parameter models using LORAs by adapting just 5% of the parameters. It is affordable and accessible now. Check the case studies at github.com/superlinked/sie. #AI #LLM #FineTuning #DataScience
1
52
Demo Bay is lovely, but the road to production is why we built SIE.
36
Move from one large model to a fleet of smaller ones and the problem changes: many models, many short requests, uneven demand. If the serving layer can't keep GPUs busy and models available, cheaper inference on paper may not become cheaper inference in production. What's your take on it?
1
78
A small model doesn't need to match a frontier model at everything. It needs to clear your quality bar on the task you give it, at a latency and cost you can live with. The useful test is on your data, inside your workflow.
2
96
Stop waiting for H100s to boot. If you use attached SSD snapshots instead of cloud storage, you bypass the bottleneck entirely. It loads in seconds. Check the SIE repo here: github.com/superlinked/sie #AI #GPU #H100 #MachineLearning
1
3
90
An AI agent rarely has one job. A contract review agent might OCR a scan, retrieve clauses, compare them, extract fields and generate an answer. One agent. Five very different inference tasks. Why send them all to the same model?
64
Superlinked retweeted
just got access to Claude Fable 5.5 👀 that was the bait for We Steal Your Tokens, which just won first place at the @AITinkerers x @superlinked Hacking Open Source Projects hackathon it scrapes the email of open-source contributors, uses qwen to write definitely-not-phishing emails from totally-not-anthropic, hides Cyrillic characters in a fake install URL, then waits for someone to paste the very secure curl | bash command into their terminal. And just like that, I still their codex/claude auth tokens... and all their usage credits don’t blindly paste links from emails into your terminal unless you want to sponsor someone else’s (my) inference thanks to @tokengobbler for organising
2
5
160
Running 160+ AI models is a dependency nightmare. Most teams struggle with slow loading and conflicting environments. The fix is a REST wrapper that standardizes every inference backend into simple YAML configs. Check out the SIE repo for the full architecture: github.c...
47
The traditional NER lifecycle: define entity types, annotate thousands of examples, fine-tune, deploy. Then someone asks you to also extract "recoupment amounts" and you're back at step one. Zero-shot NER breaks the loop. GLiNER takes the label set as part of the request, so what you extract becomes a runtime parameter instead of a training decision. One model, four domains in this example: SEC filings, CMS healthcare docs, NTSB safety reports, Supreme Court opinions. Four label sets, zero retraining. And every extracted span is validated character-for-character against the source text, so each entity carries proof it actually exists in the document. If GLiNER holds up on your text the way it holds up here, you just deleted a fine-tuning pipeline from your roadmap: superlinked.com/blog/named-e…
2
118
Small models are getting good enough that the bottleneck is shifting. The question is becoming less "can an open model do this?" and more "how do I serve 10, 20 or 50 specialised models without turning my infrastructure into a mess?" An agent might use one model for OCR, another for embeddings, another for reranking, another for extraction, another for guardrails and another for generation. That looks very different from serving one giant LLM. Traditional inference infrastructure is generally built around routing requests to a small number of large model workers. With lots of fast, small-model requests, that approach can leave GPUs sitting idle while the router tries to keep up. In our latest write-up, @Svonava breaks down the architecture we ended up with at Superlinked: Shared queues instead of top-down routing. Workers forming their own batches. Multiple models packed onto the same GPU. Models loaded and evicted as demand changes. Different runtimes hidden behind one inference layer. In our tested setup, moving to centralised queueing and worker-side batching roughly doubled cluster throughput. The models are getting smaller. The serving problem isn't. Read the full article and watch the talk below: superlinked.com/blog/serving…
1
1
2
133
Open source AI models are useless if you need a warehouse of GPUs to run them. Real accessibility means smaller footprints and smarter infrastructure handling. Stop over-engineering. Check the SIE repo: github.com/superlinked/sie #AI #OpenSource #Tech
45
Teams should not have to rebuild the production inference layer around every model, runtime and deployment. That is why we are building SIE: one production path for small, specialized models, whether you run them yourself or use a managed service.
49
Tired of endless thinking tokens? Stop paying for compute you don't need. Fine-tune your models to jump straight to the answer using QLoRA. We crushed the benchmarks for cheap. Check the SIE repo here: github.com/superlinked/sie #AI #MachineLearning #LLM
2
55
We analyzed dozens of Reddit discussions and found 40 recurring problems with running open models in real products. The biggest lesson? The model was rarely the hardest part. Keeping it fast, stable, measurable and private under real traffic was.
59
Switching models like Gemma? Don't settle for default speeds. Proper parameter tuning for speculative decoders bumped performance from 30 to 119 tokens/sec on an A100. Stop guessing configurations. Optimize your inference stack here: github.com/superlinked/sie #LLM #AI #MachineL...
2
57
What happens when your model starts reading values from the past? @f_makraduli ran into exactly that while implementing FlashNorm. The idea behind FlashNorm is surprisingly simple: RMSNorm doesn't do much arithmetic, but it can still take meaningful wall-clock time because of kernel launches, memory movement and sequential dependencies. So instead of treating normalization and matrix multiplication as a fixed sequence, FlashNorm rearranges the work. - Fold the normalization weights into the following projection. - Defer the scalar normalization. - Run the RMS calculation and matrix multiplication in parallel using CUDA streams. - The maths works. The first implementation appeared to work too. - Unit tests passed. Perplexity looked normal. - Then longer generations started repeating tokens with a one-step lag. The problem was a race condition. One CUDA stream hadn't finished writing before the next operation read the buffer, so the model was occasionally consuming stale values. Fixing it required explicit synchronization between the streams. We wrote up the full story, including how FlashNorm works, where the speedup actually comes from, and why the reported 33 to 35% improvement applies to the norm-plus-projection operation rather than the entire LLM. Read the article and watch Filip's talk below. superlinked.com/blog/flashno…
1
74