Filter
Exclude
Time range
-
Minimum likes
GCC/MENA inference teams: TaSQ compresses the KV cache to 1 bit and reports up to 14x larger batches and 1.87x higher peak throughput vs BF16 (SGLang, one RTX 6000 Ada). Bigger batches are not proportional speed, so benchmark. eucax.ai/insights/tasq-and-t… #LLMInference #KVCache
22
Chunked #KVcache compression introduces dangerous periodic blind spots into #transformer architectures, reveals new research highlighted by @HuggingPapers. Models compressing #token histories at fixed strides reliably retrieve #context at certain boundaries but fail ⚠️catastrophically at intermediate offsets. This phase-dependent #degradation can quietly flip code completions or corrupt reasoning chains without triggering obvious #runtime warnings. Engineers deploying long-context #compression must 🤔re-evaluate fixed-stride strategies to prevent subtle inference failures in mission-critical #GenerativeAI systems. 🔬 x.com/HuggingPapers/status/2…
LLMs with chunked KV-cache compression have hidden periodic weak spots Models that compress past tokens into cache entries at a fixed stride can retrieve information well at some positions but terribly at others, even flipping answers on a simple code completion. This phase sensitivity was discovered by ByteDance Seed on DeepSeek-V4 and V4.1.
2
19
How much of your inference workload is 𝒂𝒄𝒕𝒖𝒂𝒍𝒍𝒚 new? Tensormesh's #KVcache Hit Rate shows how often cached state is reused instead of recomputed, giving teams visibility into how much repeated work their models are doing. You can't 𝑜𝑝𝑡𝑖𝑚𝑖𝑧𝑒 𝑟𝑒𝑝𝑒𝑎𝑡𝑒𝑑 𝑤𝑜𝑟𝑘 if you can't see it.
1
1
60
Un #LLM escribe palabra a palabra y cada una tiene que «mirar» todo lo anterior. La #KVCache evita recalcularlo: guarda la K y la V de cada token y las reutiliza. Sin caché, cada paso recalcula más tokens; con caché, uno. El precio: ~128 KiB de memoria/token en un 8B en FP16.
25
Why KV cache makes LLM generation faster #KVCache #LLM #AIInference #Transformers #GPUServing
1
16
GCC/MENA inference teams: SparseEngine makes long-context serving a state-management problem. Reports over 10x throughput with KV eviction and over 2.5x faster decoding vs vLLM at matched concurrency. eucax.ai/insights/sparseengi… #LLMInference #KVCache #AIInfrastructure
54
That's a wrap on #TheAIConference in SF! Thanks to everyone who stopped by to talk faster, more efficient AI inference. And yes, we met a robot. Say hi to Gerald! 🤖 #AIInfrastructure #KVCache @aiconference
1
1
8
237
GCC/MENA inference teams: PulseInfer treats sparse KV offloading as an I/O scheduling problem for long-context decoding, reporting up to 4.7x decode throughput vs SGLang without treating host memory as a free lunch. eucax.ai/insights/pulseinfer… #LLMInference #KVCache #AIInfrastructure
20
Andrew Gibson retweeted
62% faster time to first token. 22% more tokens/sec. Leo X-Series puts #KVCache closer to GPUs with fabric-attached memory and Scorpio™ Smart Fabric Switches - supporting PCIe and platform-specific GPUs at rack scale. Learn more: asteralabs.com/products/leo-… #AsteraLabs #AI
1
2
9
457
Multi Head Attention vs Multi Query Attention #LLM #Transformers #KVCache #MQA #MHA #AIEngineering #Inference
34
GCC/MENA inference teams: a smaller KV cache is not a faster service. EucaX on why KV-cache work is now a memory-systems discipline. eucax.ai/insights/kv-cache-o… #LLMInference #KVCache #AIInfrastructure
15
ScaleFlux is aiming an SSD platform at NVIDIA CMX KV cache: 7 to 10+ effective DWPD, 200+ FDP write streams per drive, and 2x lower write amplification in preliminary testing. All ScaleFlux figures. storagereview.com/news/scale… #KVcache #FDP #SSD @ScaleFlux
8
769
FP8 KV cache is underrated. Weights load once, but KV grows with every token and batch. Halving KV from FP16 to FP8 roughly doubles the context you hold at fixed VRAM, and attention is memory-bound so the accuracy hit stays small. #quantization #KVcache #FP8 #inference #LLM
2
3
29
We're pulling 16 DIMMs out of a Dell PowerEdge XE7740 on purpose. With DRAM near a quarter million at list, the question is how much KV cache you can push onto flash instead. Half the RAM, full bandwidth, more stress on the flash, results soon. #KVcache #PowerEdge #SolidigmSSD @solidigm
2
8
85
20,273
AI がトレーニングから #inference へシフトすると、#memory 需要にはどのような変化が起きるのでしょうか?注目すべき 2 つのドライバーは、KV キャッシュのオフロードとエージェント型 AI です。詳細はこちら:pse.is/97cx86 🔗 #KVcache #AgenticAI
2
400
What happens to #memory demand when AI shifts from training to #inference? KV cache offloading and agentic AI are two drivers worth watching. More: pse.is/97cx86 🔗 #KVcache #AgenticAI
13
59
5,923