LLM推論2本の記事が同じ件を報じた初出 6/4 23:47
KVarN: Huaweiからの新しいKVキャッシュ量子化。3-5倍のKVキャッシュ圧縮で速度低下ではなく実際の速度向上を実現し、TurboQuantとは異なり推論でも性能を維持 (Apache 2.0, vLLM single flag)
KVarN: new KV-cache quant from Huawei. 3–5× KV cache compression with actual speed-up instead of slow-down, and unlike TurboQuant it holds up on reasoning (Apache 2.0, vLLM single flag)
https://www.reddit.com/r/LocalLLaMA/.rss2026/6/4
AI要約
Huaweiが開発したKVキャッシュ圧縮技術「KVarN」に関する情報。3-5倍の圧縮率で速度低下なく、推論性能も維持できる。vLLMとの連携も可能で、Apache 2.0ライセンス。KVキャッシュの効率化に貢献する技術。