Skip to content
#

kv-cache-compression

Here are 39 public repositories matching this topic...

Native Windows vLLM 0.25.1—prebuilt Python 3.13/CUDA 12.8 wheels for RTX 30/40/50 GPUs, no WSL or Docker. OpenAI-compatible serving with Triton/FlashAttention, 10 KV-cache compression formats, Multi-TurboQuant, and experimental persistent CPU/NVMe KV offload.

  • Updated Jul 19, 2026
  • Python

Discrete Kakeya cover for LLM KV cache: D4/E8 nested-lattice quantisation realising a Kakeya-style tube-cover over the direction sphere. 2.4x-2.8x compression at <1% perplexity loss on Qwen3, Llama-3, DeepSeek, GLM-4, Gemma. Drop-in transformers.DynamicCache. pip install kakeyalattice.

  • Updated Jun 15, 2026
  • Python

Research and training stack for AVA — a tool-using, memory-aware virtual assistant targeting 4 GB VRAM. Spans custom transformers, verifier-RL, external memory, multi-domain benchmarks, and Gemma 4 inference optimization.

  • Updated Jul 25, 2026
  • Python

Improve this page

Add a description, image, and links to the kv-cache-compression topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the kv-cache-compression topic, visit your repo's landing page and select "manage topics."

Learn more