kv-cache
Tracked open-source repos tagged kv-cache, sorted by stars.
Related topics
Topics that frequently appear alongside kv-cache on the same repo.
Recent risers
Repos created in the last 90 days, tagged kv-cache.
- #1
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,035 - #2
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 564
- #1★ 11,562+256Star change over the last 7 days
- #2★ 3,833-1Star change over the last 7 days
- #3
Unified KV Cache Compression Methods for Auto-Regressive Models
★ 1,376+1Star change over the last 7 days - #4★ 1,197+16Star change over the last 7 days
- #5
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,035+504Star change over the last 7 days - #6
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
★ 889+0Star change over the last 7 days - #7
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 667+4Star change over the last 7 days - #8
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
★ 630-1Star change over the last 7 days - #9
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation (NeurIPS 2025)
★ 581+0Star change over the last 7 days - #10
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 564+35Star change over the last 7 days - #11
[NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.
★ 532+3Star change over the last 7 days