Skip to main content
buildradar
Sign in
Topic · kv-cache

kv-cache

Tracked open-source repos tagged kv-cache, sorted by stars.

Repos
11
Total stars
22,673
Avg. stars
2,061
Share
0.00%

Topics that frequently appear alongside kv-cache on the same repo.

Recent risers

Repos created in the last 90 days, tagged kv-cache.

  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,035
  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    564
  • LMCache@LMCache

    LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

    11,562+256Star change over the last 7 days
  • godis@HDT3213

    A Golang implemented Redis Server and Cluster. Go 语言实现的 Redis 服务器和分布式集群

    3,833-1Star change over the last 7 days
  • KVCache-Factory@Zefan-Cai

    Unified KV Cache Compression Methods for Auto-Regressive Models

    1,376+1Star change over the last 7 days
  • kvpress@NVIDIA

    LLM KV cache compression made easy

    1,197+16Star change over the last 7 days
  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,035+504Star change over the last 7 days
  • llm_note@harleyszhang

    LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.

    889+0Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    667+4Star change over the last 7 days
  • Deepdive-llama3-from-scratch@therealoliver

    Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.

    630-1Star change over the last 7 days
  • mixture_of_recursions@raymin0223

    Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation (NeurIPS 2025)

    581+0Star change over the last 7 days
  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    564+35Star change over the last 7 days
  • H2O@FMInference

    [NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

    532+3Star change over the last 7 days
← Back to topics