kv-cache
kv-cache がタグ付けされた追跡中のオープンソースリポジトリを、スター数順に表示します。
関連トピック
kv-cache と同じリポジトリに頻繁に登場するトピック。
最近の急上昇
直近 90 日以内に作成され、kv-cache がタグ付けされたリポジトリ。
- #1
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,109
- #1★ 11,622+152直近 7 日のスター増減
- #2★ 3,834+1直近 7 日のスター増減
- #3
自己回帰モデル向けの統合型 KV キャッシュ圧縮手法
★ 1,376+1直近 7 日のスター増減 - #4★ 1,201+7直近 7 日のスター増減
- #5
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,109+317直近 7 日のスター増減 - #6★ 889+1直近 7 日のスター増減
- #7★ 671+13直近 7 日のスター増減
- #8
llama3の推論をステップバイステップで実現し、コア概念を理解し、プロセスの導出をマスターし、コードを実装する。
★ 630+0直近 7 日のスター増減 - #9
Mixture-of-Recursions: アダプティブなトークンレベル計算のための動的再帰深度の学習 (NeurIPS 2025)
★ 587+2直近 7 日のスター増減 - #10
教師モデルからタイルまで — フルスクラッチのLLM蒸留・サービングエンジン:カスタムTriton/CUDAカーネル、FSDP蒸留、ページ化KV連続バッチ処理、推測デコーディング、Rustゲートウェイ、JAXオラクル、および解釈可能性ツール。
★ 566+46直近 7 日のスター増減 - #11★ 532+0直近 7 日のスター増減