speculative-decoding
Tracked open-source repos tagged speculative-decoding, sorted by stars.
Related topics
Topics that frequently appear alongside speculative-decoding on the same repo.
Recent risers
Repos created in the last 90 days, tagged speculative-decoding.
- #1
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,185 - #2
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 640
- #1★ 2,829+16Star change over the last 7 days
- #2
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2,521+7Star change over the last 7 days - #3
3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
★ 1,944+173Star change over the last 7 days - #4★ 1,847+3Star change over the last 7 days
- #5
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1,619+58Star change over the last 7 days - #6
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,185+285Star change over the last 7 days - #7
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
★ 791+26Star change over the last 7 days - #8★ 675+5Star change over the last 7 days
- #9
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 640+22Star change over the last 7 days - #10
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 566+30Star change over the last 7 days