Skip to main content
buildradar
Sign in
Topic · speculative-decoding

speculative-decoding

Tracked open-source repos tagged speculative-decoding, sorted by stars.

Repos
10
Total stars
14,467
Avg. stars
1,447
Share
0.00%

Topics that frequently appear alongside speculative-decoding on the same repo.

Recent risers

Repos created in the last 90 days, tagged speculative-decoding.

  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,185
  • mlx-dspark@ARahim3

    Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

    640
  • lucebox@Luce-Org

    LLM speculative inference server for heterogeneous hardware & consumer GPUs

    2,829+16Star change over the last 7 days
  • EAGLE@SafeAILab

    Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).

    2,521+7Star change over the last 7 days
  • MTPLX@youssofal

    3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.

    1,944+173Star change over the last 7 days
  • sonar@dphnAI

    Large-scale LLM inference engine

    1,847+3Star change over the last 7 days
  • AngelSlim@Tencent

    Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

    1,619+58Star change over the last 7 days
  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,185+285Star change over the last 7 days
  • A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

    791+26Star change over the last 7 days
  • atlas@Avarok-Cybersecurity

    Pure Rust Inference Engine

    675+5Star change over the last 7 days
  • mlx-dspark@ARahim3

    Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

    640+22Star change over the last 7 days
  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    566+30Star change over the last 7 days
← Back to topics