speculative-decoding
標記 speculative-decoding 主題、收錄中的開源專案,依星數排序。
相關主題
常跟 speculative-decoding 一起出現在同一個專案上的主題。
近期新秀
近 90 天內建立、標記 speculative-decoding 主題的專案。
- #1
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,165 - #2
在 Apple Silicon 上實現高達 4 倍更快的 LLM 解碼,無損。原生 MLX 移植 DeepSeek 的 DSpark 與 z-lab 的 DFlash 推測解碼 — Gemma-4、Qwen3.8、Muse-Glimmer、Nemotron、Ornith-1.0、三進制 Bonsai-27B。
★ 637
- #1★ 2,829+38近 7 天星數變化
- #2★ 2,521+11近 7 天星數變化
- #3★ 1,944+241近 7 天星數變化
- #4★ 1,847+4近 7 天星數變化
- #5★ 1,619+71近 7 天星數變化
- #6
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,165+317近 7 天星數變化 - #7
大型語言模型策略內蒸餾(OPD)的精選論文、技術報告、框架與工具集合
★ 786+18近 7 天星數變化 - #8★ 677+5近 7 天星數變化
- #9
在 Apple Silicon 上實現高達 4 倍更快的 LLM 解碼,無損。原生 MLX 移植 DeepSeek 的 DSpark 與 z-lab 的 DFlash 推測解碼 — Gemma-4、Qwen3.8、Muse-Glimmer、Nemotron、Ornith-1.0、三進制 Bonsai-27B。
★ 637+20近 7 天星數變化 - #10
從教師模型到模型切片——從零開始的 LLM 蒸餾與服務引擎:包含自訂 Triton/CUDA 核心、FSDP 蒸餾、paged-KV 連續批次處理、推測解碼、Rust 閘道、JAX 預測器與可解釋性工具。
★ 566+46近 7 天星數變化