Skip to main content
buildradar
Sign in

mit-han-lab/omniserve

@mit-han-lab

[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Stars
856
Forks
69
Language
C++
License
Apache-2.0
Last push
2 years ago
C++

No related intel yet

This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.