mit-han-lab
mit-han-lab's tracked open-source repos, sorted by stars.
- #1
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
★ 7,268+0Star change over the last 7 days - #2
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
★ 3,624+1Star change over the last 7 days - #3
Efficient vision foundation models for high-resolution generation and perception.
★ 3,356+1Star change over the last 7 days - #4
[ICCV 2019] TSM: Temporal Shift Module for Efficient Video Understanding
★ 2,222+1Star change over the last 7 days - #5
[ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
★ 1,681+2Star change over the last 7 days - #6
A PyTorch-based framework for Quantum Classical Simulation, Quantum Machine Learning, Quantum Neural Networks, Parameterized Quantum Circuits with support for easy deployments on real quantum computers.
★ 1,663+0Star change over the last 7 days - #7
[MICRO'23, MLSys'22] TorchSparse: Efficient Training and Inference Framework for Sparse Convolution on GPUs.
★ 1,472+0Star change over the last 7 days - #8
[ICLR 2019] ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
★ 1,447+0Star change over the last 7 days - #9
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
★ 1,308+0Star change over the last 7 days - #10
[CVPR 2020] GAN Compression: Efficient Architectures for Interactive Conditional GANs
★ 1,115+0Star change over the last 7 days - #11
StreamingVLM: Real-Time Understanding for Infinite Video Streams
★ 1,076+3Star change over the last 7 days - #12
TinyChatEngine: On-Device LLM Inference Library
★ 961+0Star change over the last 7 days - #13
[NeurIPS 2020] MCUNet: Tiny Deep Learning on IoT Devices; [NeurIPS 2021] MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning; [NeurIPS 2022] MCUNetV3: On-Device Training Under 256KB Memory
★ 957+3Star change over the last 7 days - #14★ 955+67Star change over the last 7 days
- #15
[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
★ 856+2Star change over the last 7 days - #16
[CVPR 2024 Highlight] DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models
★ 728+0Star change over the last 7 days - #17
[IJCV] FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
★ 714-1Star change over the last 7 days - #18
[NeurIPS 2020] MCUNet: Tiny Deep Learning on IoT Devices; [NeurIPS 2021] MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
★ 713+1Star change over the last 7 days - #19★ 649+2Star change over the last 7 days
- #20
[NeurIPS 2025] Radial Attention: O(nlogn) Sparse Attention with Energy Decay for Long Video Generation
★ 607-1Star change over the last 7 days - #21
A sparse attention kernel supporting mix sparse patterns
★ 549+1Star change over the last 7 days - #22
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
★ 539+0Star change over the last 7 days - #23
On-Device Training Under 256KB Memory [NeurIPS'22]
★ 524+0Star change over the last 7 days