Skip to main content
buildradar
Sign in

NVIDIA/Model-Optimizer

@NVIDIA

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Stars
3,697
Forks
578
Language
Python
License
Apache-2.0
Last push
4 days ago
Python

No related intel yet

This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.