Skip to main content
buildradar
Sign in

NVIDIA/TensorRT-LLM

@NVIDIA

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Stars
14,536
Forks
2,716
Language
Python
License
NOASSERTION
Last push
4 days ago
Pythoncudapytorchllm-servingmoeblackwell

No related intel yet

This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.