NVIDIA/TensorRT-LLM
@NVIDIATensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Stars
14,536
Forks
2,716
Language
Python
License
NOASSERTION
Last push
4 days ago
Related intel (0)
No related intel yet
This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.