NVIDIA/TransformerEngine
@NVIDIAA library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Stars
3,509
Forks
814
Language
Python
License
Apache-2.0
Last push
3 days ago
Related intel (0)
No related intel yet
This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.