NVIDIA/TransformerEngine
@NVIDIAA library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
星数
3,518
Fork 数
816
语言
Python
许可
Apache-2.0
最后推送
4天前
相关情报(0)
还没有相关情报
radar 追踪的来源里还没有出现过这个 repo。收集器按计划运行——等它覆盖到这个 repo 再回来看看。