跳到主要内容
buildradar
Sign in

NVIDIA/TransformerEngine

@NVIDIA

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.

星数
3,518
Fork 数
816
语言
Python
许可
Apache-2.0
最后推送
4天前
Pythonpythonmachine-learningdeep-learningcudapytorchjaxgpufp8fp4

还没有相关情报

radar 追踪的来源里还没有出现过这个 repo。收集器按计划运行——等它覆盖到这个 repo 再回来看看。