cuda
Tracked open-source repos tagged cuda, sorted by stars.
- #91
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
★ 1,655+0Star change over the last 7 days - #92★ 1,625+0Star change over the last 7 days
- #93
ThunderSVM: A Fast SVM Library on GPUs and CPUs
★ 1,622+0Star change over the last 7 days - #94★ 1,617+0Star change over the last 7 days
- #95
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1,613+0Star change over the last 7 days - #96
an implementation of 3D Ken Burns Effect from a Single Image using PyTorch
★ 1,569+0Star change over the last 7 days - #97
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
★ 1,545+0Star change over the last 7 days - #98
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
★ 1,502+0Star change over the last 7 days - #99★ 1,492+0Star change over the last 7 days
- #100
🔥🔥🔥TensorRT for YOLOv8、YOLOv8-Pose、YOLOv8-Seg、YOLOv8-Cls、YOLOv7、YOLOv6、YOLOv5、YOLONAS......🚀🚀🚀CUDA IS ALL YOU NEED.🍎🍎🍎
★ 1,460+0Star change over the last 7 days - #101★ 1,444+0Star change over the last 7 days
- #102
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice cloning, and large audiobook-scale text processing. Runs accelerated on NVIDIA (CUDA), AMD (ROCm), and CPU.
★ 1,427+0Star change over the last 7 days - #103★ 1,426+0Star change over the last 7 days
- #104
A fast, ergonomic and portable tensor library in Nim with a deep learning focus for CPU, GPU and embedded devices via OpenMP, Cuda and OpenCL backends
★ 1,407+0Star change over the last 7 days - #105★ 1,358+0Star change over the last 7 days
- #106★ 1,323+0Star change over the last 7 days
- #107
https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA GPUs.
★ 1,304+0Star change over the last 7 days - #108★ 1,270+0Star change over the last 7 days
- #109★ 1,269+0Star change over the last 7 days
- #110
The FFmpeg build script provides an easy way to build a static FFmpeg on OSX and Linux with non-free codecs included.
★ 1,215+0Star change over the last 7 days - #111★ 1,214+0Star change over the last 7 days
- #112
Samples for Intel® oneAPI Toolkits
★ 1,162+0Star change over the last 7 days - #113★ 1,150+0Star change over the last 7 days
- #114★ 1,131+0Star change over the last 7 days
- #115
Fast Clojure Matrix Library
★ 1,128+0Star change over the last 7 days - #116
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 1,118+56Star change over the last 7 days - #117
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1,100+10Star change over the last 7 days - #118★ 1,096+0Star change over the last 7 days
- #119★ 1,064+1Star change over the last 7 days
- #120
High-Performance Rendering Framework on Stream Architectures
★ 1,044+1Star change over the last 7 days