Skip to main content
buildradar
Sign in
Topic · cuda

cuda

Tracked open-source repos tagged cuda, sorted by stars.

216 repos
  • tt-metal@tenstorrent

    :metal: TT-NN operator library, and TT-Metalium low level kernel programming model.

    1,655+0Star change over the last 7 days
  • gpu-hot@psalias2006

    🔥 Real-time NVIDIA GPU dashboard

    1,625+0Star change over the last 7 days
  • thundersvm@Xtra-Computing

    ThunderSVM: A Fast SVM Library on GPUs and CPUs

    1,622+0Star change over the last 7 days
  • mnn-llm@wangzhaode

    llm deploy project based mnn. This project has merged into MNN.

    1,617+0Star change over the last 7 days
  • InferenceX@SemiAnalysisAI

    Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

    1,613+0Star change over the last 7 days
  • 3d-ken-burns@sniklaus

    an implementation of 3D Ken Burns Effect from a Single Image using PyTorch

    1,569+0Star change over the last 7 days
  • autokernel@RightNow-AI

    Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.

    1,545+0Star change over the last 7 days
  • uccl@uccl-project

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    1,502+0Star change over the last 7 days
  • TornadoVM@beehive-lab

    Write Java. Run on GPUs. Fast.

    1,492+0Star change over the last 7 days
  • 🔥🔥🔥TensorRT for YOLOv8、YOLOv8-Pose、YOLOv8-Seg、YOLOv8-Cls、YOLOv7、YOLOv6、YOLOv5、YOLONAS......🚀🚀🚀CUDA IS ALL YOU NEED.🍎🍎🍎

    1,460+0Star change over the last 7 days
  • MatX@NVIDIA

    An efficient C++20 GPU numerical computing library with Python-like syntax

    1,444+0Star change over the last 7 days
  • Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice cloning, and large audiobook-scale text processing. Runs accelerated on NVIDIA (CUDA), AMD (ROCm), and CPU.

    1,427+0Star change over the last 7 days
  • CUDA.jl@JuliaGPU

    CUDA programming in Julia.

    1,426+0Star change over the last 7 days
  • Arraymancer@mratsim

    A fast, ergonomic and portable tensor library in Nim with a deep learning focus for CPU, GPU and embedded devices via OpenMP, Cuda and OpenCL backends

    1,407+0Star change over the last 7 days
  • flux@bytedance

    A fast communication-overlapping library for tensor/expert parallelism on GPUs.

    1,358+0Star change over the last 7 days
  • LuxCore@LuxCoreRender

    LuxCore source repository

    1,323+0Star change over the last 7 days
  • stable-fast@chengzeyi

    https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA GPUs.

    1,304+0Star change over the last 7 days
  • stdgpu@stotko

    stdgpu: Efficient STL-like Data Structures on the GPU

    1,270+0Star change over the last 7 days
  • graphvite@DeepGraphLearning

    GraphVite: A General and High-performance Graph Embedding System

    1,269+0Star change over the last 7 days
  • ffmpeg-build-script@markus-perl

    The FFmpeg build script provides an easy way to build a static FFmpeg on OSX and Linux with non-free codecs included.

    1,215+0Star change over the last 7 days
  • cudarc@chelsea0x3b

    Safe rust wrapper around CUDA toolkit

    1,214+0Star change over the last 7 days
  • oneAPI-samples@oneapi-src

    Samples for Intel® oneAPI Toolkits

    1,162+0Star change over the last 7 days
  • pyopencl@inducer

    OpenCL integration for Python, plus shiny features

    1,150+0Star change over the last 7 days
  • juice@fff-rs

    The Hacker's Machine Learning Engine

    1,131+0Star change over the last 7 days
  • neanderthal@uncomplicate

    Fast Clojure Matrix Library

    1,128+0Star change over the last 7 days
  • sglang-omni@sgl-project

    SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

    1,118+56Star change over the last 7 days
  • tiny-vllm@jmaczan

    Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

    1,100+10Star change over the last 7 days
  • gunrock@gunrock

    Programmable CUDA/C++ GPU Graph Analytics

    1,096+0Star change over the last 7 days
  • cupoch@neka-nat

    Robotics with GPU computing

    1,064+1Star change over the last 7 days
  • LuisaCompute@LuisaGroup

    High-Performance Rendering Framework on Stream Architectures

    1,044+1Star change over the last 7 days
← Back to topics