Skip to main content
buildradar
Sign in
Topic · cuda

cuda

Tracked open-source repos tagged cuda, sorted by stars.

216 repos
  • rust-cuda@Rust-GPU

    Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.

    5,334+0Star change over the last 7 days
  • cuml@NVIDIA

    NVIDIA cuML: GPU-Accelerated Machine Learning

    5,269+0Star change over the last 7 days
  • kaolin@NVIDIAGameWorks

    A PyTorch Library for Accelerating 3D Deep Learning Research

    5,163+0Star change over the last 7 days
  • nccl@NVIDIA

    Optimized primitives for collective multi-GPU communication

    5,041+0Star change over the last 7 days
  • arrayfire@arrayfire

    ArrayFire: a general purpose GPU library.

    4,903+0Star change over the last 7 days
  • CTranslate2@OpenNMT

    Fast inference engine for Transformer models

    4,658+0Star change over the last 7 days
  • Tengine@OAID

    Tengine is a lite, high performance, modular inference engine for embedded device

    4,533+0Star change over the last 7 days
  • tiny-cuda-nn@NVlabs

    Lightning fast C++/CUDA neural network framework

    4,530+0Star change over the last 7 days
  • hip@ROCm

    HIP: C++ Heterogeneous-Compute Interface for Portability

    4,395+0Star change over the last 7 days
  • iree@iree-org

    A retargetable MLIR-based machine learning compiler and runtime toolkit.

    3,914+0Star change over the last 7 days
  • SageAttention@thu-ml

    [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

    3,699+0Star change over the last 7 days
  • Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.

    3,645+0Star change over the last 7 days
  • A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.

    3,518+0Star change over the last 7 days
  • viseron@roflcoopter

    Self-hosted, local only NVR and AI Computer Vision software. With features such as object detection, motion detection, face recognition and more, it gives you the power to keep an eye on your home, office or any other place you want to monitor.

    3,468+0Star change over the last 7 days
  • lygia@patriciogonzalezvivo

    LYGIA, it's a granular and multi-language (GLSL, HLSL, Metal, WGSL, WEGL and CUDA) shader library designed for performance and flexibility

    3,423+0Star change over the last 7 days
  • Remotery@Celtoys

    Single C file, Realtime CPU/GPU Profiler with Remote Web Viewer

    3,312+0Star change over the last 7 days
  • how to optimize some algorithm in cuda.

    3,242+0Star change over the last 7 days
  • jittor@Jittor

    Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.

    3,232+0Star change over the last 7 days
  • lc0@LeelaChessZero

    Open source neural network chess engine with GPU acceleration and broad hardware support.

    3,197+0Star change over the last 7 days
  • skills@NVIDIA

    Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

    3,180+0Star change over the last 7 days
  • cuda-oxide@NVlabs

    cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.

    3,118+0Star change over the last 7 days
  • aresdb@uber

    A GPU-powered real-time analytics storage and query engine.

    3,076+0Star change over the last 7 days
  • heavydb@heavyai

    HeavyDB (formerly MapD/OmniSciDB)

    3,059+0Star change over the last 7 days
  • ramalama@containers

    RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

    3,031+0Star change over the last 7 days
  • TensorRT@pytorch

    PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT

    2,989+0Star change over the last 7 days
  • ao@pytorch

    PyTorch native quantization for training and inference

    2,961+0Star change over the last 7 days
  • MinkowskiEngine@NVIDIA

    Minkowski Engine is an auto-diff neural network library for high-dimensional sparse tensors

    2,957+0Star change over the last 7 days
  • NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes

    2,857+0Star change over the last 7 days
  • lucebox@Luce-Org

    LLM speculative inference server for heterogeneous hardware & consumer GPUs

    2,829+0Star change over the last 7 days
  • futhark@diku-dk

    :boom::computer::boom: A data-parallel functional programming language

    2,794+0Star change over the last 7 days
← Back to topics