Skip to main content
buildradar
Sign in
Topic · cuda-kernels

cuda-kernels

Tracked open-source repos tagged cuda-kernels, sorted by stars.

Repos
13
Total stars
47,266
Avg. stars
3,636
Share
0.00%

Topics that frequently appear alongside cuda-kernels on the same repo.

Recent risers

Repos created in the last 90 days, tagged cuda-kernels.

No new repos tagged with this topic in the last 90 days.

  • LeetCUDA@xlite-dev

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    11,850+47Star change over the last 7 days
  • cuda-samples@NVIDIA

    Samples for CUDA Developers which demonstrates features in CUDA Toolkit

    9,566+31Star change over the last 7 days
  • lmdeploy@InternLM

    LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

    8,033+20Star change over the last 7 days
  • rust-cuda@Rust-GPU

    Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.

    5,331+5Star change over the last 7 days
  • lucebox@Luce-Org

    LLM speculative inference server for heterogeneous hardware & consumer GPUs

    2,813+35Star change over the last 7 days
  • cccl@NVIDIA

    CUDA Core Compute Libraries

    2,495+12Star change over the last 7 days
  • dfdx@chelsea0x3b

    Deep learning in Rust, with shape checked tensors and neural networks

    1,933+4Star change over the last 7 days
  • cudarc@chelsea0x3b

    Safe rust wrapper around CUDA toolkit

    1,215+4Star change over the last 7 days
  • framework@mni-ml

    A machine learning library with a TypeScript API and Rust backend. CUDA and WebGPU compatibility. Built to understand how ML frameworks and models work internally.

    998+10Star change over the last 7 days
  • nvbench@NVIDIA

    CUDA Kernel Benchmarking Library

    924+4Star change over the last 7 days
  • cudnn-frontend@NVIDIA

    cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

    918+5Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    660+4Star change over the last 7 days
  • FlashRT@flashrt-project

    FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B

    537+16Star change over the last 7 days
← Back to topics