cuda
Tracked open-source repos tagged cuda, sorted by stars.
- #181
Best practices & guides on how to write distributed pytorch training code
★ 631+2Star change over the last 7 days - #182★ 622+2Star change over the last 7 days
- #183
High-Performance Cross-Platform Monte Carlo Renderer Based on LuisaCompute
★ 618+1Star change over the last 7 days - #184★ 616+0Star change over the last 7 days
- #185
AutoDock for GPUs and other accelerators
★ 610+1Star change over the last 7 days - #186
A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.
★ 604+0Star change over the last 7 days - #187
An app and guide to easily configure Windows, Linux, MacOS, Google TV, Stremio, Home Assistant and more (including WSL2, GPU drivers & development tools). Improve your UX & productivity.
★ 602+0Star change over the last 7 days - #188
Go with your own intelligence - Write Go applications that directly integrate llama.cpp for local inference using hardware acceleration on Linux, macOS, Windows, & WebAssembly.
★ 599+17Star change over the last 7 days - #189★ 597+1Star change over the last 7 days
- #190★ 594+1Star change over the last 7 days
- #191★ 592+0Star change over the last 7 days
- #192
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
★ 584+4Star change over the last 7 days - #193
Radar Simulator built with Python and C++
★ 576+3Star change over the last 7 days - #194
A single-header C++ library for simplifying the use of CUDA Runtime Compilation (NVRTC).
★ 574+0Star change over the last 7 days - #195
校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。
★ 573+0Star change over the last 7 days - #196
An open collection of methodologies to help with successful training of large language models.
★ 569+2Star change over the last 7 days - #197
Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruction.
★ 568+0Star change over the last 7 days - #198
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 566+0Star change over the last 7 days - #199
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
★ 560+14Star change over the last 7 days - #200★ 560+1Star change over the last 7 days
- #201
LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。
★ 559+20Star change over the last 7 days - #202
Use your NVIDIA GPU's VRAM as swap space on Linux. Built for laptops with soldered memory and no upgrade path. If you have an RTX card sitting there with 8GB of VRAM and you're getting swapped to SSD, this puts that VRAM to work
★ 555+4Star change over the last 7 days - #203★ 552+2Star change over the last 7 days
- #204
A Slurm cluster using docker-compose
★ 544+2Star change over the last 7 days - #205★ 532+0Star change over the last 7 days
- #206
a reimplementation of Holistically-Nested Edge Detection in PyTorch
★ 525+0Star change over the last 7 days - #207★ 518-1Star change over the last 7 days
- #208★ 518+0Star change over the last 7 days
- #209
Docker Image for Ubuntu Desktop which support HW GPU accelerated GUI apps. you can access the Container with ssh or remote desktop, just like Cloud VM.
★ 515+0Star change over the last 7 days - #210
Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WS
★ 512+0Star change over the last 7 days