gpu
Tracked open-source repos tagged gpu, sorted by stars.
- #121★ 1,664+2Star change over the last 7 days
- #122
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
★ 1,655+11Star change over the last 7 days - #123
Awesome Lists for Tenure-Track Assistant Professors and PhD students. (助理教授/博士生生存指南)
★ 1,646+0Star change over the last 7 days - #124★ 1,625+0Star change over the last 7 days
- #125
ThunderSVM: A Fast SVM Library on GPUs and CPUs
★ 1,622+0Star change over the last 7 days - #126
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1,613+33Star change over the last 7 days - #127★ 1,600+79Star change over the last 7 days
- #128★ 1,577+1Star change over the last 7 days
- #129
A lightweight 2D graphics library for modern GPUs, delivering high-performance text, image, and vector rendering across major platforms.
★ 1,576+1Star change over the last 7 days - #130
a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.
★ 1,550+1Star change over the last 7 days - #131
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
★ 1,545+7Star change over the last 7 days - #132
Next-gen fast plotting library running on WGPU using the pygfx rendering engine
★ 1,524+0Star change over the last 7 days - #133
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
★ 1,502+3Star change over the last 7 days - #134
A GPU-accelerated computing library for running physics simulations and other GPGPU computations in a web browser.
★ 1,483+0Star change over the last 7 days - #135
Visual SLAM/odometry package based on NVIDIA-accelerated cuVSLAM
★ 1,454+5Star change over the last 7 days - #136★ 1,444+0Star change over the last 7 days
- #137
Intel® Graphics Compute Runtime for oneAPI Level Zero and OpenCL™ Driver
★ 1,434+2Star change over the last 7 days - #138★ 1,426+2Star change over the last 7 days
- #139
🌊 Julia software for fast, friendly, flexible, ocean-flavored fluid dynamics on CPUs and GPUs
★ 1,412+0Star change over the last 7 days - #140
Pre-built Mesa3D drivers for Windows
★ 1,410+3Star change over the last 7 days - #141
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
★ 1,406+0Star change over the last 7 days - #142★ 1,367+2Star change over the last 7 days
- #143★ 1,358+3Star change over the last 7 days
- #144
Extension for Scikit-learn is a seamless way to speed up your Scikit-learn application
★ 1,356+0Star change over the last 7 days - #145★ 1,319+0Star change over the last 7 days
- #146
🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring, Leader-Worker architecture, constant-size memory.
★ 1,288+3Star change over the last 7 days - #147★ 1,270+0Star change over the last 7 days
- #148★ 1,269+0Star change over the last 7 days
- #149
Examples of programs built using Modal
★ 1,264+2Star change over the last 7 days - #150
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
★ 1,227+7Star change over the last 7 days