NVIDIA
NVIDIA's tracked open-source repos, sorted by stars.
- #1
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
★ 22,308+73Star change over the last 7 days - #2
Ongoing research training transformer models at scale
★ 17,669+159Star change over the last 7 days - #3
NVIDIA Linux open GPU kernel module source
★ 17,321+15Star change over the last 7 days - #4
Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.
★ 15,196+340Star change over the last 7 days - #5
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
★ 14,842+1Star change over the last 7 days - #6
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
★ 14,498+58Star change over the last 7 days - #7
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
★ 13,305+32Star change over the last 7 days - #8
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11,669+85Star change over the last 7 days - #9
PersonaPlex code.
★ 10,396+12Star change over the last 7 days - #10★ 10,345+64Star change over the last 7 days
- #11★ 9,739+6Star change over the last 7 days
- #12
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
★ 9,566+31Star change over the last 7 days - #13★ 9,071+181Star change over the last 7 days
- #14★ 8,996+5Star change over the last 7 days
- #15★ 8,417+101Star change over the last 7 days
- #16
NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7,951+69Star change over the last 7 days - #17★ 7,053+30Star change over the last 7 days
- #18★ 6,924+1Star change over the last 7 days
- #19
Transformer related optimization, including BERT, GPT
★ 6,447+4Star change over the last 7 days - #20
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
★ 5,738+4Star change over the last 7 days - #21★ 5,295-1Star change over the last 7 days
- #22★ 5,270+13Star change over the last 7 days
- #23★ 5,032+12Star change over the last 7 days
- #24
Build and run containers leveraging NVIDIA GPUs
★ 4,531+14Star change over the last 7 days - #25
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
★ 4,165+10Star change over the last 7 days - #26
NVIDIA device plugin for Kubernetes
★ 3,862+6Star change over the last 7 days - #27
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
★ 3,568+103Star change over the last 7 days - #28
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
★ 3,509+10Star change over the last 7 days - #29
CUDA Python: Performance meets Productivity
★ 3,367+24Star change over the last 7 days - #30
Pytorch implementation of FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
★ 3,289+0Star change over the last 7 days - #31
Open-source deep-learning framework for building, training, and fine-tuning deep learning models using state-of-the-art Physics-ML methods
★ 3,208+24Star change over the last 7 days - #32
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
★ 3,137+83Star change over the last 7 days - #33
NeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice. NeMo Retriever Library uses specialized NVIDIA NIM microservices to find, contextualize, and extract text, tables, charts and images that you can use in downstream generative applications.
★ 2,971+7Star change over the last 7 days - #34
Minkowski Engine is an auto-diff neural network library for high-dimensional sparse tensors
★ 2,956+1Star change over the last 7 days - #35
NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes
★ 2,853+9Star change over the last 7 days - #36
The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
★ 2,605+11Star change over the last 7 days - #37★ 2,495+12Star change over the last 7 days
- #38
CUDA Library Samples
★ 2,488+4Star change over the last 7 days - #39
`std::execution`, the standard C++ framework for asynchronous and parallel programming.
★ 2,425+9Star change over the last 7 days - #40
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
★ 2,138+6Star change over the last 7 days - #41
TensorRT Extension for Stable Diffusion Web UI
★ 1,989+0Star change over the last 7 days - #42
NVIDIA curated collection of educational resources related to general purpose GPU programming.
★ 1,949+26Star change over the last 7 days - #43★ 1,918+2Star change over the last 7 days
- #44
NVIDIA GPU metrics exporter for Prometheus leveraging DCGM
★ 1,849+11Star change over the last 7 days - #45
Simple samples for TensorRT programming
★ 1,668+2Star change over the last 7 days - #46
NCCL Tests
★ 1,639+11Star change over the last 7 days - #47
This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?
★ 1,610+1Star change over the last 7 days - #48★ 1,469+2Star change over the last 7 days
- #49★ 1,444+2Star change over the last 7 days
- #50
NVIDIA DLSS is a new and improved deep learning neural network that boosts frame rates and generates beautiful, sharp images for your games
★ 1,426+10Star change over the last 7 days - #51★ 1,412+2Star change over the last 7 days
- #52
Documentation of NVIDIA chip/hardware interfaces
★ 1,360+1Star change over the last 7 days - #53
Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.
★ 1,306+20Star change over the last 7 days - #54★ 1,226+0Star change over the last 7 days
- #55★ 1,196+16Star change over the last 7 days
- #56
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
★ 1,182+1Star change over the last 7 days - #57
NVIDIA container runtime library
★ 1,122+1Star change over the last 7 days - #58
C++ and Python support for the CUDA Quantum programming model for heterogeneous quantum-classical workflows
★ 1,122+4Star change over the last 7 days - #59
Open-source deep-learning framework for exploring, building and deploying AI weather/climate workflows.
★ 1,105+12Star change over the last 7 days - #60
A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.
★ 1,095+3Star change over the last 7 days - #61
A Python library that enables the use of Jetson's GPIOs
★ 1,075+0Star change over the last 7 days - #62
Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
★ 1,066+17Star change over the last 7 days - #63
RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form building blocks for more easily writing high performance applications.
★ 1,037+1Star change over the last 7 days - #64★ 1,032+11Star change over the last 7 days
- #65
CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA tensor core units.
★ 1,019+7Star change over the last 7 days - #66
NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs
★ 998+2Star change over the last 7 days - #67★ 961+4Star change over the last 7 days
- #68★ 924+4Star change over the last 7 days
- #69
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
★ 918+5Star change over the last 7 days - #70
Examples demonstrating available options to program multiple GPUs in a single node or a cluster
★ 911+0Star change over the last 7 days - #71★ 842+3Star change over the last 7 days
- #72★ 806+5Star change over the last 7 days
- #73
NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs
★ 784+2Star change over the last 7 days - #74
A tool for bandwidth measurements on NVIDIA GPUs.
★ 758+4Star change over the last 7 days - #75
Differentiable signal processing on the sphere for PyTorch
★ 695-1Star change over the last 7 days - #76★ 690+3Star change over the last 7 days
- #77★ 661+0Star change over the last 7 days
- #78★ 641+0Star change over the last 7 days
- #79
DreamGen: Nvidia GEAR Lab's initiative to solve the robotics data problem using world models
★ 604+5Star change over the last 7 days - #80
NVIDIA Math Libraries for the Python Ecosystem
★ 597+1Star change over the last 7 days - #81
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
★ 578+2Star change over the last 7 days - #82
A single-header C++ library for simplifying the use of CUDA Runtime Compilation (NVRTC).
★ 574+0Star change over the last 7 days - #83
The NVIDIA® Tools Extension SDK (NVTX) is a C-based Application Programming Interface (API) for annotating events, code ranges, and resources in your applications.
★ 556+1Star change over the last 7 days - #84
SOMA BVH to humanoid robot motion retargeting library built with Newton and NVIDIA Warp
★ 549+9Star change over the last 7 days - #85★ 538+0Star change over the last 7 days
- #86★ 535+0Star change over the last 7 days
- #87
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
★ 528+10Star change over the last 7 days - #88
HPC Container Maker
★ 517+2Star change over the last 7 days