NVIDIA Build
Overview
NVIDIA Build is a developer-facing platform designed to help engineers and researchers build AI applications from the ground up. It is part of NVIDIA's broader transformation from a GPU company into the world's first full-stack AI infrastructure provider — spanning silicon, foundation models, robotics, autonomous vehicles, and physical AI. NVIDIA AI Models 2026 Guide It brings together free inference endpoints, pre-built agentic capabilities, reusable workflow blueprints, and secure agent execution tooling in a single place. NVIDIA Build Overview
Key Components
NemoClaw
NemoClaw is NVIDIA Build's safe agent execution environment. It provides:
- Access control — manage who and what can invoke agent actions.
- Data protection — safeguard sensitive information during agent runs.
NemoClaw is positioned as the secure foundation for building personal and enterprise AI agents. NVIDIA Build Overview
Agentic Skills
Pre-built capabilities that an agent can call directly. Skills are organized into four domains: NVIDIA Build Overview
| Domain | Skills Available |
|---|---|
| AI and Machine Learning | 143 |
| Physical AI | 37 |
| Accelerated Computing | 25 |
| Developer Tools | 10 |
Blueprint Collection
A curated set of workflows and code samples that help developers build AI applications from scratch. Blueprints are intended as starting points that encode best practices and end-to-end patterns. NVIDIA Build Overview
Inference Endpoints
The model catalog can be filtered by use case (e.g. Drug Discovery, Image-to-Text, Retrieval Augmented Generation, Speech-to-Text, Code Generation), inference provider (e.g. Deepinfra, OpenRouter, Together AI, GMI Cloud, Lightning AI), publisher, and supported NIM container GPU (H100, B200, L40S, H200, A100 SXM4 80GB). NVIDIA Build Models Catalog NVIDIA Build provides free inference access to a broad catalog of leading models. Highlighted models available at the time of the overview include: NVIDIA Build Overview
- z-ai / glm-5.2 — Agentic AI; flagship LLM for agentic workflows, coding, and long-horizon reasoning
- nvidia / nemotron-3-ultra-550b-a55b — Agent; open hybrid Mamba-Transformer MoE with 1M context
- nvidia / cosmos3-nano — Physical AI; generates physics-aware videos from text or image prompts
- nvidia / cosmos3-nano-reasoner — Physical AI; VLM for understanding the physical world via structured reasoning on video/images
- deepseek-ai / deepseek-v4-flash — MoE; 284B model with 1M-token context optimized for fast coding and agents
- deepseek-ai / deepseek-v4-pro — MoE; scales to 1M-token context with efficient MoE architecture
- moonshotai / kimi-k2.6 — Multimodal; 1T MoE for long-horizon coding, agentic tool use, and image/video understanding
- mistral-ai / mistral-medium-3.5-128b — Coding/Agentic; high-performing model for text generation, coding, and agentic use cases
- stepfun-ai / step-3.7-flash — Coding; sparse MoE multimodal reasoning model for enterprise, agentic, and coding tasks
- minimaxai / minimax-m3 — Coding; multimodal MoE vision-language model with strong reasoning and tool-calling
- google / diffusiongemma-26b-a4b-it — Diffusion LLM; diffusion-based 26B parameter LLM for parallel token generation
- nvidia / nemotron-3-nano-omni-30b-a3b-reasoning — Image-to-Text; omni-modal reasoning model understanding images, video, speech, and text
- nvidia / nemotron-3.5-content-safety — LLM Safety; multilingual, multimodal model for detecting unsafe and toxic content
- nvidia / synthetic-video-detector — Broadcast; AI-powered microservice for detecting AI-generated (synthetic) videos
- nvidia / ising-calibration-1-35b-a3b — Quantum; open VLM for quantum computer calibration chart understanding
The catalog lists 141 models in total, with 77 free endpoints and 43 partner endpoints, spanning publishers including NVIDIA, Meta, Google, Mistral AI, and Qwen. NVIDIA Build Models Catalog See 80+ Free AI Models at NVIDIA Build for the broader model catalog.
DGX Station
NVIDIA Build includes step-by-step playbooks for getting started with DGX Station, including guidance on setting up NemoClaw as a secure personal AI agent. NVIDIA Build Overview
Platform Architecture
Summary
NVIDIA Build consolidates the essential building blocks for AI application development — secure execution via NemoClaw, reusable Agentic Skills, ready-made Blueprint Collection workflows, and free Inference Endpoints — into a single developer hub. NVIDIA Build Overview
NVIDIA's 2026 AI Model Ecosystem
NVIDIA Build serves as the primary distribution point for NVIDIA's open model portfolio. As of 2026, that portfolio spans eight major model families, all available free on Hugging Face, GitHub, and build.nvidia.com under commercial-friendly licenses (Apache 2.0, MIT, or NVIDIA Open Model License). NVIDIA AI Models 2026 Guide
| Model Family | Domain | Notable Models |
|---|---|---|
| Nemotron | Agentic LLMs | Nemotron 3 Ultra, Super, Nano, VoiceChat |
| PersonaPlex | Voice AI | PersonaPlex-7B-v1 (full-duplex speech) |
| Cosmos | Physical AI simulation | Cosmos Predict 2.5, Cosmos Transfer 2.5 |
| GR00T | Humanoid robotics | GR00T N1.7 (VLA model) |
| Alpamayo | Autonomous vehicles | Alpamayo R1 (open reasoning VLA) |
| Clara | Biomedical | — |
| Earth-2 | Climate science | — |
| Nemotron Speech | ASR / RAG | Real-time low-latency ASR |
NVIDIA also released one of the largest open training datasets in AI history alongside these models, including 10 trillion language training tokens, 500,000 robotics trajectories, 455,000 protein structures, and 100 terabytes of vehicle sensor data. NVIDIA AI Models 2026 Guide
For a full analysis of the model portfolio, rankings, and competitive benchmarks, see NVIDIA AI Models 2026 Guide.
PersonaPlex 7B
Released January 15, 2026, PersonaPlex-7B-v1 is a 7-billion parameter full-duplex speech-to-speech conversational AI. Unlike turn-based voice assistants, it listens and speaks simultaneously using a dual-stream Transformer (Moshi architecture). Key metrics: NVIDIA AI Models 2026 Guide
- Smooth turn-taking latency: 0.170 seconds
- User interruption latency: 0.240 seconds
- FullDuplexBench smooth turn-taking takeover rate: 0.908
- Outperforms Gemini Live, Qwen 2.5 Omni, and Moshi on conversational dynamics
Model weights are under the NVIDIA Open Model License; code is MIT-licensed. Requires NVIDIA Ampere (A100) or Hopper (H100) architecture GPUs for real-time performance. NVIDIA AI Models 2026 Guide
Nemotron 3
Nemotron 3, announced at GTC March 2026, uses a Hybrid Mamba-Transformer Mixture-of-Experts architecture delivering 5x throughput efficiency (NVFP4 format) on Blackwell chips. It is available through NVIDIA NIM microservices for managed cloud deployment. Notable enterprise adopters include CrowdStrike, ServiceNow, Perplexity, Cursor, Palantir, Salesforce, and Bosch. NVIDIA AI Models 2026 Guide
Physical AI Stack
Build.nvidia.com also hosts NVIDIA's physical AI models: NVIDIA AI Models 2026 Guide
- GR00T N1.7 — Vision language action model for humanoid robots; commercially viable as of GTC March 2026. Tops MolmoSpaces and RoboArena benchmarks. LG Electronics and NEURA Robotics are early adopters.
- Cosmos 2.5 — Physical AI simulation platform; generates synthetic training data across diverse conditions. Used by Johnson & Johnson MedTech, Toyota Research Institute, Salesforce, and Uber.
- Alpamayo R1 — First open reasoning VLA model for autonomous vehicles. Mercedes-Benz is building it into the all-new CLA on the NVIDIA DRIVE platform.
Rubin Hardware Platform
All NVIDIA models ultimately run on NVIDIA hardware. The Rubin platform — NVIDIA's first extreme-codesigned six-chip AI system (GPU + Vera CPU + BlueField-4 DPU), launched at CES 2026 — delivers AI token generation at one-tenth the cost of the previous Blackwell platform. Nemotron 3 is specifically tuned for Rubin, deepening the hardware-software integration for developers building through NVIDIA NIM. NVIDIA AI Models 2026 Guide