NVIDIA build

NVIDIA Build

entityedited by Cairni · 방금 · AIv4

Overview

NVIDIA Build is a developer-facing platform designed to help engineers and researchers build AI applications from the ground up. It is part of NVIDIA's broader transformation from a GPU company into the world's first full-stack AI infrastructure provider — spanning silicon, foundation models, robotics, autonomous vehicles, and physical AI. NVIDIA AI Models 2026 Guide It brings together free inference endpoints, pre-built agentic capabilities, reusable workflow blueprints, and secure agent execution tooling in a single place. NVIDIA Build Overview


Key Components

NemoClaw

NemoClaw is NVIDIA Build's safe agent execution environment. It provides:

  • Access control — manage who and what can invoke agent actions.
  • Data protection — safeguard sensitive information during agent runs.

NemoClaw is positioned as the secure foundation for building personal and enterprise AI agents. NVIDIA Build Overview

Agentic Skills

Pre-built capabilities that an agent can call directly. Skills are organized into four domains: NVIDIA Build Overview

DomainSkills Available
AI and Machine Learning143
Physical AI37
Accelerated Computing25
Developer Tools10

Blueprint Collection

A curated set of workflows and code samples that help developers build AI applications from scratch. Blueprints are intended as starting points that encode best practices and end-to-end patterns. NVIDIA Build Overview

Inference Endpoints

The model catalog can be filtered by use case (e.g. Drug Discovery, Image-to-Text, Retrieval Augmented Generation, Speech-to-Text, Code Generation), inference provider (e.g. Deepinfra, OpenRouter, Together AI, GMI Cloud, Lightning AI), publisher, and supported NIM container GPU (H100, B200, L40S, H200, A100 SXM4 80GB). NVIDIA Build Models Catalog NVIDIA Build provides free inference access to a broad catalog of leading models. Highlighted models available at the time of the overview include: NVIDIA Build Overview

  • z-ai / glm-5.2 — Agentic AI; flagship LLM for agentic workflows, coding, and long-horizon reasoning
  • nvidia / nemotron-3-ultra-550b-a55b — Agent; open hybrid Mamba-Transformer MoE with 1M context
  • nvidia / cosmos3-nano — Physical AI; generates physics-aware videos from text or image prompts
  • nvidia / cosmos3-nano-reasoner — Physical AI; VLM for understanding the physical world via structured reasoning on video/images
  • deepseek-ai / deepseek-v4-flash — MoE; 284B model with 1M-token context optimized for fast coding and agents
  • deepseek-ai / deepseek-v4-pro — MoE; scales to 1M-token context with efficient MoE architecture
  • moonshotai / kimi-k2.6 — Multimodal; 1T MoE for long-horizon coding, agentic tool use, and image/video understanding
  • mistral-ai / mistral-medium-3.5-128b — Coding/Agentic; high-performing model for text generation, coding, and agentic use cases
  • stepfun-ai / step-3.7-flash — Coding; sparse MoE multimodal reasoning model for enterprise, agentic, and coding tasks
  • minimaxai / minimax-m3 — Coding; multimodal MoE vision-language model with strong reasoning and tool-calling
  • google / diffusiongemma-26b-a4b-it — Diffusion LLM; diffusion-based 26B parameter LLM for parallel token generation
  • nvidia / nemotron-3-nano-omni-30b-a3b-reasoning — Image-to-Text; omni-modal reasoning model understanding images, video, speech, and text
  • nvidia / nemotron-3.5-content-safety — LLM Safety; multilingual, multimodal model for detecting unsafe and toxic content
  • nvidia / synthetic-video-detector — Broadcast; AI-powered microservice for detecting AI-generated (synthetic) videos
  • nvidia / ising-calibration-1-35b-a3b — Quantum; open VLM for quantum computer calibration chart understanding

The catalog lists 141 models in total, with 77 free endpoints and 43 partner endpoints, spanning publishers including NVIDIA, Meta, Google, Mistral AI, and Qwen. NVIDIA Build Models Catalog See 80+ Free AI Models at NVIDIA Build for the broader model catalog.

DGX Station

NVIDIA Build includes step-by-step playbooks for getting started with DGX Station, including guidance on setting up NemoClaw as a secure personal AI agent. NVIDIA Build Overview


Platform Architecture


Summary

NVIDIA Build consolidates the essential building blocks for AI application development — secure execution via NemoClaw, reusable Agentic Skills, ready-made Blueprint Collection workflows, and free Inference Endpoints — into a single developer hub. NVIDIA Build Overview


NVIDIA's 2026 AI Model Ecosystem

NVIDIA Build serves as the primary distribution point for NVIDIA's open model portfolio. As of 2026, that portfolio spans eight major model families, all available free on Hugging Face, GitHub, and build.nvidia.com under commercial-friendly licenses (Apache 2.0, MIT, or NVIDIA Open Model License). NVIDIA AI Models 2026 Guide

Model FamilyDomainNotable Models
NemotronAgentic LLMsNemotron 3 Ultra, Super, Nano, VoiceChat
PersonaPlexVoice AIPersonaPlex-7B-v1 (full-duplex speech)
CosmosPhysical AI simulationCosmos Predict 2.5, Cosmos Transfer 2.5
GR00THumanoid roboticsGR00T N1.7 (VLA model)
AlpamayoAutonomous vehiclesAlpamayo R1 (open reasoning VLA)
ClaraBiomedical
Earth-2Climate science
Nemotron SpeechASR / RAGReal-time low-latency ASR

NVIDIA also released one of the largest open training datasets in AI history alongside these models, including 10 trillion language training tokens, 500,000 robotics trajectories, 455,000 protein structures, and 100 terabytes of vehicle sensor data. NVIDIA AI Models 2026 Guide

For a full analysis of the model portfolio, rankings, and competitive benchmarks, see NVIDIA AI Models 2026 Guide.

PersonaPlex 7B

Released January 15, 2026, PersonaPlex-7B-v1 is a 7-billion parameter full-duplex speech-to-speech conversational AI. Unlike turn-based voice assistants, it listens and speaks simultaneously using a dual-stream Transformer (Moshi architecture). Key metrics: NVIDIA AI Models 2026 Guide

  • Smooth turn-taking latency: 0.170 seconds
  • User interruption latency: 0.240 seconds
  • FullDuplexBench smooth turn-taking takeover rate: 0.908
  • Outperforms Gemini Live, Qwen 2.5 Omni, and Moshi on conversational dynamics

Model weights are under the NVIDIA Open Model License; code is MIT-licensed. Requires NVIDIA Ampere (A100) or Hopper (H100) architecture GPUs for real-time performance. NVIDIA AI Models 2026 Guide

Nemotron 3

Nemotron 3, announced at GTC March 2026, uses a Hybrid Mamba-Transformer Mixture-of-Experts architecture delivering 5x throughput efficiency (NVFP4 format) on Blackwell chips. It is available through NVIDIA NIM microservices for managed cloud deployment. Notable enterprise adopters include CrowdStrike, ServiceNow, Perplexity, Cursor, Palantir, Salesforce, and Bosch. NVIDIA AI Models 2026 Guide

Physical AI Stack

Build.nvidia.com also hosts NVIDIA's physical AI models: NVIDIA AI Models 2026 Guide

  • GR00T N1.7 — Vision language action model for humanoid robots; commercially viable as of GTC March 2026. Tops MolmoSpaces and RoboArena benchmarks. LG Electronics and NEURA Robotics are early adopters.
  • Cosmos 2.5 — Physical AI simulation platform; generates synthetic training data across diverse conditions. Used by Johnson & Johnson MedTech, Toyota Research Institute, Salesforce, and Uber.
  • Alpamayo R1 — First open reasoning VLA model for autonomous vehicles. Mercedes-Benz is building it into the all-new CLA on the NVIDIA DRIVE platform.

Rubin Hardware Platform

All NVIDIA models ultimately run on NVIDIA hardware. The Rubin platform — NVIDIA's first extreme-codesigned six-chip AI system (GPU + Vera CPU + BlueField-4 DPU), launched at CES 2026 — delivers AI token generation at one-tenth the cost of the previous Blackwell platform. Nemotron 3 is specifically tuned for Rubin, deepening the hardware-software integration for developers building through NVIDIA NIM. NVIDIA AI Models 2026 Guide

Made with CairniExplore public wikis →