NVIDIA NIM Inference Providers
Overview
Inference Providers are third-party platforms listed within the NVIDIA Build models catalog that host and serve NVIDIA NIM (Inference Microservices)-compatible models. Alongside NVIDIA's own Free Endpoint and Downloadable options, these partner providers give developers the flexibility to run models on infrastructure beyond NVIDIA's direct hosting. NVIDIA Build Models Catalog
The catalog exposes inference providers as a first-class filter dimension, sitting alongside filters for use case, publisher, NIM container GPU type, and access tier (free endpoint vs. downloadable). NVIDIA Build Models Catalog
Providers Listed in the Catalog
The following inference providers appear as selectable filters in the NVIDIA Build models catalog: NVIDIA Build Models Catalog
| Provider | Models Available |
|---|---|
| Deepinfra | 34 |
| OpenRouter | 29 |
| Together AI | 23 |
| GMI Cloud | 15 |
| Lightning AI | 7 |
These figures reflect the catalog snapshot in the source and may change as new models are added. The catalog notes "Show more" beyond the five listed above, indicating additional providers may exist. NVIDIA Build Models Catalog
Relationship to Access Tiers
The NVIDIA Build Inference Endpoints catalog distinguishes models by access tier:
- Free Endpoint — 77 models hosted directly by NVIDIA at no cost.
- Partner Endpoint — 43 models served through partner/inference-provider infrastructure.
- Download Available — 108 models that can be pulled as NIM containers for self-hosting.
NVIDIA Build Models Catalog
Inference providers primarily fulfil the Partner Endpoint tier, giving publishers an alternative path to reach developers without requiring NVIDIA-hosted infrastructure. This tiering is explored further in Model Tiering Strategy and Free vs. Paid AI Model Access.
Publisher Landscape
Models served through inference providers come from a broad set of publishers. The catalog's publisher filter lists the following top contributors across all access tiers: NVIDIA Build Models Catalog
| Publisher | Models in Catalog |
|---|---|
| NVIDIA | 76 |
| Meta | 11 |
| 6 | |
| Mistral AI | 6 |
| Qwen | 5 |
Other publishers represented in the catalog include Z.ai (publisher of GLM-5.2), DeepSeek AI, Moonshotai, Stepfun-ai, Minimaxai, and Resemble.AI, among others. NVIDIA Build Models Catalog
GPU Infrastructure
Models delivered through the catalog — whether via NVIDIA endpoints or inference providers — are optimized for specific GPU hardware. The catalog's NIM Container GPU filter lists: NVIDIA Build Models Catalog
- H100 80GB HBM3 — 16 models
- B200 — 15 models
- L40S — 15 models
- H200 — 14 models
- A100 SXM4 80GB — 13 models
For a deeper look at GPU compatibility requirements, see NVIDIA NIM GPU Compatibility.
Use Cases Covered
Inference providers collectively support the full breadth of use cases cataloged in NIM Use-Case Taxonomy, including: NVIDIA Build Models Catalog
- Drug Discovery (13 models)
- Image-to-Text (10 models)
- Retrieval Augmented Generation (9 models)
- Speech-to-Text (9 models)
- Code Generation (8 models)
Related Pages
- NVIDIA Build — The platform hosting the models catalog.
- NVIDIA NIM (Inference Microservices) — The underlying microservice standard models conform to.
- NVIDIA Build Inference Endpoints — Details on endpoint types and access.
- OpenAI-Compatible API — The API standard used to query models across providers.
- NVIDIA NIM Framework Integrations — How NIM models integrate with developer frameworks.
- 80+ Free AI Models at NVIDIA Build — Guide to the free-tier model offering.
- NVIDIA NIM Free Models Guide — Broader guide to accessing NIM models at no cost.