NVIDIA NIM: 102 Free Model Endpoints on integrate.api.nvidia.com
One API key. One OpenAI-compatible endpoint. 102 models โ including Nemotron 4 340B, Llama 3.3 70B, DeepSeek V4 Flash, Mistral Large, GPT-OSS 120B and GLM-5.2. NVIDIA's NIM free tier is the quietest free AI deal on the internet, and it's perfect for prototyping.
The Big Picture
NVIDIA NIM (NVIDIA Inference Microservices) is NVIDIA's model-serving layer for AI. It wraps foundation models in optimized, production-grade runtimes โ and the hosted API at integrate.api.nvidia.com gives you free access to the same catalog that powers build.nvidia.com. You don't need a GPU, a cluster, or even a credit card. You need one thing: an API key.
Getting the Free API Key
- Go to build.nvidia.com and sign in (GitHub or email works).
- Open any model page and click "Get API Key" โ a personal NVIDIA API key is generated instantly.
- Use it as a bearer token against
https://integrate.api.nvidia.com/v1.
The free tier is designed for prototyping and evaluation โ it's rate-limited, but for hackathons, side projects, and benchmark testing it's genuinely unlimited-feeling. This is the same key infrastructure used by the NGC catalog.
Quickstart โ OpenAI-Compatible curl
Because the endpoint speaks the OpenAI protocol, it works with openai-python, LangChain, LiteLLM, and anything else that speaks OpenAI chat completions:
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer $NVIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-4-340b-instruct",
"messages": [{"role":"user","content":"Write a Python function to detect palindromes."}],
"max_tokens": 1024
}'
model field to switch between coding, reasoning, embedding, vision, translation, and retrieval models โ no other config changes.Standout Models on the Free Tier
| Model ID | Why it matters |
|---|---|
nvidia/nemotron-4-340b-instruct | NVIDIA's flagship 340B instruct model โ strong reasoning and math. |
nvidia/llama-3.1-nemotron-ultra-253b-v1 | Ultra-tier Nemotron built on Llama 3.1 โ the top of NVIDIA's hosted stack. |
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning | Free open multimodal model, also runs locally (see our earlier review). |
meta/llama-3.3-70b-instruct | Meta's latest 70B flagship, served free. |
meta/llama-3.1-70b-instruct / 8b | Reliable general-purpose workhorses. |
deepseek-ai/deepseek-v4-flash / v4-pro | DeepSeek's agent-grade models โ the V4 Flash we benchmarked at 90.8% GPQA. |
mistralai/mistral-large | Mistral's flagship, free on NIM. |
mistralai/mixtral-8x22b-v0.1 | MoE classic with 22B active params. |
openai/gpt-oss-120b / 20b | OpenAI's open-weights GPT-OSS models, hosted by NVIDIA. |
z-ai/glm-5.2 | Our 9.6-rated open model โ 1M context โ available free here too. |
google/gemma-4-31b-it | Google's compact but capable Gemma 4. |
moonshotai/kimi-k2.6 | Kimi's frontier reasoning model on NIM. |
nvidia/embed-qa-4 | Top-tier embedding model for RAG pipelines. |
nvidia/riva-translate-4b-instruct | Speech/translation via Riva. |
Full Catalog (102 models, live-verified)
The complete list currently served on the free tier, in catalog order:
Use Cases
- Prototyping & Hackathons: Test 100+ models side-by-side without provisioning anything.
- Benchmarking: Reproduce our OpenAdapter, DeepSeek, and GLM benchmark runs on the same hosted infra.
- RAG Pipelines: Pair
nvidia/embed-qa-4(embedding) with any instruct model on the same key. - Vision & Multimodal: Llama 3.2 Vision, Gemma 3, Kosmos-2, and Nemotron Nano VL are all free.
- Translation: Riva translate models for multilingual pipelines.
Limits & Caveats
GET /v1/models before shipping code that depends on a specific ID.Invites & Discounts
Two of the models above are also available through paid platforms with active discount codes โ same models, lower price:
OpenAdapter โ Robin & 75+ Models
Get premium access to the Robin model, GLM 5.2, and 75+ other models at 20% off with this invite link.
Z.ai โ GLM-5.2 Subscription
GLM-5.2 (the 1M-context open model on this list) with a 10% discount and extra credits on subscription.
No credit card ยท Key in 60 seconds ยท 102 models on one endpoint