๐Ÿ—“๏ธ 2026-08-03 Free API โฑ 6 min read

NVIDIA NIM: 102 Free Model Endpoints on integrate.api.nvidia.com

One API key. One OpenAI-compatible endpoint. 102 models โ€” including Nemotron 4 340B, Llama 3.3 70B, DeepSeek V4 Flash, Mistral Large, GPT-OSS 120B and GLM-5.2. NVIDIA's NIM free tier is the quietest free AI deal on the internet, and it's perfect for prototyping.

NVIDIA NIM Free API LLM OpenAI-Compatible

The Big Picture

NVIDIA NIM (NVIDIA Inference Microservices) is NVIDIA's model-serving layer for AI. It wraps foundation models in optimized, production-grade runtimes โ€” and the hosted API at integrate.api.nvidia.com gives you free access to the same catalog that powers build.nvidia.com. You don't need a GPU, a cluster, or even a credit card. You need one thing: an API key.

๐Ÿ’ก Verified today: I hit the catalog endpoint live and pulled the model list โ€” 102 models are currently served on the free tier. This article's numbers are from that live pull, not a press release.

Getting the Free API Key

  1. Go to build.nvidia.com and sign in (GitHub or email works).
  2. Open any model page and click "Get API Key" โ€” a personal NVIDIA API key is generated instantly.
  3. Use it as a bearer token against https://integrate.api.nvidia.com/v1.

The free tier is designed for prototyping and evaluation โ€” it's rate-limited, but for hackathons, side projects, and benchmark testing it's genuinely unlimited-feeling. This is the same key infrastructure used by the NGC catalog.

Quickstart โ€” OpenAI-Compatible curl

Because the endpoint speaks the OpenAI protocol, it works with openai-python, LangChain, LiteLLM, and anything else that speaks OpenAI chat completions:

curl https://integrate.api.nvidia.com/v1/chat/completions \
  -H "Authorization: Bearer $NVIDIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/nemotron-4-340b-instruct",
    "messages": [{"role":"user","content":"Write a Python function to detect palindromes."}],
    "max_tokens": 1024
  }'
One endpoint, 102 models. Swap the model field to switch between coding, reasoning, embedding, vision, translation, and retrieval models โ€” no other config changes.

Standout Models on the Free Tier

Model IDWhy it matters
nvidia/nemotron-4-340b-instructNVIDIA's flagship 340B instruct model โ€” strong reasoning and math.
nvidia/llama-3.1-nemotron-ultra-253b-v1Ultra-tier Nemotron built on Llama 3.1 โ€” the top of NVIDIA's hosted stack.
nvidia/nemotron-3-nano-omni-30b-a3b-reasoningFree open multimodal model, also runs locally (see our earlier review).
meta/llama-3.3-70b-instructMeta's latest 70B flagship, served free.
meta/llama-3.1-70b-instruct / 8bReliable general-purpose workhorses.
deepseek-ai/deepseek-v4-flash / v4-proDeepSeek's agent-grade models โ€” the V4 Flash we benchmarked at 90.8% GPQA.
mistralai/mistral-largeMistral's flagship, free on NIM.
mistralai/mixtral-8x22b-v0.1MoE classic with 22B active params.
openai/gpt-oss-120b / 20bOpenAI's open-weights GPT-OSS models, hosted by NVIDIA.
z-ai/glm-5.2Our 9.6-rated open model โ€” 1M context โ€” available free here too.
google/gemma-4-31b-itGoogle's compact but capable Gemma 4.
moonshotai/kimi-k2.6Kimi's frontier reasoning model on NIM.
nvidia/embed-qa-4Top-tier embedding model for RAG pipelines.
nvidia/riva-translate-4b-instructSpeech/translation via Riva.

Full Catalog (102 models, live-verified)

The complete list currently served on the free tier, in catalog order:

Use Cases

Limits & Caveats

Read the fine print: the free tier is rate-limited and intended for development, not production traffic. For production, NVIDIA's paid NIM API or self-hosted NIM microservices (via NVIDIA AI Enterprise) are the path. Also note: some models rotate in and out of the catalog, so re-check GET /v1/models before shipping code that depends on a specific ID.

Invites & Discounts

Two of the models above are also available through paid platforms with active discount codes โ€” same models, lower price:

20% OFF ยท Invite

OpenAdapter โ€” Robin & 75+ Models

Get premium access to the Robin model, GLM 5.2, and 75+ other models at 20% off with this invite link.

Invite: BDPBCR3R
Claim 20% Off โ†’
Plans from $7/mo (Go) to $79/mo (Max).
10% Discount + Credits

Z.ai โ€” GLM-5.2 Subscription

GLM-5.2 (the 1M-context open model on this list) with a 10% discount and extra credits on subscription.

Code: R0K78RJKNW
Subscribe with Discount โ†’
$0.77/1M input tokens on the API.
๐Ÿค Disclosure: these are affiliate/invite links โ€” I may earn a referral credit if you sign up, at no extra cost to you. They help keep this blog free.
Get Your Free NVIDIA API Key โ†’

No credit card ยท Key in 60 seconds ยท 102 models on one endpoint


#NVIDIA #NIM #FreeAPI #LLM #OpenAICompatible #Nemotron #Llama #DeepSeek #GPT-OSS #GLM