# NVIDIA API Catalog
The NVIDIA API Catalog is the hosted side of NVIDIA NIM, served from build.nvidia.com. You sign in, generate a key, and call open-weight models on NVIDIA's GPUs for free, with no credit card. The menu is wide: [[NVIDIA Nemotron]], [[Kimi K3]], GLM-5.3, [[DeepSeek V4.1 Flash]], gpt-oss, Gemma, Mistral, plus a long tail of speech, vision, biology and route-optimization models.
Why would NVIDIA give inference away? Because it's a trial, and the terms say so in plain words: without a paid subscription you may use it "for internal testing and evaluation purposes, not in production". The intended path is to prototype against the hosted endpoint, then run the same NIM container on your own GPUs under an NVIDIA AI Enterprise license. NVIDIA earns nothing on the free tokens. It earns on the GPUs and the license that come after.
## What NIM means here
NIM (NVIDIA Inference Microservices) are containers that bundle a model with an optimized inference engine (TensorRT, TensorRT-LLM) and expose standard APIs. The name covers two things:
- **Hosted endpoints** on build.nvidia.com (this note). Labeled "Free Endpoint" in the catalog.
- **Downloadable containers** you run yourself. NVIDIA Developer Program members can self-host them for free for development, testing and research, on up to 2 nodes or 16 GPUs (announced July 2024). Labeled "Download Available".
## API format
- **Base URL**: `https://integrate.api.nvidia.com/v1`
- **Auth**: `Authorization: Bearer $NVIDIA_API_KEY`. Keys start with `nvapi-` and are generated at build.nvidia.com/settings.
- **Format**: every LLM endpoint implements the OpenAI Chat Completions API (`POST /v1/chat/completions`). Any OpenAI SDK client works once you change the base URL and the model name.
- **Model list**: `GET /v1/models` answers without a key.
- **What's missing**: no OpenAI Responses API and no Anthropic Messages API. On 2026-10-03, `/v1/responses` and `/v1/messages` both returned 404, while `/v1/chat/completions` asked for a key.
```bash
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer $NVIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"nvidia/nemotron-3-super-120b-a12b","messages":[{"role":"user","content":"Hello"}]}'
```
## Models
On 2026-10-03, the catalog index (`build.nvidia.com/models.md`) listed 101 entries, including the non-LLM ones, and `/v1/models` returned 81 model IDs. The chat models included `moonshotai/kimi-k3`, `z-ai/glm-5.3`, `z-ai/glm-5.3-flash`, `deepseek-ai/deepseek-v4.1-flash`, `nvidia/nemotron-3-super-120b-a12b`, `nvidia/nemotron-3-ultra-550b-a55b`, `openai/gpt-oss-20b`, `mistralai/mistral-large` and `google/gemma-4-31b-it`.
The catalog rotates, and that's the first thing to know before building anything on it:
- The MiniMax M2.7 page is still up but flagged as deprecated.
- An HN user noticed GLM-5.2 had disappeared from the catalog (August 2026).
- A Reddit developer probing the API found that `/v1/models` over-reports (5 of the 24 models he tried were actually served) and that retired models answer with `410 Gone` and an end-of-life date.
- The API reference still documents models (`z-ai/glm5.1`, `z-ai/glm4.7`, `minimaxai/minimax-m2.7`, `openai/gpt-oss-120b`, `sarvamai/sarvam-m`, `moonshotai/kimi-k2-instruct`) that the live `/v1/models` no longer returns.
So check the live list (and send one test request) before hardcoding a model ID anywhere.
## Free tier, rate limits and pricing
- **Sign-up**: free NVIDIA account, no credit card. People report phone verification, and SMS codes not arriving in some countries.
- **Credits are gone.** NVIDIA used to grant 1,000 credits at sign-up, and up to 5,000 with a business email and a 90-day AI Enterprise trial (NVIDIA forum, September 2024). In April 2025 an NVIDIA moderator wrote that the credit system had been removed, and in September 2025 she confirmed it was replaced by rate limits.
- **Rate limits**: the site's own copy says "Up to 40 rpm" and "10,000 requests per day". Limits vary per model and with the traffic of other users. NVIDIA doesn't publish per-model limits; your current maximum shows in the top right corner of build.nvidia.com. On the free tier there is no official way to get a higher limit.
- **Duration**: an NVIDIA moderator said the trial isn't limited in time (September 2025). The terms still describe the service as "limited use for a limited time" and let NVIDIA pull any pre-release service whenever it wants.
- **In practice**: in August 2026 a forum thread filled up with users stuck on `429 Too Many Requests` for GLM-5.2 for hours, sometimes a full day, even with minutes between calls. Reddit threads complain about slow, queued responses on popular models. Everyone shares the same free GPUs, so the busiest models get the worst service.
- **Pricing**: the hosted endpoints have no paid per-token tier. Production means NVIDIA AI Enterprise: a list price of $4,500 per GPU per year ($1,125 for education and Inception members), on your own hardware or through partners, with a 90-day trial license to start.
## Terms worth reading
The NVIDIA API Trial Terms of Service (version of September 19, 2025) apply to every hosted endpoint:
- Trial use only; neither the API nor its output may be used in production.
- Don't send confidential information, personal data, health data or payment card data.
- Section 3.3: NVIDIA collects your prompts and the generated output "to improve NVIDIA products and services, including AI models" (without identifying specific users), and logs usage for security and abuse monitoring.
- Each model adds its own license on top. MiniMax M2.7, for example, falls under the NVIDIA Software and Model Evaluation license and is marked "for research and development only".
In other words: don't point it at client code, company documents or your private notes.
## Using it with coding agents
- **[[OpenCode]]** ships a built-in NVIDIA provider: run `/connect`, pick NVIDIA, paste the key (or set `NVIDIA_API_KEY`).
- **[[Hermes Agent]]** lists NVIDIA NIM among its supported providers.
- **[[NemoClaw]]** uses Nemotron 3 Super on build.nvidia.com as its default backend.
- **Anything with a custom OpenAI-compatible provider setting**: base URL, key, model ID. That's it.
- **[[Claude Code]] does NOT work out of the box.** Claude Code talks to `ANTHROPIC_BASE_URL` using the Anthropic Messages format (`/v1/messages`), and NVIDIA only exposes Chat Completions. You need a translating proxy such as [[LiteLLM Claude Code Proxy|LiteLLM]] in between, and Anthropic's docs state that it doesn't support routing Claude Code to non-Claude models through a gateway.
## Alternatives
- **[[OpenRouter]]**: one key for hundreds of models, paid by credits. Free models are capped at 20 requests per minute and 50 requests per day, or 1,000 per day once you've bought at least $10 of credits.
- **[[Groq]]**: its own LPU chips, very fast open-model inference, OpenAI-compatible, free tier with per-model limits.
- **[[Cloudflare Workers AI]]**: 10,000 Neurons per day free, then $0.011 per 1,000 Neurons on the Workers Paid plan.
- **Together AI**: no free trial; fully prepaid, $5 minimum.
My take: the NVIDIA API Catalog gives you the widest free menu of open models, which makes it great for trying a new release the week it lands, or for comparing a few models on the same prompt. For anything you'd bill a customer for, the terms rule it out, and the 429s would make the call for you anyway.
## Claims in the viral X post, checked
A long post by @iam_elias1 (May 2, 2026, about 73k views) pushed the "free NVIDIA API" story. Against NVIDIA's own sources:
- **"1,000 free inference credits on signup"**: outdated. NVIDIA removed credits by April 2025.
- **"40 requests per minute"**: matches the site's "Up to 40 rpm", but limits vary per model and with load.
- **"No credit card, no expiry"**: correct per NVIDIA's FAQ and moderator, with the caveat that the terms let NVIDIA stop at any time.
- **Base URL, `nvapi-` keys, OpenAI format**: correct.
- **MiniMax M2.7 specs** (230B total parameters, 256 experts with 8 active per token, 204,800-token context): match NVIDIA's model card, which also lists 10B active parameters. The model is now flagged as deprecated.
- **[[GLM-5.1]], GLM-4.7, Kimi K2, gpt-oss-120b, Sarvam-M**: documented in the API reference, but absent from the live model list today. DeepSeek V3.2, Llama 4 Maverick and Qwen3-Coder appear in neither.
- **"100+ models"**: 101 catalog entries, many of them not LLMs; 81 IDs via `/v1/models`.
- **"Claude Code works without any code changes"**: false (see above).
- **The funnel analysis** (prototype free, then pay for NVIDIA AI Enterprise): consistent with NVIDIA's own wording.
## References
- X post by @iam_elias1 (May 2, 2026), read via the fxtwitter API: https://x.com/iam_elias1/status/2050545769787371710
- NVIDIA API Catalog: https://build.nvidia.com/models
- build.nvidia.com llms.txt (base URL, auth, FAQ): https://build.nvidia.com/llms.txt
- Catalog index in markdown: https://build.nvidia.com/models.md
- Live model list: https://integrate.api.nvidia.com/v1/models
- MiniMax M2.7 model card on NVIDIA: https://build.nvidia.com/minimaxai/minimax-m2.7.md
- NVIDIA NIM LLM API reference: https://docs.api.nvidia.com/nim/reference/llm-apis
- NVIDIA API Trial Terms of Service (v. September 19, 2025): https://assets.ngc.nvidia.com/products/api-catalog/legal/NVIDIA%20API%20Trial%20Terms%20of%20Service.pdf
- NVIDIA blog, Access to NVIDIA NIM Now Available Free to Developer Program Members (July 29, 2024): https://developer.nvidia.com/blog/access-to-nvidia-nim-now-available-free-to-developer-program-members/
- NVIDIA AI Enterprise: https://www.nvidia.com/en-us/data-center/products/ai-enterprise/
- NVIDIA AI Enterprise pricing: https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/pricing.html
- NVIDIA forum, API credits for build.nvidia.com (2024): https://forums.developer.nvidia.com/t/api-credits-for-build-nvidia-com/306633
- NVIDIA forum, Unable to view my account credit (credits removed, 2025): https://forums.developer.nvidia.com/t/unable-to-view-my-account-credit/328471
- NVIDIA forum, "Request More" (+4,000 credits) option (rate limits, trial duration, 2025): https://forums.developer.nvidia.com/t/request-more-4-000-credits-option-on-build-nvidia-com/344567
- NVIDIA forum, API Error 429 (2026): https://forums.developer.nvidia.com/t/api-error-429-help/376113
- NVIDIA forum, rate-limit increase requests (2026): https://forums.developer.nvidia.com/t/credit-rate-limit-increase-request-1-000-5-000-credits-40-200-rpm/379108
- Claude Code docs, LLM gateway protocol: https://code.claude.com/docs/en/llm-gateway-protocol
- Claude Code docs, LLM gateway: https://code.claude.com/docs/en/llm-gateway
- OpenCode docs, providers (NVIDIA): https://opencode.ai/docs/providers/
- OpenRouter docs, limits: https://openrouter.ai/docs/api/reference/limits
- Groq docs, rate limits: https://console.groq.com/docs/rate-limits
- Cloudflare Workers AI pricing: https://developers.cloudflare.com/workers-ai/platform/pricing/
- Together AI docs, credits: https://docs.together.ai/docs/billing
- HN comment, GLM-5.2 removed from the catalog: https://news.ycombinator.com/item?id=49497740
- HN comment by Simon Willison, account and phone verification: https://news.ycombinator.com/item?id=48480481
- HN comment, account verification stuck: https://news.ycombinator.com/item?id=47738452
- HN, Show HN: Free Unlimited Claude Code Usage with Nvidia NIM Models (proxy project): https://news.ycombinator.com/item?id=46917761
- Reddit r/ClaudeCode, code review panel on NVIDIA's free tier (`/v1/models` over-reports, 410 Gone): https://reddit.com/r/ClaudeCode/comments/1vzhrea/
- Reddit r/opencode, Kimi K3 and DeepSeek V4 Pro on NVIDIA NIM (speed complaints): https://reddit.com/r/opencode/comments/1w0sx6u/
- Reddit r/LocalLLM, phone verification issues: https://reddit.com/r/LocalLLM/comments/1w0h1ny/
- Reddit r/better_claw, free LLM API list: https://reddit.com/r/better_claw/comments/1vef1nz/
## Related
- [[NVIDIA Nemotron]]
- [[NemoClaw]]
- [[OpenRouter]]
- [[Groq]]
- [[Cloudflare Workers AI]]
- [[OpenCode]]
- [[Claude Code]]
- [[LiteLLM Claude Code Proxy]]
- [[AI Inference]]
- [[Cheaper Inference]]