NVIDIA hosts models from build.nvidia.com behind an OpenAI-compatible API at https://integrate.api.nvidia.com/v1, including its own Nemotron family. The model list is public.
Create a key at build.nvidia.com. Send it as Authorization: Bearer <key>. Listing models needs no key. In the examples it's written {{NVIDIA_API_KEY}}.
Full reference: https://docs.api.nvidia.com/nim/reference/llm-apis
No key needed.
curl https://integrate.api.nvidia.com/v1/models
{
"object": "list",
"data": [
{
"id": "01-ai/yi-large",
"object": "model",
"created": 735790403,
"owned_by": "01-ai"
},
{
"id": "adept/fuyu-8b",
"object": "model",
"created": 735790403,
"owned_by": "adept"
}
]
}
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer {{NVIDIA_API_KEY}}" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3-super-120b-a12b", "messages": [{"role": "user", "content": "Explain what an API rate limit is in two sentences."}]}'
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer {{NVIDIA_API_KEY}}" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3-ultra-550b-a55b", "messages": [{"role": "user", "content": "What is a Mixture-of-Experts model?"}]}'
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer {{NVIDIA_API_KEY}}" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3.5-lightning-30b-a3b", "messages": [{"role": "user", "content": "Give me three names for a taco truck."}]}'