Hugging Face's Inference Providers route one OpenAI-compatible API at https://router.huggingface.co/v1 to many hosting providers (Groq, Together, Cerebras, and others). One Hugging Face token covers all of them.
Create a fine-grained token at huggingface.co/settings/tokens with the Make calls to Inference Providers permission. Send it as Authorization: Bearer <token>. In the examples it's written {{HF_TOKEN}}. Listing models needs no token.
Full reference: https://huggingface.co/docs/inference-providers/index
No token needed. Each model lists which providers serve it, with price per million tokens, latency, and throughput.
curl https://router.huggingface.co/v1/models
{
"object": "list",
"data": [
{
"id": "openai/gpt-oss-120b",
"object": "model",
"created": 1754346786,
"owned_by": "openai",
"architecture": {
"input_modalities": [
"text"
],
"output_modalities": [
"text"
]
},
"providers": [
{
"provider": "groq",
"status": "live",
"context_length": 131072,
"pricing": {
"input": 0.15,
"output": 0.75
},
"is_free": false,
"supports_tools": true,
"supports_structured_output": true,
"first_token_latency_ms": 270,
"throughput": 427.9487559328392,
"is_model_author": false
},
{
"provider": "cerebras",
"status": "live",
"context_length": 131072,
"pricing": {
"input": 0.35,
"output": 0.75
},
"is_free": false,
"supports_tools": true,
"supports_structured_output": true,
"first_token_latency_ms": 166.2,
"throughput": 1144.740272686181,
"is_model_author": false
}
]
}
]
}
Any chat model from the list. By default the fastest live provider serves it.
curl https://router.huggingface.co/v1/chat/completions \
-H "Authorization: Bearer {{HF_TOKEN}}" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b",
"messages": [{"role": "user", "content": "How many Gs are in huggingface?"}]
}'
Append :provider (e.g. :groq) or a policy — :cheapest, :fastest, :preferred — to the model id.
curl https://router.huggingface.co/v1/chat/completions \
-H "Authorization: Bearer {{HF_TOKEN}}" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b:cheapest",
"messages": [{"role": "user", "content": "Summarize what an API gateway does."}]
}'
The account, orgs, and permissions behind a token.
curl https://huggingface.co/api/whoami-v2 \
-H "Authorization: Bearer {{HF_TOKEN}}"