How to hit the Hugging Face API

Hugging Face's Inference Providers route one OpenAI-compatible API at https://router.huggingface.co/v1 to many hosting providers (Groq, Together, Cerebras, and others). One Hugging Face token covers all of them.

Authentication

Create a fine-grained token at huggingface.co/settings/tokens with the Make calls to Inference Providers permission. Send it as Authorization: Bearer <token>. In the examples it's written {{HF_TOKEN}}. Listing models needs no token.

Full reference: https://huggingface.co/docs/inference-providers/index

List models and providers

No token needed. Each model lists which providers serve it, with price per million tokens, latency, and throughput.

curl https://router.huggingface.co/v1/models
Run in PostTaco
Example response (trimmed)
{
  "object": "list",
  "data": [
    {
      "id": "openai/gpt-oss-120b",
      "object": "model",
      "created": 1754346786,
      "owned_by": "openai",
      "architecture": {
        "input_modalities": [
          "text"
        ],
        "output_modalities": [
          "text"
        ]
      },
      "providers": [
        {
          "provider": "groq",
          "status": "live",
          "context_length": 131072,
          "pricing": {
            "input": 0.15,
            "output": 0.75
          },
          "is_free": false,
          "supports_tools": true,
          "supports_structured_output": true,
          "first_token_latency_ms": 270,
          "throughput": 427.9487559328392,
          "is_model_author": false
        },
        {
          "provider": "cerebras",
          "status": "live",
          "context_length": 131072,
          "pricing": {
            "input": 0.35,
            "output": 0.75
          },
          "is_free": false,
          "supports_tools": true,
          "supports_structured_output": true,
          "first_token_latency_ms": 166.2,
          "throughput": 1144.740272686181,
          "is_model_author": false
        }
      ]
    }
  ]
}

Chat completion

Any chat model from the list. By default the fastest live provider serves it.

curl https://router.huggingface.co/v1/chat/completions \
  -H "Authorization: Bearer {{HF_TOKEN}}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b",
    "messages": [{"role": "user", "content": "How many Gs are in huggingface?"}]
  }'
Run in PostTaco

Pick a provider or policy

Append :provider (e.g. :groq) or a policy — :cheapest, :fastest, :preferred — to the model id.

curl https://router.huggingface.co/v1/chat/completions \
  -H "Authorization: Bearer {{HF_TOKEN}}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b:cheapest",
    "messages": [{"role": "user", "content": "Summarize what an API gateway does."}]
  }'
Run in PostTaco

Check your token

The account, orgs, and permissions behind a token.

curl https://huggingface.co/api/whoami-v2 \
  -H "Authorization: Bearer {{HF_TOKEN}}"
Run in PostTaco