How to hit the Groq API

Groq runs open models on its own inference hardware and exposes them through an OpenAI-compatible API at https://api.groq.com/openai/v1. If you have OpenAI code, swapping the base URL and key is usually enough.

Authentication

Create a key at console.groq.com/keys and send it as Authorization: Bearer <key>. In the examples it's written {{GROQ_API_KEY}}.

Full reference: https://console.groq.com/docs/api-reference

Chat completion

Llama 3.3 70B. The reply is in choices[0].message.content; usage includes timing.

curl https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer {{GROQ_API_KEY}}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-3.3-70b-versatile",
    "messages": [{"role": "user", "content": "Explain rate limiting in two sentences."}]
  }'
Run in PostTaco

Use an OpenAI open-weights model

Same endpoint, different model id.

curl https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer {{GROQ_API_KEY}}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-20b",
    "messages": [{"role": "user", "content": "Write a regex that matches an HTTP status line."}]
  }'
Run in PostTaco

Stream the response

"stream": true returns server-sent events, one chunk per data: line.

curl https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer {{GROQ_API_KEY}}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-3.1-8b-instant",
    "stream": true,
    "messages": [{"role": "user", "content": "Count to ten."}]
  }'
Run in PostTaco

List models

Every model available to your key, with context window sizes.

curl https://api.groq.com/openai/v1/models \
  -H "Authorization: Bearer {{GROQ_API_KEY}}"
Run in PostTaco