Groq runs open models on its own inference hardware and exposes them through an OpenAI-compatible API at https://api.groq.com/openai/v1. If you have OpenAI code, swapping the base URL and key is usually enough.
Create a key at console.groq.com/keys and send it as Authorization: Bearer <key>. In the examples it's written {{GROQ_API_KEY}}.
Full reference: https://console.groq.com/docs/api-reference
Llama 3.3 70B. The reply is in choices[0].message.content; usage includes timing.
curl https://api.groq.com/openai/v1/chat/completions \
-H "Authorization: Bearer {{GROQ_API_KEY}}" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.3-70b-versatile",
"messages": [{"role": "user", "content": "Explain rate limiting in two sentences."}]
}'
Same endpoint, different model id.
curl https://api.groq.com/openai/v1/chat/completions \
-H "Authorization: Bearer {{GROQ_API_KEY}}" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-20b",
"messages": [{"role": "user", "content": "Write a regex that matches an HTTP status line."}]
}'
"stream": true returns server-sent events, one chunk per data: line.
curl https://api.groq.com/openai/v1/chat/completions \
-H "Authorization: Bearer {{GROQ_API_KEY}}" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.1-8b-instant",
"stream": true,
"messages": [{"role": "user", "content": "Count to ten."}]
}'
Every model available to your key, with context window sizes.
curl https://api.groq.com/openai/v1/models \
-H "Authorization: Bearer {{GROQ_API_KEY}}"