문서 · 연동
채팅 완성
POST /v1/chat/completions sends a list of messages and returns the model's reply. Every chat model in the catalog uses this endpoint.
본문은 현재 영어로만 제공됩니다.
curl https://<your-endpoint>/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4.1-mini",
"messages": [{"role": "user", "content": "Hello!"}]
}'Multi-turn conversations and system messages
The API is stateless: send the full conversation history with every request. Put a system message first to set the role and rules.
resp = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is a vector database?"},
{"role": "assistant", "content": "A database optimized for similarity search over embeddings."},
{"role": "user", "content": "Give me two use cases."},
],
temperature=0.3,
max_tokens=800,
)
print(resp.choices[0].message.content)
print(resp.usage)Common parameters
| Parameter | Description |
|---|---|
model | Required. A model ID from the catalog (case-sensitive) |
messages | Required. Array of messages with role system / user / assistant / tool |
max_tokens / max_completion_tokens | Output token cap; reasoning models take max_completion_tokens |
temperature / top_p | Sampling randomness; usually not supported by reasoning models |
stop | Stop sequences |
stream | true to stream over SSE — see Streaming |
tools / tool_choice | Function calling — see Tool calling |
response_format | JSON output — see Structured output |
reasoning_effort | Reasoning depth — see Reasoning |
Parameters are passed through to the model. Unsupported parameters may be ignored or rejected with a 400, depending on the model.
Response
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "gpt-4.1-mini",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Hello! How can I help you today?" },
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 9, "completion_tokens": 10, "total_tokens": 19 }
}| finish_reason | Meaning |
|---|---|
stop | Finished normally |
length | Hit max_tokens; the output is truncated |
tool_calls | The model wants to call tools |
content_filter | Blocked by a safety policy |
usage reports input and output tokens; cost follows the model price.