POST /openai/v1/chat/completions. The model returns the next assistant message.
Messages and roles
messages is the conversation so far, oldest first. Each message has a role and content.
content is a string, or an array of text parts that are joined in order:
Multi-turn conversations
The API is stateless: it does not store conversations, so each request must include the full history you want the model to see. To continue a conversation, append the assistant’s reply and the next user message, then send the whole list again.Python
400 error with code context_length_exceeded.
Control generation
stop is accepted but has no effect: generation doesn’t stop at stop sequences. See OpenAI compatibility.
See Create a chat completion for every parameter.
Read the response
The reply is inchoices[0].message. Always check finish_reason to see why generation ended:
usage reports the tokens used by the request:
Next steps
- Stream the response to show it as it is generated.
- Adjust reasoning to trade answer quality for latency.
- Get structured output when your code needs to parse the reply.