Skip to main content
To generate text, send a conversation to POST /openai/v1/chat/completions. The model returns the next assistant message.

Messages and roles

messages is the conversation so far, oldest first. Each message has a role and content. content is a string, or an array of text parts that are joined in order:
Only text input is supported.

Multi-turn conversations

The API is stateless: it does not store conversations, so each request must include the full history you want the model to see. To continue a conversation, append the assistant’s reply and the next user message, then send the whole list again.
Python
Longer histories use more input tokens, which count toward rate limits and usage. The API doesn’t truncate a history that’s too long: a request longer than the model accepts returns a 400 error with code context_length_exceeded.

Control generation

stop is accepted but has no effect: generation doesn’t stop at stop sequences. See OpenAI compatibility. See Create a chat completion for every parameter.

Read the response

The reply is in choices[0].message. Always check finish_reason to see why generation ended: usage reports the tokens used by the request:

Next steps