stream: true to receive the response incrementally instead of waiting for the whole completion. Streaming lets you show text as soon as the model produces it.
Event format
The response is a stream of server-sent events with content typetext/event-stream. Each event is a data: line carrying a JSON chat completion chunk, and a completed stream ends with data: [DONE]:
id and created. Each chunk’s choices[0].delta holds the new part of the message:
With reasoning enabled, the model’s reasoning streams in
delta.reasoning_content before the answer starts in delta.content, so a loop that prints only content can wait silently at first. Read reasoning_content too if you want to show progress; in Python, use getattr(delta, "reasoning_content", None).finish_reason is null until the final chunk for the choice, which carries the reason generation ended. Concatenate the content values in order to rebuild the full answer.
Token usage
Setstream_options.include_usage to true to receive usage for the whole request. The stream then sends one more chunk before data: [DONE], with usage set and an empty choices array. Guard for the empty array when you read choices[0].
Padding
To protect against side-channel attacks that infer content from traffic, streamed events have their payloads and timings normalized. Each event carries anobfuscation field of random characters that pads it to a fixed length. Ignore the field’s value.
Turn padding off with "stream_options": {"include_obfuscation": false} only where traffic can’t be observed, for example on a loopback connection. Doing so may slightly increase the rate at which chunks arrive.
Streaming tool calls
When the model calls a tool in a streamed response, each tool call arrives in fragments acrossdelta.tool_calls. Every fragment has an index identifying the call; the call’s id, type, and function.name arrive first, then pieces of function.arguments. Accumulate the fragments by index and parse arguments as JSON once finish_reason is tool_calls.
Python
Handling interruptions
If the connection drops beforedata: [DONE], the response is incomplete. Treat it as failed and retry the request; the API does not resume partial streams.