Skip to main content
Set stream: true to receive the response incrementally instead of waiting for the whole completion. Streaming lets you show text as soon as the model produces it.

Event format

The response is a stream of server-sent events with content type text/event-stream. Each event is a data: line carrying a JSON chat completion chunk, and a completed stream ends with data: [DONE]:
Every chunk in a stream has the same id and created. Each chunk’s choices[0].delta holds the new part of the message:
With reasoning enabled, the model’s reasoning streams in delta.reasoning_content before the answer starts in delta.content, so a loop that prints only content can wait silently at first. Read reasoning_content too if you want to show progress; in Python, use getattr(delta, "reasoning_content", None).
finish_reason is null until the final chunk for the choice, which carries the reason generation ended. Concatenate the content values in order to rebuild the full answer.

Token usage

Set stream_options.include_usage to true to receive usage for the whole request. The stream then sends one more chunk before data: [DONE], with usage set and an empty choices array. Guard for the empty array when you read choices[0].

Padding

To protect against side-channel attacks that infer content from traffic, streamed events have their payloads and timings normalized. Each event carries an obfuscation field of random characters that pads it to a fixed length. Ignore the field’s value. Turn padding off with "stream_options": {"include_obfuscation": false} only where traffic can’t be observed, for example on a loopback connection. Doing so may slightly increase the rate at which chunks arrive.

Streaming tool calls

When the model calls a tool in a streamed response, each tool call arrives in fragments across delta.tool_calls. Every fragment has an index identifying the call; the call’s id, type, and function.name arrive first, then pieces of function.arguments. Accumulate the fragments by index and parse arguments as JSON once finish_reason is tool_calls.
Python

Handling interruptions

If the connection drops before data: [DONE], the response is incomplete. Treat it as failed and retry the request; the API does not resume partial streams.