> ## Documentation Index
> Fetch the complete documentation index at: https://developers.reflection.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Receive a chat completion as it is generated, with server-sent events

Set `stream: true` to receive the response incrementally instead of waiting for the whole completion. Streaming lets you show text as soon as the model produces it.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.reflection.ai/openai/v1/chat/completions \
    -H "Authorization: Bearer $REFLECTION_API_KEY" \
    -H "Content-Type: application/json" \
    -N \
    -d '{
      "model": "Beam-501B-A23B",
      "messages": [{"role": "user", "content": "Why is the sky blue?"}],
      "stream": true,
      "stream_options": {"include_usage": true}
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.reflection.ai/openai/v1",
      api_key=os.environ["REFLECTION_API_KEY"],
  )

  stream = client.chat.completions.create(
      model="Beam-501B-A23B",
      messages=[{"role": "user", "content": "Why is the sky blue?"}],
      stream=True,
      stream_options={"include_usage": True},
  )

  for chunk in stream:
      if chunk.choices:
          delta = chunk.choices[0].delta
          if delta.content:
              print(delta.content, end="", flush=True)
      if chunk.usage:
          print(f"\n\nTotal tokens: {chunk.usage.total_tokens}")
  ```

  ```typescript TypeScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.reflection.ai/openai/v1",
    apiKey: process.env.REFLECTION_API_KEY,
  });

  const stream = await client.chat.completions.create({
    model: "Beam-501B-A23B",
    messages: [{ role: "user", content: "Why is the sky blue?" }],
    stream: true,
    stream_options: { include_usage: true },
  });

  for await (const chunk of stream) {
    const content = chunk.choices[0]?.delta?.content;
    if (content) process.stdout.write(content);
    if (chunk.usage) console.log(`\n\nTotal tokens: ${chunk.usage.total_tokens}`);
  }
  ```
</CodeGroup>

## Event format

The response is a stream of [server-sent events](https://html.spec.whatwg.org/multipage/server-sent-events.html) with content type `text/event-stream`. Each event is a `data:` line carrying a JSON chat completion chunk, and a completed stream ends with `data: [DONE]`:

```text theme={null}
data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"wGMmHWX0D2a"}

data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"reasoning_content":"The user wants a one-sentence explanation of Rayleigh scattering."},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"lCav4U"}

data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"content":"Air molecules scatter short blue wavelengths of sunlight far more"},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"cVmEhIa5zitBUV3O"}

data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"content":" than longer red ones, so blue light reaches your eyes from every"},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"6WhBZTNBLlolaThi"}

data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"content":" part of the sky."},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"cGv4yGc4Dx-S3LMCyfoevWzBWsxtzk7cCIthcSCHxpJ4yLTophhyxeC6O2Fas9Q-"}

data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"kxtEv0FWmRsjv6IHiQoowTE2U-ZgKvc_EmkfM2CwbV3xz_mxCobC3dFUgFzknbzBzZunVlqbQG1jCUa4YFxVcRtdghX"}

data: {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":{"prompt_tokens":23,"completion_tokens":45,"total_tokens":68,"completion_tokens_details":{"reasoning_tokens":14}},"obfuscation":"QkHVHPHgk3xZMnPQqH_DnJvrzXV"}

data: [DONE]
```

Every chunk in a stream has the same `id` and `created`. Each chunk's `choices[0].delta` holds the new part of the message:

| Delta field | Contains |
| - | - |
| `role` | `assistant`, on the first chunk. |
| `content` | The next piece of the answer. |
| `reasoning_content` | The next piece of the model's [reasoning](/reasoning), reported separately from `content`. |
| `tool_calls` | Fragments of [tool calls](/tool-calling). See [Streaming tool calls](#streaming-tool-calls). |

<Note>
  With reasoning enabled, the model's [reasoning](/reasoning) streams in `delta.reasoning_content` before the answer starts in `delta.content`, so a loop that prints only `content` can wait silently at first. Read `reasoning_content` too if you want to show progress; in Python, use `getattr(delta, "reasoning_content", None)`.
</Note>

`finish_reason` is `null` until the final chunk for the choice, which carries the reason generation ended. Concatenate the `content` values in order to rebuild the full answer.

## Token usage

Set `stream_options.include_usage` to `true` to receive usage for the whole request. The stream then sends one more chunk before `data: [DONE]`, with `usage` set and an empty `choices` array. Guard for the empty array when you read `choices[0]`.

## Padding

To protect against side-channel attacks that infer content from traffic, streamed events have their payloads and timings normalized. Each event carries an `obfuscation` field of random characters that pads it to a fixed length. Ignore the field's value.

Turn padding off with `"stream_options": {"include_obfuscation": false}` only where traffic can't be observed, for example on a loopback connection. Doing so may slightly increase the rate at which chunks arrive.

## Streaming tool calls

When the model calls a tool in a streamed response, each tool call arrives in fragments across `delta.tool_calls`. Every fragment has an `index` identifying the call; the call's `id`, `type`, and `function.name` arrive first, then pieces of `function.arguments`. Accumulate the fragments by `index` and parse `arguments` as JSON once `finish_reason` is `tool_calls`.

```python Python theme={null}
tool_calls = {}

for chunk in stream:
    if not chunk.choices:
        continue
    for fragment in chunk.choices[0].delta.tool_calls or []:
        call = tool_calls.setdefault(
            fragment.index, {"id": None, "name": "", "arguments": ""}
        )
        if fragment.id:
            call["id"] = fragment.id
        if fragment.function and fragment.function.name:
            call["name"] += fragment.function.name
        if fragment.function and fragment.function.arguments:
            call["arguments"] += fragment.function.arguments
```

## Handling interruptions

If the connection drops before `data: [DONE]`, the response is incomplete. Treat it as failed and retry the request; the API does not resume partial streams.
