> ## Documentation Index
> Fetch the complete documentation index at: https://developers.reflection.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Text generation

> Generate a response to a conversation with the Chat Completions API

To generate text, send a conversation to `POST /openai/v1/chat/completions`. The model returns the next assistant message.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.reflection.ai/openai/v1/chat/completions \
    -H "Authorization: Bearer $REFLECTION_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "Beam-501B-A23B",
      "messages": [
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "Why is the sky blue?"}
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.reflection.ai/openai/v1",
      api_key=os.environ["REFLECTION_API_KEY"],
  )

  completion = client.chat.completions.create(
      model="Beam-501B-A23B",
      messages=[
          {"role": "system", "content": "Answer in one sentence."},
          {"role": "user", "content": "Why is the sky blue?"},
      ],
  )
  print(completion.choices[0].message.content)
  ```

  ```typescript TypeScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.reflection.ai/openai/v1",
    apiKey: process.env.REFLECTION_API_KEY,
  });

  const completion = await client.chat.completions.create({
    model: "Beam-501B-A23B",
    messages: [
      { role: "system", content: "Answer in one sentence." },
      { role: "user", content: "Why is the sky blue?" },
    ],
  });
  console.log(completion.choices[0].message.content);
  ```
</CodeGroup>

## Messages and roles

`messages` is the conversation so far, oldest first. Each message has a `role` and `content`.

| Role | Use it for |
| - | - |
| `system` or `developer` | Instructions that set the model's behavior for the whole conversation, such as tone, format, or constraints. Put them first. |
| `user` | Input from the person or application using the model. |
| `assistant` | The model's earlier replies. Include them to continue a conversation. |
| `tool` | The result of a tool call. See [Tool calling](/tool-calling). |

`content` is a string, or an array of text parts that are joined in order:

```json theme={null}
{
  "role": "user",
  "content": [
    {"type": "text", "text": "Summarize this report:"},
    {"type": "text", "text": "<report text>"}
  ]
}
```

Only text input is supported.

## Multi-turn conversations

The API is stateless: it does not store conversations, so each request must include the full history you want the model to see. To continue a conversation, append the assistant's reply and the next user message, then send the whole list again.

```python Python theme={null}
messages = [{"role": "user", "content": "Name a primary color."}]

completion = client.chat.completions.create(
    model="Beam-501B-A23B", messages=messages
)
messages.append(completion.choices[0].message)
messages.append({"role": "user", "content": "Name another one."})

completion = client.chat.completions.create(
    model="Beam-501B-A23B", messages=messages
)
print(completion.choices[0].message.content)
```

Longer histories use more input tokens, which count toward [rate limits](/rate-limits) and usage. The API doesn't truncate a history that's too long: a request longer than the model accepts returns a `400` error with code `context_length_exceeded`.

## Control generation

| Parameter | Effect |
| - | - |
| `max_completion_tokens` | Upper bound on generated tokens, including [reasoning tokens](/reasoning). The prompt plus this value should fit within the model's [context window](/models). Replaces the deprecated `max_tokens`; don't set both. |
| `temperature` | Sampling randomness from `0` to `2`. Lower values are more focused and repeatable. |
| `top_p` | Nucleus sampling, from `0` to `1`. Change this or `temperature`, not both. |
| `frequency_penalty` | From `-2` to `2`. Positive values penalize tokens by how often they have appeared so far, making verbatim repetition less likely. |
| `presence_penalty` | From `-2` to `2`. Positive values penalize tokens that have appeared at all so far, making new topics more likely. |
| `seed` | Best-effort deterministic sampling for repeated identical requests. Not guaranteed. |
| `reasoning_effort` | How much the model reasons before answering. See [Reasoning](/reasoning). |

`stop` is accepted but has no effect: generation doesn't stop at stop sequences. See [OpenAI compatibility](/openai-compatibility#stop-sequences).

See [Create a chat completion](/api-reference/chat/create-a-chat-completion) for every parameter.

## Read the response

The reply is in `choices[0].message`. Always check `finish_reason` to see why generation ended:

| `finish_reason` | Meaning |
| - | - |
| `stop` | The model finished. |
| `length` | The `max_completion_tokens` limit was reached. The output is cut off, and `content` can be `null` if the limit was reached during reasoning. |
| `tool_calls` | The model called a tool. `content` is `null` when the model made only tool calls. See [Tool calling](/tool-calling). |
| `content_filter` | Content was omitted by a content filter. |

`usage` reports the tokens used by the request:

```json theme={null}
"usage": {
  "prompt_tokens": 23,
  "completion_tokens": 45,
  "total_tokens": 68,
  "prompt_tokens_details": { "cached_tokens": 0 },
  "completion_tokens_details": { "reasoning_tokens": 14 }
}
```

## Next steps

* [Stream](/streaming) the response to show it as it is generated.
* [Adjust reasoning](/reasoning) to trade answer quality for latency.
* [Get structured output](/structured-outputs) when your code needs to parse the reply.
