> ## Documentation Index
> Fetch the complete documentation index at: https://developers.reflection.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Reasoning

> Control how much the model reasons before it answers, and read its reasoning

Reflection models can reason before they answer: they generate reasoning tokens to work through a problem, then write the final answer. More reasoning tends to improve answers to hard problems, such as math, code, and multi-step planning, at the cost of latency and tokens.

## Set reasoning effort

`Beam-501B-A23B` always reasons. Use `reasoning_effort` to choose how much. It supports these values:

| Value | Behavior |
| - | - |
| `low` | Less reasoning, for lower latency. |
| `medium` | A balance of quality and latency. |
| `high` | More reasoning, for harder problems. |
| `xhigh` | More reasoning than `high`. |
| `max` | The most reasoning, for the hardest problems. |
| Omitted | `medium`, the model's default effort. |

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.reflection.ai/openai/v1/chat/completions \
    -H "Authorization: Bearer $REFLECTION_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "Beam-501B-A23B",
      "messages": [{"role": "user", "content": "How many primes are there below 100?"}],
      "reasoning_effort": "high"
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.reflection.ai/openai/v1",
      api_key=os.environ["REFLECTION_API_KEY"],
  )

  completion = client.chat.completions.create(
      model="Beam-501B-A23B",
      messages=[{"role": "user", "content": "How many primes are there below 100?"}],
      reasoning_effort="high",
  )

  message = completion.choices[0].message
  print("Reasoning:", getattr(message, "reasoning_content", None))
  print("Answer:", message.content)
  ```

  ```typescript TypeScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.reflection.ai/openai/v1",
    apiKey: process.env.REFLECTION_API_KEY,
  });

  const completion = await client.chat.completions.create({
    model: "Beam-501B-A23B",
    messages: [{ role: "user", content: "How many primes are there below 100?" }],
    reasoning_effort: "high",
  });

  const message = completion.choices[0].message as typeof completion.choices[0]["message"] & {
    reasoning_content?: string | null;
  };
  console.log("Reasoning:", message.reasoning_content);
  console.log("Answer:", message.content);
  ```
</CodeGroup>

Any other value returns a `400` error with code `unsupported_value`, and the message lists the values the model accepts. This includes OpenAI effort names that the model doesn't support, such as `none` and `minimal`. An integer effort returns a `400` error with code `invalid_type`.

Start with the default, `medium`, and raise the effort only if answers aren't good enough. Lower it to `low` for simple tasks where latency matters more.

## Discover supported efforts

The [models endpoints](/models) describe each model's efforts in its `reasoning` object:

```json theme={null}
"reasoning": {
  "supported_efforts": ["max", "xhigh", "high", "medium", "low"],
  "default_effort": "medium",
  "mandatory": true
}
```

| Field | Description |
| - | - |
| `supported_efforts` | The `reasoning_effort` values the model accepts, highest effort first. |
| `default_effort` | The effort applied when a request omits `reasoning_effort`. Omitted when the model chooses its own. |
| `mandatory` | Whether reasoning is always on. When `true`, no `reasoning_effort` value turns reasoning off. |

A model that doesn't announce its efforts has no `reasoning` object.

## Read the reasoning

The model's reasoning is returned separately from its answer, in `message.reasoning_content`. The answer is in `message.content` as usual.

```json theme={null}
"message": {
  "role": "assistant",
  "content": "There are 25 primes below 100.",
  "refusal": null,
  "reasoning_content": "List the primes: 2, 3, 5, 7, 11, ..."
}
```

When [streaming](/streaming), reasoning arrives in `delta.reasoning_content`, before the answer arrives in `delta.content`.

<Tip>
  Reasoning is useful for debugging and for showing progress, but it is not written for end users. Show `content` as the answer.
</Tip>

## Reasoning tokens and limits

Reasoning tokens are counted in `usage.completion_tokens` and reported separately in `usage.completion_tokens_details.reasoning_tokens`. They count toward `max_completion_tokens` and toward [rate limits](/rate-limits).

If `max_completion_tokens` is too low, the model can run out of tokens while reasoning. The response then has `finish_reason: "length"` and `content` may be `null`. Raise `max_completion_tokens` or lower `reasoning_effort`.

## Reasoning in multi-turn conversations

Assistant messages in a request accept `reasoning_content`, so you can send a prior turn back exactly as you received it. During [tool calling](/tool-calling), return the assistant message with its `tool_calls` and `reasoning_content` unchanged.

```json theme={null}
{
  "role": "assistant",
  "content": null,
  "reasoning_content": "I need the current weather before I can answer...",
  "tool_calls": [ ... ]
}
```
