Skip to main content
Reflection models can reason before they answer: they generate reasoning tokens to work through a problem, then write the final answer. More reasoning tends to improve answers to hard problems, such as math, code, and multi-step planning, at the cost of latency and tokens.

Set reasoning effort

Beam-501B-A23B always reasons. Use reasoning_effort to choose how much. It supports these values:
Any other value returns a 400 error with code unsupported_value, and the message lists the values the model accepts. This includes OpenAI effort names that the model doesn’t support, such as none and minimal. An integer effort returns a 400 error with code invalid_type. Start with the default, medium, and raise the effort only if answers aren’t good enough. Lower it to low for simple tasks where latency matters more.

Discover supported efforts

The models endpoints describe each model’s efforts in its reasoning object:
A model that doesn’t announce its efforts has no reasoning object.

Read the reasoning

The model’s reasoning is returned separately from its answer, in message.reasoning_content. The answer is in message.content as usual.
When streaming, reasoning arrives in delta.reasoning_content, before the answer arrives in delta.content.
Reasoning is useful for debugging and for showing progress, but it is not written for end users. Show content as the answer.

Reasoning tokens and limits

Reasoning tokens are counted in usage.completion_tokens and reported separately in usage.completion_tokens_details.reasoning_tokens. They count toward max_completion_tokens and toward rate limits. If max_completion_tokens is too low, the model can run out of tokens while reasoning. The response then has finish_reason: "length" and content may be null. Raise max_completion_tokens or lower reasoning_effort.

Reasoning in multi-turn conversations

Assistant messages in a request accept reasoning_content, so you can send a prior turn back exactly as you received it. During tool calling, return the assistant message with its tool_calls and reasoning_content unchanged.