Set reasoning effort
Beam-501B-A23B always reasons. Use reasoning_effort to choose how much. It supports these values:
400 error with code unsupported_value, and the message lists the values the model accepts. This includes OpenAI effort names that the model doesn’t support, such as none and minimal. An integer effort returns a 400 error with code invalid_type.
Start with the default, medium, and raise the effort only if answers aren’t good enough. Lower it to low for simple tasks where latency matters more.
Discover supported efforts
The models endpoints describe each model’s efforts in itsreasoning object:
A model that doesn’t announce its efforts has no
reasoning object.
Read the reasoning
The model’s reasoning is returned separately from its answer, inmessage.reasoning_content. The answer is in message.content as usual.
delta.reasoning_content, before the answer arrives in delta.content.
Reasoning tokens and limits
Reasoning tokens are counted inusage.completion_tokens and reported separately in usage.completion_tokens_details.reasoning_tokens. They count toward max_completion_tokens and toward rate limits.
If max_completion_tokens is too low, the model can run out of tokens while reasoning. The response then has finish_reason: "length" and content may be null. Raise max_completion_tokens or lower reasoning_effort.
Reasoning in multi-turn conversations
Assistant messages in a request acceptreasoning_content, so you can send a prior turn back exactly as you received it. During tool calling, return the assistant message with its tool_calls and reasoning_content unchanged.