curl -N https://api.reflection.ai/openai/v1/chat/completions \
-H "Authorization: Bearer $REFLECTION_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Beam-501B-A23B",
"messages": [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "Why is the sky blue?"}
],
"stream": true,
"stream_options": {"include_usage": true}
}'{
"id": "chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f",
"object": "chat.completion",
"created": 1791198000,
"model": "Beam-501B-A23B",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Air molecules scatter short blue wavelengths of sunlight far more than longer red ones, so blue light reaches your eyes from every part of the sky.",
"refusal": null,
"reasoning_content": "The user wants a one-sentence explanation of Rayleigh scattering."
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 45,
"total_tokens": 68,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 14
}
}
}Create a chat completion
Creates a model response for the given chat conversation. Learn more in the text generation guide.
Returns a chat completion object or, if the request is streamed, a sequence of server-sent events; chat completion chunks arrive as data: lines, and a completed stream ends with data: [DONE].
curl -N https://api.reflection.ai/openai/v1/chat/completions \
-H "Authorization: Bearer $REFLECTION_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Beam-501B-A23B",
"messages": [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "Why is the sky blue?"}
],
"stream": true,
"stream_options": {"include_usage": true}
}'{
"id": "chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f",
"object": "chat.completion",
"created": 1791198000,
"model": "Beam-501B-A23B",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Air molecules scatter short blue wavelengths of sunlight far more than longer red ones, so blue light reaches your eyes from every part of the sky.",
"refusal": null,
"reasoning_content": "The user wants a one-sentence explanation of Rayleigh scattering."
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 45,
"total_tokens": 68,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 14
}
}
}Streaming responses
Withstream: true, the response is a sequence of server-sent events
instead of one JSON object. Each event is a data: line holding a
chat.completion.chunk. Read the generated text from
choices[0].delta.content on each chunk, not from message, and
concatenate the pieces in order. The final chunk for each choice
carries a finish_reason. With stream_options.include_usage, one
more chunk with empty choices reports usage for the whole
request. A completed stream ends with data: [DONE].Authorizations
A Reflection API key, sent in the Authorization header as Bearer <API key>. Create a key for your project from the Reflection platform.
Body
Parameters not listed here are not supported.
ID of the model to use. Use List models to see the models available to you.
1A list of messages comprising the conversation so far.
1Show child attributes
Show child attributes
Whether to stream the model response data to the client as it is generated using server-sent events. See the streaming guide for how to handle the streamed events.
Options for streaming response. Only set this when you set stream: true.
Show child attributes
Show child attributes
The sampling temperature, from 0 to 2. Higher values such as 0.8 make the output more varied; lower values such as 0.2 make it more focused and consistent. Adjust this or top_p, not both.
0 <= x <= 2The nucleus sampling threshold, from 0 to 1. The model samples only from the most likely tokens whose probabilities add up to top_p, so 0.1 limits it to the top 10% of probability mass. Adjust this or temperature, not both.
0 <= x <= 1The maximum number of tokens that can be generated in the chat completion, including reasoning tokens.
Deprecated in favor of max_completion_tokens. The two cannot be set together.
x >= 1An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens. The prompt plus this value should fit within the model's context_length. Cannot be set together with max_tokens.
x >= 1Sequences at which to stop generating further tokens: a non-empty string, or an array of 1 to 4 non-empty strings. Accepted but has no effect: generation does not stop at these sequences.
1If specified, the system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed.
Constrains effort on reasoning: a level name of 1 to 64 characters, such as low, medium, high, xhigh, or max. A model that lists reasoning.supported_efforts accepts only those values. When omitted, the model's reasoning.default_effort applies if it lists one; otherwise the model applies its own effort. See the reasoning guide.
1 - 64A list of tools the model may call, each a function tool with a name, an optional description and JSON Schema parameters. A call the model makes is returned in message.tool_calls with finish_reason: tool_calls.
128Show child attributes
Show child attributes
Controls which (if any) tool is called by the model: none, auto, required, or an object naming one function the model must call. Allowed only when tools is supplied.
none, auto, required Whether to enable parallel function calling during tool use. Allowed only when tools is supplied.
How many chat completion choices to generate for each input message. Only 1 is supported.
1 Whether to return log probabilities of the output tokens. Only false is supported.
false Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
-2 <= x <= 2Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
-2 <= x <= 2Modifies the likelihood of specified tokens appearing in the completion. Not supported: only an empty object is accepted.
An object specifying the format that the model must output, selected by its type: plain text, a valid JSON object (json_object), or JSON matching the schema in json_schema (json_schema).
- Option 1
- Option 2
- Option 3
Show child attributes
Show child attributes
Specifies the processing type used for serving the request. Only auto and default are supported. Both select standard processing.
auto, default Whether to store the output of this chat completion request. Only false is supported. Completions are not stored.
false Output types that you would like the model to generate. Only ["text"] is supported.
1 elementtext Constrains the verbosity of the model's response. Only medium is supported.
medium Response
A chat completion, or a stream of chat completion chunks when stream is true.
Represents a chat completion response returned by the model, based on the provided input.
A unique identifier for the chat completion.
The object type, which is always chat.completion.
chat.completion The Unix timestamp (in seconds) of when the chat completion was created.
The model used for the chat completion.
A list of chat completion choices.
Show child attributes
Show child attributes
Usage statistics for the completion request.
Show child attributes
Show child attributes