Skip to main content
POST

Streaming responses

With stream: true, the response is a sequence of server-sent events instead of one JSON object. Each event is a data: line holding a chat.completion.chunk. Read the generated text from choices[0].delta.content on each chunk, not from message, and concatenate the pieces in order. The final chunk for each choice carries a finish_reason. With stream_options.include_usage, one more chunk with empty choices reports usage for the whole request. A completed stream ends with data: [DONE].

Authorizations

Authorization
string
header
required

A Reflection API key, sent in the Authorization header as Bearer <API key>. Create a key for your project from the Reflection platform.

Body

application/json

Parameters not listed here are not supported.

model
string
required

ID of the model to use. Use List models to see the models available to you.

Minimum string length: 1
messages
object[]
required

A list of messages comprising the conversation so far.

Minimum array length: 1
stream
boolean | null
default:false

Whether to stream the model response data to the client as it is generated using server-sent events. See the streaming guide for how to handle the streamed events.

stream_options
object | null

Options for streaming response. Only set this when you set stream: true.

temperature
number | null

The sampling temperature, from 0 to 2. Higher values such as 0.8 make the output more varied; lower values such as 0.2 make it more focused and consistent. Adjust this or top_p, not both.

Required range: 0 <= x <= 2
top_p
number | null

The nucleus sampling threshold, from 0 to 1. The model samples only from the most likely tokens whose probabilities add up to top_p, so 0.1 limits it to the top 10% of probability mass. Adjust this or temperature, not both.

Required range: 0 <= x <= 1
max_tokens
integer | null
deprecated

The maximum number of tokens that can be generated in the chat completion, including reasoning tokens. Deprecated in favor of max_completion_tokens. The two cannot be set together.

Required range: x >= 1
max_completion_tokens
integer | null

An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens. The prompt plus this value should fit within the model's context_length. Cannot be set together with max_tokens.

Required range: x >= 1
stop

Sequences at which to stop generating further tokens: a non-empty string, or an array of 1 to 4 non-empty strings. Accepted but has no effect: generation does not stop at these sequences.

Minimum string length: 1
seed
integer<int64> | null

If specified, the system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed.

reasoning_effort
string | null

Constrains effort on reasoning: a level name of 1 to 64 characters, such as low, medium, high, xhigh, or max. A model that lists reasoning.supported_efforts accepts only those values. When omitted, the model's reasoning.default_effort applies if it lists one; otherwise the model applies its own effort. See the reasoning guide.

Required string length: 1 - 64
tools
object[] | null

A list of tools the model may call, each a function tool with a name, an optional description and JSON Schema parameters. A call the model makes is returned in message.tool_calls with finish_reason: tool_calls.

Maximum array length: 128
tool_choice

Controls which (if any) tool is called by the model: none, auto, required, or an object naming one function the model must call. Allowed only when tools is supplied.

Available options:
none,
auto,
required
parallel_tool_calls
boolean | null

Whether to enable parallel function calling during tool use. Allowed only when tools is supplied.

n
enum<integer> | null
Only 1

How many chat completion choices to generate for each input message. Only 1 is supported.

Available options:
1
logprobs
enum<boolean> | null
Only false

Whether to return log probabilities of the output tokens. Only false is supported.

Available options:
false
frequency_penalty
number | null

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

Required range: -2 <= x <= 2
presence_penalty
number | null

Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

Required range: -2 <= x <= 2
logit_bias
object | null
Only an empty object

Modifies the likelihood of specified tokens appearing in the completion. Not supported: only an empty object is accepted.

response_format
object

An object specifying the format that the model must output, selected by its type: plain text, a valid JSON object (json_object), or JSON matching the schema in json_schema (json_schema).

service_tier
enum<string> | null
Only auto or default

Specifies the processing type used for serving the request. Only auto and default are supported. Both select standard processing.

Available options:
auto,
default
store
enum<boolean> | null
Only false

Whether to store the output of this chat completion request. Only false is supported. Completions are not stored.

Available options:
false
modalities
enum<string>[] | null
Only ["text"]

Output types that you would like the model to generate. Only ["text"] is supported.

Required array length: 1 element
Available options:
text
verbosity
enum<string> | null
Only medium

Constrains the verbosity of the model's response. Only medium is supported.

Available options:
medium

Response

A chat completion, or a stream of chat completion chunks when stream is true.

Represents a chat completion response returned by the model, based on the provided input.

id
string
required

A unique identifier for the chat completion.

object
enum<string>
required

The object type, which is always chat.completion.

Available options:
chat.completion
created
integer<int64>
required

The Unix timestamp (in seconds) of when the chat completion was created.

model
string
required

The model used for the chat completion.

choices
object[]
required

A list of chat completion choices.

usage
object
required

Usage statistics for the completion request.