> ## Documentation Index
> Fetch the complete documentation index at: https://developers.reflection.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create a chat completion

> Creates a model response for the given chat conversation. Learn more in the [text generation guide](/text-generation).
Returns a chat completion object or, if the request is streamed, a sequence of server-sent events; chat completion chunks arrive as `data:` lines, and a completed stream ends with `data: [DONE]`.

## Streaming responses

With `stream: true`, the response is a sequence of server-sent events
instead of one JSON object. Each event is a `data:` line holding a
`chat.completion.chunk`. Read the generated text from
`choices[0].delta.content` on each chunk, not from `message`, and
concatenate the pieces in order. The final chunk for each choice
carries a `finish_reason`. With `stream_options.include_usage`, one
more chunk with empty `choices` reports `usage` for the whole
request. A completed stream ends with `data: [DONE]`.


## OpenAPI

````yaml /api-reference/openapi.yaml post /openai/v1/chat/completions
openapi: 3.0.3
info:
  title: Reflection API
  version: 0.1.0
  description: >-
    Reflection's hosted inference API. The `/openai/v1` routes are
    OpenAI-compatible: chat completions and the models that serve them.
    Responses, except CORS preflights, include identical `x-request-id` and
    `x-server-request-id` values independent of client-supplied IDs.

    Once a request is authenticated and admitted for rate limiting, its
    response, including an error response, carries `x-ratelimit-` headers
    reporting the organization's remaining request capacity, and for chat
    completions its token capacity. Headers ending in `-day` report the UTC
    calendar day; the others report a per-minute limit. Token capacity counts
    the input estimated at admission, not the completion's final usage. Reset
    durations measure full replenishment without further traffic, not when a
    rejected request may be retried; use `Retry-After` for that. The headers are
    omitted when capacity cannot be reported.
servers:
  - url: https://api.reflection.ai
    description: The Reflection API.
security:
  - apiKey: []
  - playgroundCredential: []
tags:
  - name: Models
    description: List and retrieve the models you can use with the API.
  - name: Chat
    description: Generate a model response for a conversation.
paths:
  /openai/v1/chat/completions:
    post:
      tags:
        - Chat
      summary: Create a chat completion
      description: >-
        Creates a model response for the given chat conversation. Learn more in
        the [text generation guide](/text-generation).

        Returns a chat completion object or, if the request is streamed, a
        sequence of server-sent events; chat completion chunks arrive as `data:`
        lines, and a completed stream ends with `data: [DONE]`.
      operationId: generateChatCompletion
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ChatCompletionRequest'
            examples:
              Basic request:
                summary: Basic request
                value:
                  model: Beam-501B-A23B
                  messages:
                    - role: system
                      content: Answer in one sentence.
                    - role: user
                      content: Why is the sky blue?
              Streaming request:
                summary: Streaming request
                value:
                  model: Beam-501B-A23B
                  messages:
                    - role: system
                      content: Answer in one sentence.
                    - role: user
                      content: Why is the sky blue?
                  stream: true
                  stream_options:
                    include_usage: true
              Function calling:
                summary: Function calling
                value:
                  model: Beam-501B-A23B
                  messages:
                    - role: user
                      content: What's the weather in Paris right now?
                  tools:
                    - type: function
                      function:
                        name: get_weather
                        description: Get the current weather for a city.
                        parameters:
                          type: object
                          properties:
                            city:
                              type: string
                              description: The city name.
                          required:
                            - city
                          additionalProperties: false
                  tool_choice: auto
              Function calling, returning the result:
                summary: Function calling, returning the result
                value:
                  model: Beam-501B-A23B
                  messages:
                    - role: user
                      content: What's the weather in Paris right now?
                    - role: assistant
                      content: null
                      tool_calls:
                        - id: call-7f3a9c2e-5b1d-4e8a-9c36-2f0d8b4a71e5
                          type: function
                          function:
                            name: get_weather
                            arguments: '{"city": "Paris"}'
                    - role: tool
                      tool_call_id: call-7f3a9c2e-5b1d-4e8a-9c36-2f0d8b4a71e5
                      content: '{"temperature_c": 18, "conditions": "light rain"}'
                  tools:
                    - type: function
                      function:
                        name: get_weather
                        description: Get the current weather for a city.
                        parameters:
                          type: object
                          properties:
                            city:
                              type: string
                              description: The city name.
                          required:
                            - city
                          additionalProperties: false
              Structured output:
                summary: Structured output
                value:
                  model: Beam-501B-A23B
                  messages:
                    - role: user
                      content: >-
                        Extract the event: Team offsite on March 3 in Lisbon for
                        the platform group.
                  response_format:
                    type: json_schema
                    json_schema:
                      name: event
                      strict: true
                      schema:
                        type: object
                        properties:
                          name:
                            type: string
                          date:
                            type: string
                          city:
                            type: string
                        required:
                          - name
                          - date
                          - city
                        additionalProperties: false
      responses:
        '200':
          description: >-
            A chat completion, or a stream of chat completion chunks when
            `stream` is true.
          headers:
            x-request-id:
              $ref: '#/components/headers/x-request-id'
            x-server-request-id:
              $ref: '#/components/headers/x-server-request-id'
            x-ratelimit-limit-requests:
              $ref: '#/components/headers/x-ratelimit-limit-requests'
            x-ratelimit-remaining-requests:
              $ref: '#/components/headers/x-ratelimit-remaining-requests'
            x-ratelimit-reset-requests:
              $ref: '#/components/headers/x-ratelimit-reset-requests'
            x-ratelimit-limit-requests-day:
              $ref: '#/components/headers/x-ratelimit-limit-requests-day'
            x-ratelimit-remaining-requests-day:
              $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
            x-ratelimit-reset-requests-day:
              $ref: '#/components/headers/x-ratelimit-reset-requests-day'
            x-ratelimit-limit-tokens:
              $ref: '#/components/headers/x-ratelimit-limit-tokens'
            x-ratelimit-remaining-tokens:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens'
            x-ratelimit-reset-tokens:
              $ref: '#/components/headers/x-ratelimit-reset-tokens'
            x-ratelimit-limit-tokens-day:
              $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
            x-ratelimit-remaining-tokens-day:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
            x-ratelimit-reset-tokens-day:
              $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ChatCompletionResponse'
              examples:
                Chat completion:
                  summary: Chat completion
                  value:
                    id: chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f
                    object: chat.completion
                    created: 1791198000
                    model: Beam-501B-A23B
                    choices:
                      - index: 0
                        message:
                          role: assistant
                          content: >-
                            Air molecules scatter short blue wavelengths of
                            sunlight far more than longer red ones, so blue
                            light reaches your eyes from every part of the sky.
                          refusal: null
                          reasoning_content: >-
                            The user wants a one-sentence explanation of
                            Rayleigh scattering.
                        logprobs: null
                        finish_reason: stop
                    usage:
                      prompt_tokens: 23
                      completion_tokens: 45
                      total_tokens: 68
                      prompt_tokens_details:
                        cached_tokens: 0
                      completion_tokens_details:
                        reasoning_tokens: 14
                Function call:
                  summary: Function call
                  value:
                    id: chatcmpl-9b2e7d41-c0a3-4f6e-8a15-d3b7c6f02e19
                    object: chat.completion
                    created: 1791198060
                    model: Beam-501B-A23B
                    choices:
                      - index: 0
                        message:
                          role: assistant
                          content: null
                          refusal: null
                          tool_calls:
                            - id: call-7f3a9c2e-5b1d-4e8a-9c36-2f0d8b4a71e5
                              type: function
                              function:
                                name: get_weather
                                arguments: '{"city": "Paris"}'
                        logprobs: null
                        finish_reason: tool_calls
                    usage:
                      prompt_tokens: 96
                      completion_tokens: 21
                      total_tokens: 117
                Structured output:
                  summary: Structured output
                  value:
                    id: chatcmpl-2d6f0a8b-3e91-4c7b-9f54-a1e8c3d7b260
                    object: chat.completion
                    created: 1791198120
                    model: Beam-501B-A23B
                    choices:
                      - index: 0
                        message:
                          role: assistant
                          content: >-
                            {"name": "Team offsite", "date": "March 3", "city":
                            "Lisbon"}
                          refusal: null
                        logprobs: null
                        finish_reason: stop
                    usage:
                      prompt_tokens: 41
                      completion_tokens: 19
                      total_tokens: 60
            text/event-stream:
              schema:
                $ref: '#/components/schemas/ChatCompletionChunk'
              examples:
                Streaming:
                  summary: A streamed response
                  value: >+
                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"wGMmHWX0D2a"}


                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"reasoning_content":"The
                    user wants a one-sentence explanation of Rayleigh
                    scattering."},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"lCav4U"}


                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"content":"Air
                    molecules scatter short blue wavelengths of sunlight far
                    more"},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"cVmEhIa5zitBUV3O"}


                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"content":"
                    than longer red ones, so blue light reaches your eyes from
                    every"},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"6WhBZTNBLlolaThi"}


                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{"content":"
                    part of the
                    sky."},"finish_reason":null}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"cGv4yGc4Dx-S3LMCyfoevWzBWsxtzk7cCIthcSCHxpJ4yLTophhyxeC6O2Fas9Q-"}


                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":null,"obfuscation":"kxtEv0FWmRsjv6IHiQoowTE2U-ZgKvc_EmkfM2CwbV3xz_mxCobC3dFUgFzknbzBzZunVlqbQG1jCUa4YFxVcRtdghX"}


                    data:
                    {"id":"chatcmpl-4cfa1b54-d28b-4be5-b095-5c6c52b2d27f","choices":[],"created":1791198000,"model":"Beam-501B-A23B","object":"chat.completion.chunk","usage":{"prompt_tokens":23,"completion_tokens":45,"total_tokens":68,"completion_tokens_details":{"reasoning_tokens":14}},"obfuscation":"QkHVHPHgk3xZMnPQqH_DnJvrzXV"}


                    data: [DONE]

        '400':
          description: >-
            The request is invalid, and `error.param` names the parameter at
            fault. Correct the request before retrying.
          headers:
            x-request-id:
              $ref: '#/components/headers/x-request-id'
            x-server-request-id:
              $ref: '#/components/headers/x-server-request-id'
            x-ratelimit-limit-requests:
              $ref: '#/components/headers/x-ratelimit-limit-requests'
            x-ratelimit-remaining-requests:
              $ref: '#/components/headers/x-ratelimit-remaining-requests'
            x-ratelimit-reset-requests:
              $ref: '#/components/headers/x-ratelimit-reset-requests'
            x-ratelimit-limit-requests-day:
              $ref: '#/components/headers/x-ratelimit-limit-requests-day'
            x-ratelimit-remaining-requests-day:
              $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
            x-ratelimit-reset-requests-day:
              $ref: '#/components/headers/x-ratelimit-reset-requests-day'
            x-ratelimit-limit-tokens:
              $ref: '#/components/headers/x-ratelimit-limit-tokens'
            x-ratelimit-remaining-tokens:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens'
            x-ratelimit-reset-tokens:
              $ref: '#/components/headers/x-ratelimit-reset-tokens'
            x-ratelimit-limit-tokens-day:
              $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
            x-ratelimit-remaining-tokens-day:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
            x-ratelimit-reset-tokens-day:
              $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                Value out of range:
                  summary: Value out of range
                  value:
                    error:
                      message: >-
                        Invalid value for 'temperature': number must be at most
                        2.
                      type: invalid_request_error
                      param: temperature
                      code: invalid_value
                Unrecognized parameter:
                  summary: Unrecognized parameter
                  value:
                    error:
                      message: 'Unrecognized request argument supplied: ignore_eos.'
                      type: invalid_request_error
                      param: ignore_eos
                      code: unsupported_parameter
                Context window exceeded:
                  summary: Context window exceeded
                  value:
                    error:
                      message: >-
                        This model's maximum context length is 262144 tokens.
                        However, your messages resulted in 262168 tokens. Please
                        reduce the length of the messages.
                      type: invalid_request_error
                      param: messages
                      code: context_length_exceeded
        '401':
          $ref: '#/components/responses/Unauthorized'
        '402':
          $ref: '#/components/responses/PaymentRequired'
        '403':
          $ref: '#/components/responses/ChatForbidden'
        '404':
          $ref: '#/components/responses/ChatModelNotFound'
        '409':
          $ref: '#/components/responses/ChatBillingConflict'
        '413':
          description: The request body is too large. Reduce its size before retrying.
          headers:
            x-request-id:
              $ref: '#/components/headers/x-request-id'
            x-server-request-id:
              $ref: '#/components/headers/x-server-request-id'
            x-ratelimit-limit-requests:
              $ref: '#/components/headers/x-ratelimit-limit-requests'
            x-ratelimit-remaining-requests:
              $ref: '#/components/headers/x-ratelimit-remaining-requests'
            x-ratelimit-reset-requests:
              $ref: '#/components/headers/x-ratelimit-reset-requests'
            x-ratelimit-limit-requests-day:
              $ref: '#/components/headers/x-ratelimit-limit-requests-day'
            x-ratelimit-remaining-requests-day:
              $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
            x-ratelimit-reset-requests-day:
              $ref: '#/components/headers/x-ratelimit-reset-requests-day'
            x-ratelimit-limit-tokens:
              $ref: '#/components/headers/x-ratelimit-limit-tokens'
            x-ratelimit-remaining-tokens:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens'
            x-ratelimit-reset-tokens:
              $ref: '#/components/headers/x-ratelimit-reset-tokens'
            x-ratelimit-limit-tokens-day:
              $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
            x-ratelimit-remaining-tokens-day:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
            x-ratelimit-reset-tokens-day:
              $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                Request body too large:
                  summary: Request body too large
                  value:
                    error:
                      message: The request body exceeds the maximum size.
                      type: invalid_request_error
                      param: null
                      code: request_too_large
        '415':
          description: >-
            The request body is not declared as JSON. Set the `Content-Type`
            header to `application/json` before retrying.
          headers:
            x-request-id:
              $ref: '#/components/headers/x-request-id'
            x-server-request-id:
              $ref: '#/components/headers/x-server-request-id'
            x-ratelimit-limit-requests:
              $ref: '#/components/headers/x-ratelimit-limit-requests'
            x-ratelimit-remaining-requests:
              $ref: '#/components/headers/x-ratelimit-remaining-requests'
            x-ratelimit-reset-requests:
              $ref: '#/components/headers/x-ratelimit-reset-requests'
            x-ratelimit-limit-requests-day:
              $ref: '#/components/headers/x-ratelimit-limit-requests-day'
            x-ratelimit-remaining-requests-day:
              $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
            x-ratelimit-reset-requests-day:
              $ref: '#/components/headers/x-ratelimit-reset-requests-day'
            x-ratelimit-limit-tokens:
              $ref: '#/components/headers/x-ratelimit-limit-tokens'
            x-ratelimit-remaining-tokens:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens'
            x-ratelimit-reset-tokens:
              $ref: '#/components/headers/x-ratelimit-reset-tokens'
            x-ratelimit-limit-tokens-day:
              $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
            x-ratelimit-remaining-tokens-day:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
            x-ratelimit-reset-tokens-day:
              $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                Unsupported media type:
                  summary: Unsupported media type
                  value:
                    error:
                      message: >-
                        Unsupported Content-Type. Send the request body as JSON
                        with the header Content-Type: application/json.
                      type: invalid_request_error
                      param: null
                      code: unsupported_media_type
        '429':
          $ref: '#/components/responses/ChatTooManyRequests'
        '502':
          description: >-
            The inference service failed to produce a response. The request may
            be retried.
          headers:
            x-request-id:
              $ref: '#/components/headers/x-request-id'
            x-server-request-id:
              $ref: '#/components/headers/x-server-request-id'
            x-ratelimit-limit-requests:
              $ref: '#/components/headers/x-ratelimit-limit-requests'
            x-ratelimit-remaining-requests:
              $ref: '#/components/headers/x-ratelimit-remaining-requests'
            x-ratelimit-reset-requests:
              $ref: '#/components/headers/x-ratelimit-reset-requests'
            x-ratelimit-limit-requests-day:
              $ref: '#/components/headers/x-ratelimit-limit-requests-day'
            x-ratelimit-remaining-requests-day:
              $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
            x-ratelimit-reset-requests-day:
              $ref: '#/components/headers/x-ratelimit-reset-requests-day'
            x-ratelimit-limit-tokens:
              $ref: '#/components/headers/x-ratelimit-limit-tokens'
            x-ratelimit-remaining-tokens:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens'
            x-ratelimit-reset-tokens:
              $ref: '#/components/headers/x-ratelimit-reset-tokens'
            x-ratelimit-limit-tokens-day:
              $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
            x-ratelimit-remaining-tokens-day:
              $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
            x-ratelimit-reset-tokens-day:
              $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                Inference service unavailable:
                  summary: Inference service unavailable
                  value:
                    error:
                      message: The inference service is temporarily unavailable
                      type: upstream_unavailable_error
                      param: null
                      code: null
        '503':
          $ref: '#/components/responses/ChatServiceUnavailable'
      x-codeSamples:
        - lang: python
          label: Python
          source: |
            import os

            from openai import OpenAI

            client = OpenAI(
                base_url="https://api.reflection.ai/openai/v1",
                api_key=os.environ["REFLECTION_API_KEY"],
            )

            completion = client.chat.completions.create(
                model="Beam-501B-A23B",
                messages=[
                    {"role": "system", "content": "Answer in one sentence."},
                    {"role": "user", "content": "Why is the sky blue?"},
                ],
            )
            print(completion.choices[0].message.content)
        - lang: python
          label: Python (streaming)
          source: |
            import os

            from openai import OpenAI

            client = OpenAI(
                base_url="https://api.reflection.ai/openai/v1",
                api_key=os.environ["REFLECTION_API_KEY"],
            )

            stream = client.chat.completions.create(
                model="Beam-501B-A23B",
                messages=[
                    {"role": "system", "content": "Answer in one sentence."},
                    {"role": "user", "content": "Why is the sky blue?"},
                ],
                stream=True,
                stream_options={"include_usage": True},
            )
            for chunk in stream:
                if chunk.choices and chunk.choices[0].delta.content:
                    print(chunk.choices[0].delta.content, end="", flush=True)
                if chunk.usage:
                    print(f"\n{chunk.usage.total_tokens} tokens")
        - lang: typescript
          label: TypeScript
          source: |
            import OpenAI from "openai";

            const client = new OpenAI({
              baseURL: "https://api.reflection.ai/openai/v1",
              apiKey: process.env.REFLECTION_API_KEY,
            });

            const completion = await client.chat.completions.create({
              model: "Beam-501B-A23B",
              messages: [
                { role: "system", content: "Answer in one sentence." },
                { role: "user", content: "Why is the sky blue?" },
              ],
            });
            console.log(completion.choices[0].message.content);
        - lang: bash
          label: cURL (streaming)
          source: |
            curl -N https://api.reflection.ai/openai/v1/chat/completions \
              -H "Authorization: Bearer $REFLECTION_API_KEY" \
              -H "Content-Type: application/json" \
              -d '{
                "model": "Beam-501B-A23B",
                "messages": [
                  {"role": "system", "content": "Answer in one sentence."},
                  {"role": "user", "content": "Why is the sky blue?"}
                ],
                "stream": true,
                "stream_options": {"include_usage": true}
              }'
components:
  schemas:
    ChatCompletionRequest:
      type: object
      additionalProperties: true
      required:
        - model
        - messages
      properties:
        model:
          type: string
          minLength: 1
          description: >-
            ID of the model to use. Use [List models](/models) to see the models
            available to you.
        messages:
          type: array
          minItems: 1
          description: A list of messages comprising the conversation so far.
          items:
            $ref: '#/components/schemas/ChatCompletionMessage'
        stream:
          type: boolean
          nullable: true
          default: false
          description: >-
            Whether to stream the model response data to the client as it is
            generated using server-sent events. See the [streaming
            guide](/streaming) for how to handle the streamed events.
        stream_options:
          $ref: '#/components/schemas/ChatCompletionStreamOptions'
        temperature:
          type: number
          minimum: 0
          maximum: 2
          nullable: true
          description: >-
            The sampling temperature, from 0 to 2. Higher values such as 0.8
            make the output more varied; lower values such as 0.2 make it more
            focused and consistent. Adjust this or `top_p`, not both.
        top_p:
          type: number
          minimum: 0
          maximum: 1
          nullable: true
          description: >-
            The nucleus sampling threshold, from 0 to 1. The model samples only
            from the most likely tokens whose probabilities add up to `top_p`,
            so 0.1 limits it to the top 10% of probability mass. Adjust this or
            `temperature`, not both.
        max_tokens:
          type: integer
          minimum: 1
          nullable: true
          deprecated: true
          description: >-
            The maximum number of tokens that can be generated in the chat
            completion, including reasoning tokens.

            Deprecated in favor of `max_completion_tokens`. The two cannot be
            set together.
        max_completion_tokens:
          type: integer
          minimum: 1
          nullable: true
          description: >-
            An upper bound for the number of tokens that can be generated for a
            completion, including visible output tokens and [reasoning
            tokens](/reasoning). The prompt plus this value should fit within
            the model's `context_length`. Cannot be set together with
            `max_tokens`.
        stop:
          $ref: '#/components/schemas/ChatCompletionStop'
        seed:
          type: integer
          format: int64
          nullable: true
          description: >-
            If specified, the system will make a best effort to sample
            deterministically, such that repeated requests with the same `seed`
            and parameters should return the same result. Determinism is not
            guaranteed.
        reasoning_effort:
          $ref: '#/components/schemas/ChatCompletionReasoningEffort'
        tools:
          type: array
          nullable: true
          maxItems: 128
          description: >-
            A list of tools the model may call, each a `function` tool with a
            `name`, an optional `description` and JSON Schema `parameters`. A
            call the model makes is returned in `message.tool_calls` with
            `finish_reason: tool_calls`.
          items:
            type: object
            additionalProperties: true
            required:
              - type
              - function
            properties:
              type:
                type: string
                enum:
                  - function
                description: The kind of tool, which is always `function`.
              function:
                type: object
                additionalProperties: true
                required:
                  - name
                properties:
                  name:
                    type: string
                    minLength: 1
                    maxLength: 64
                    description: The function's name, as the model will call it.
                  strict:
                    type: boolean
                    nullable: true
                    description: >-
                      Whether to enable strict schema adherence when generating
                      the call. When `true`, `parameters` must be an object
                      schema whose objects set `additionalProperties: false` and
                      require every property. **Not supported:** `allOf`,
                      `oneOf`, `not`, `if`, `then`, `else`, `contains`,
                      `minContains`, `maxContains`, `uniqueItems`,
                      `unevaluatedItems`, `propertyNames`, `minProperties`,
                      `maxProperties`, `unevaluatedProperties`,
                      `dependentRequired`, `dependentSchemas`,
                      `patternProperties`, `multipleOf`, `exclusiveMinimum` or
                      `exclusiveMaximum`. Some keyword combinations and
                      `pattern` constructs are also refused.
                description: >-
                  The function definition, with its `name`, optional
                  `description`, JSON Schema `parameters` and optional `strict`
                  flag.
        tool_choice:
          $ref: '#/components/schemas/ChatCompletionToolChoice'
        parallel_tool_calls:
          type: boolean
          nullable: true
          description: >-
            Whether to enable parallel function calling during tool use. Allowed
            only when `tools` is supplied.
        'n':
          type: integer
          nullable: true
          enum:
            - 1
          x-reflection-default-only: '1'
          description: >-
            How many chat completion choices to generate for each input message.
            **Only `1` is supported.**
          x-mint:
            post:
              - Only 1
        logprobs:
          type: boolean
          nullable: true
          enum:
            - false
          x-reflection-default-only: 'false'
          description: >-
            Whether to return log probabilities of the output tokens. **Only
            `false` is supported.**
          x-mint:
            post:
              - Only false
        frequency_penalty:
          type: number
          minimum: -2
          maximum: 2
          nullable: true
          description: >-
            Number between -2.0 and 2.0. Positive values penalize new tokens
            based on their existing frequency in the text so far, decreasing the
            model's likelihood to repeat the same line verbatim.
        presence_penalty:
          type: number
          minimum: -2
          maximum: 2
          nullable: true
          description: >-
            Number between -2.0 and 2.0. Positive values penalize new tokens
            based on whether they appear in the text so far, increasing the
            model's likelihood to talk about new topics.
        logit_bias:
          type: object
          nullable: true
          maxProperties: 0
          x-reflection-default-only: an empty object
          description: >-
            Modifies the likelihood of specified tokens appearing in the
            completion. **Not supported:** only an empty object is accepted.
          x-mint:
            post:
              - Only an empty object
        response_format:
          $ref: '#/components/schemas/ChatCompletionResponseFormat'
        service_tier:
          type: string
          nullable: true
          enum:
            - auto
            - default
          x-reflection-default-only: auto or default
          description: >-
            Specifies the processing type used for serving the request. **Only
            `auto` and `default` are supported.** Both select standard
            processing.
          x-mint:
            post:
              - Only auto or default
        store:
          type: boolean
          nullable: true
          enum:
            - false
          x-reflection-default-only: 'false'
          description: >-
            Whether to store the output of this chat completion request. **Only
            `false` is supported.** Completions are not stored.
          x-mint:
            post:
              - Only false
        modalities:
          type: array
          nullable: true
          minItems: 1
          maxItems: 1
          items:
            type: string
            enum:
              - text
          x-reflection-default-only: '["text"]'
          description: >-
            Output types that you would like the model to generate. **Only
            `["text"]` is supported.**
          x-mint:
            post:
              - Only ["text"]
        verbosity:
          type: string
          nullable: true
          enum:
            - medium
          x-reflection-default-only: medium
          description: >-
            Constrains the verbosity of the model's response. **Only `medium` is
            supported.**
          x-mint:
            post:
              - Only medium
      description: Parameters not listed here are not supported.
    ChatCompletionResponse:
      type: object
      description: >-
        Represents a chat completion response returned by the model, based on
        the provided input.
      required:
        - id
        - object
        - created
        - model
        - choices
        - usage
      properties:
        id:
          type: string
          description: A unique identifier for the chat completion.
        object:
          type: string
          enum:
            - chat.completion
          description: The object type, which is always `chat.completion`.
        created:
          type: integer
          format: int64
          description: >-
            The Unix timestamp (in seconds) of when the chat completion was
            created.
        model:
          type: string
          description: The model used for the chat completion.
        choices:
          type: array
          description: A list of chat completion choices.
          items:
            $ref: '#/components/schemas/ChatCompletionChoice'
        usage:
          $ref: '#/components/schemas/CompletionUsage'
    ChatCompletionChunk:
      type: object
      description: >-
        Represents a streamed chunk of a chat completion response returned by
        the model, based on the provided input. A completed stream ends with a
        `data: [DONE]` message. `service_tier` and `system_fingerprint` never
        appear on a chunk.
      required:
        - id
        - object
        - created
        - model
        - choices
      properties:
        id:
          type: string
          description: >-
            A unique identifier for the chat completion. Each chunk has the same
            ID.
        object:
          type: string
          enum:
            - chat.completion.chunk
          description: The object type, which is always `chat.completion.chunk`.
        created:
          type: integer
          format: int64
          description: >-
            The Unix timestamp (in seconds) of when the chat completion was
            created. Each chunk has the same timestamp.
        model:
          type: string
          description: The model used for the chat completion.
        choices:
          type: array
          description: >-
            A list of chat completion choices. Empty on the final usage chunk
            when `stream_options.include_usage` is true.
          items:
            $ref: '#/components/schemas/ChatCompletionChunkChoice'
        usage:
          nullable: true
          description: >-
            Token usage statistics for the entire request. Present only on the
            final chunk, and only when `stream_options.include_usage` is true;
            null or absent otherwise.
          allOf:
            - $ref: '#/components/schemas/CompletionUsage'
        obfuscation:
          type: string
          description: >-
            Random filler that pads the event to the stream's fixed length,
            hiding the lengths of the tokens it carries; a larger event is
            padded in 64-byte steps, revealing its length to within 64 bytes.
            The value carries no information and must be ignored, and may be
            empty. Absent only when `include_obfuscation` was false.
    ErrorResponse:
      type: object
      description: The body returned with every error status.
      required:
        - error
      properties:
        error:
          $ref: '#/components/schemas/Error'
    ChatCompletionMessage:
      type: object
      additionalProperties: false
      description: >-
        A message in the conversation. `content` is required except on an
        assistant turn, where it may be null or absent when the turn produced
        only reasoning or tool calls, as a prior completion may have returned.
      required:
        - role
      properties:
        role:
          type: string
          enum:
            - developer
            - system
            - user
            - assistant
            - tool
          description: The role of the message's author.
        content:
          $ref: '#/components/schemas/ChatCompletionMessageContent'
        name:
          type: string
          description: >-
            An optional name for the participant, distinguishing participants
            with the same role. On a `tool` message, the name of the tool that
            produced the result.
        tool_calls:
          type: array
          nullable: true
          description: >-
            The tool calls of a prior assistant turn, as returned in that
            completion's `message.tool_calls`. Accepted only on assistant
            messages. Each call needs a distinct `id`, and the messages that
            directly follow the turn must be `tool` messages answering every
            call exactly once, in any order, before any other message.
          items:
            $ref: '#/components/schemas/ChatCompletionMessageToolCall'
        tool_call_id:
          type: string
          nullable: true
          minLength: 1
          description: >-
            The `id` of the tool call this message answers. Required on, and
            accepted only on, `tool` messages. The call must belong to the
            assistant message this `tool` message follows, directly or after
            that message's other results, and must not already be answered.
        reasoning_content:
          type: string
          nullable: true
          description: >-
            The reasoning of a prior assistant turn, as returned in that
            completion's `message.reasoning_content`. Accepted only on assistant
            messages.
        refusal:
          type: string
          nullable: true
          description: >-
            The refusal of a prior assistant turn, as returned in that
            completion's `message.refusal`. Accepted only on assistant messages.
    ChatCompletionStreamOptions:
      type: object
      nullable: true
      additionalProperties: false
      description: >-
        Options for streaming response. Only set this when you set `stream:
        true`.
      properties:
        include_usage:
          type: boolean
          default: false
          description: >-
            Whether to stream an additional chunk before the `data: [DONE]`
            message. The `usage` field on this chunk shows the token usage
            statistics for the entire request, and the `choices` field will
            always be an empty array.
        include_obfuscation:
          type: boolean
          default: true
          description: >-
            Whether streamed events have their payloads and timings normalized
            in a manner that mitigates specific side-channel attacks. When
            enabled, each event carries an `obfuscation` property that pads
            events to a fixed length with random characters. Setting to `false`
            may slightly increase the rate at which chunks are received.
            Obfuscation should only be disabled in environments where traffic is
            not observable.
    ChatCompletionStop:
      nullable: true
      description: >-
        Sequences at which to stop generating further tokens: a non-empty
        string, or an array of 1 to 4 non-empty strings. **Accepted but has no
        effect:** generation does not stop at these sequences.
      oneOf:
        - type: string
          minLength: 1
        - type: array
          minItems: 1
          maxItems: 4
          items:
            type: string
            minLength: 1
    ChatCompletionReasoningEffort:
      type: string
      nullable: true
      minLength: 1
      maxLength: 64
      description: >-
        Constrains effort on reasoning: a level name of 1 to 64 characters, such
        as `low`, `medium`, `high`, `xhigh`, or `max`. A model that lists
        `reasoning.supported_efforts` accepts only those values. When omitted,
        the model's `reasoning.default_effort` applies if it lists one;
        otherwise the model applies its own effort. See the [reasoning
        guide](/reasoning).
    ChatCompletionToolChoice:
      nullable: true
      description: >-
        Controls which (if any) tool is called by the model: `none`, `auto`,
        `required`, or an object naming one function the model must call.
        Allowed only when `tools` is supplied.
      oneOf:
        - type: string
          enum:
            - none
            - auto
            - required
        - type: object
          additionalProperties: false
          required:
            - type
            - function
          properties:
            type:
              type: string
              enum:
                - function
              description: The kind of tool to force, which is always `function`.
            function:
              type: object
              additionalProperties: false
              required:
                - name
              properties:
                name:
                  type: string
                  minLength: 1
                  description: The name of the function the model must call.
              description: The function the model must call.
    ChatCompletionResponseFormat:
      description: >-
        An object specifying the format that the model must output, selected by
        its `type`: plain `text`, a valid JSON object (`json_object`), or JSON
        matching the schema in `json_schema` (`json_schema`).
      oneOf:
        - $ref: '#/components/schemas/ChatCompletionResponseFormatText'
        - $ref: '#/components/schemas/ChatCompletionResponseFormatJsonObject'
        - $ref: '#/components/schemas/ChatCompletionResponseFormatJsonSchema'
      discriminator:
        propertyName: type
        mapping:
          text: '#/components/schemas/ChatCompletionResponseFormatText'
          json_object: '#/components/schemas/ChatCompletionResponseFormatJsonObject'
          json_schema: '#/components/schemas/ChatCompletionResponseFormatJsonSchema'
    ChatCompletionChoice:
      type: object
      description: One of the completions generated for the request.
      required:
        - index
        - message
        - finish_reason
        - logprobs
      properties:
        index:
          type: integer
          description: The index of the choice in the list of choices.
        message:
          $ref: '#/components/schemas/AssistantMessage'
        logprobs:
          type: object
          nullable: true
          additionalProperties: true
          description: >-
            Log probability information for the choice. Always null, as log
            probabilities are not supported.
        finish_reason:
          type: string
          nullable: true
          enum:
            - stop
            - length
            - tool_calls
            - content_filter
          description: >-
            The reason the model stopped generating tokens. This will be `stop`
            if the model hit a natural stop point, `length` if the maximum
            number of tokens specified in the request was reached,
            `content_filter` if content was omitted due to a content filter, or
            `tool_calls` if the model called a tool.
    CompletionUsage:
      type: object
      description: Usage statistics for the completion request.
      required:
        - prompt_tokens
        - completion_tokens
        - total_tokens
      properties:
        prompt_tokens:
          type: integer
          minimum: 0
          description: Number of tokens in the prompt.
        completion_tokens:
          type: integer
          minimum: 0
          description: Number of tokens in the generated completion.
        total_tokens:
          type: integer
          minimum: 0
          description: Total number of tokens used in the request (prompt + completion).
        prompt_tokens_details:
          $ref: '#/components/schemas/PromptTokensDetails'
        completion_tokens_details:
          $ref: '#/components/schemas/CompletionTokensDetails'
    ChatCompletionChunkChoice:
      type: object
      description: >-
        One choice's portion of a streamed chunk. `matched_stop` is not part of
        a streamed choice and never appears.
      required:
        - index
        - delta
        - finish_reason
      properties:
        index:
          type: integer
          description: The index of the choice in the list of choices.
        delta:
          $ref: '#/components/schemas/ChatCompletionDelta'
        logprobs:
          type: object
          nullable: true
          additionalProperties: true
          description: >-
            Log probability information for the choice. Always null, as log
            probabilities are not supported.
        finish_reason:
          type: string
          nullable: true
          enum:
            - stop
            - length
            - tool_calls
            - content_filter
          x-enum-varnames:
            - ChatCompletionChunkFinishReasonStop
            - ChatCompletionChunkFinishReasonLength
            - ChatCompletionChunkFinishReasonToolCalls
            - ChatCompletionChunkFinishReasonContentFilter
          description: >-
            The reason the model stopped generating tokens, or null until the
            choice is finished. This will be `stop` if the model hit a natural
            stop point, `length` if the requested maximum number of tokens was
            reached, `content_filter` if content was omitted due to a content
            filter, or `tool_calls` if the model called a tool.
    Error:
      type: object
      description: An error returned by the API.
      required:
        - message
        - type
        - param
        - code
      properties:
        message:
          type: string
          description: A human-readable description of the error.
        type:
          type: string
          description: The error category, such as `invalid_request_error`.
        param:
          type: string
          nullable: true
          description: The request parameter the error relates to, or null.
        code:
          type: string
          nullable: true
          description: A machine-readable error code, or null.
    ChatCompletionMessageContent:
      nullable: true
      description: >-
        The contents of the message, either a string or a non-empty array of
        text parts, or null on an assistant message only.
      oneOf:
        - type: string
        - type: array
          minItems: 1
          description: >-
            An array of text parts, concatenated in order into the text the
            model receives.
          items:
            $ref: '#/components/schemas/ChatCompletionContentPartText'
    ChatCompletionMessageToolCall:
      type: object
      additionalProperties: true
      description: A tool call made in a prior assistant turn.
      required:
        - type
        - id
        - function
      properties:
        type:
          type: string
          enum:
            - function
          description: The kind of tool call, which is always `function`.
        id:
          type: string
          minLength: 1
          description: >-
            The ID of the tool call, which the `tool` message answering it gives
            as its `tool_call_id`.
        function:
          type: object
          additionalProperties: true
          required:
            - name
            - arguments
          properties:
            name:
              type: string
              description: The name of the function the model called.
            arguments:
              type: string
              description: >-
                The arguments the model generated for the call, as a JSON
                string.
          description: The function the model called.
    ChatCompletionResponseFormatText:
      type: object
      additionalProperties: false
      required:
        - type
      description: Plain text output.
      properties:
        type:
          type: string
          enum:
            - text
          description: The type of response format, which is always `text`.
    ChatCompletionResponseFormatJsonObject:
      type: object
      additionalProperties: false
      required:
        - type
      description: Output that is a valid JSON object, with no schema applied.
      properties:
        type:
          type: string
          enum:
            - json_object
          description: The type of response format, which is always `json_object`.
    ChatCompletionResponseFormatJsonSchema:
      type: object
      additionalProperties: false
      required:
        - type
        - json_schema
      description: Output that is JSON matching the supplied schema.
      properties:
        type:
          type: string
          enum:
            - json_schema
          description: The type of response format, which is always `json_schema`.
        json_schema:
          type: object
          additionalProperties: true
          required:
            - name
            - schema
          properties:
            name:
              type: string
              minLength: 1
              maxLength: 64
              description: The name of the response format.
            schema:
              type: object
              additionalProperties: true
              description: The JSON Schema the output must satisfy.
            strict:
              type: boolean
              nullable: true
              description: >-
                Whether to enable strict schema adherence when generating the
                output. When `true`, `schema` must be an object schema whose
                objects set `additionalProperties: false` and require every
                property. **Not supported:** `allOf`, `oneOf`, `not`, `if`,
                `then`, `else`, `contains`, `minContains`, `maxContains`,
                `uniqueItems`, `unevaluatedItems`, `propertyNames`,
                `minProperties`, `maxProperties`, `unevaluatedProperties`,
                `dependentRequired`, `dependentSchemas`, `patternProperties`,
                `multipleOf`, `exclusiveMinimum` or `exclusiveMaximum`. Some
                keyword combinations and `pattern` constructs are also refused.
          description: >-
            The schema definition for `json_schema`: a `name`, an optional
            `description`, the JSON `schema`, and an optional `strict` flag.
    AssistantMessage:
      type: object
      description: A chat completion message generated by the model.
      required:
        - role
        - content
        - refusal
      properties:
        role:
          type: string
          enum:
            - assistant
          description: The role of the author of this message.
        content:
          type: string
          nullable: true
          description: >-
            The contents of the message. Null when the model made only tool
            calls, or when the token limit was reached before any visible output
            was generated.
        refusal:
          type: string
          nullable: true
          description: The refusal message generated by the model.
        reasoning_content:
          type: string
          nullable: true
          description: >-
            The reasoning the model produced before the final answer, reported
            separately from `content`. Its tokens are counted in
            `usage.completion_tokens_details.reasoning_tokens`.
        tool_calls:
          type: array
          nullable: true
          description: >-
            The tool calls the model made: each has an `id`, `type: function`,
            and a `function` with `name` and JSON `arguments`. Present when
            `finish_reason` is `tool_calls`.
          items:
            type: object
            additionalProperties: true
    PromptTokensDetails:
      type: object
      description: Breakdown of tokens used in the prompt.
      properties:
        cached_tokens:
          type: integer
          minimum: 0
          description: Cached tokens present in the prompt.
    CompletionTokensDetails:
      type: object
      description: Breakdown of tokens used in a completion.
      properties:
        reasoning_tokens:
          type: integer
          minimum: 0
          description: >-
            Tokens generated by the model for reasoning. They count toward the
            request's token limit but are not returned as `content`.
    ChatCompletionDelta:
      type: object
      description: A chat completion delta generated by streamed model responses.
      properties:
        role:
          type: string
          enum:
            - assistant
          description: The role of the author of this message.
        content:
          type: string
          nullable: true
          description: The contents of the chunk message.
        reasoning_content:
          type: string
          nullable: true
          description: >-
            The reasoning contents carried by this chunk, reported separately
            from `content`.
        tool_calls:
          type: array
          description: >-
            Tool-call fragments carried by this chunk: each has an `index` and,
            over successive chunks, the call's `id`, `type`, function `name` and
            `arguments` pieces.
          items:
            type: object
            additionalProperties: true
    ChatCompletionContentPartText:
      type: object
      description: A text part of a message.
      additionalProperties: false
      required:
        - type
        - text
      properties:
        type:
          type: string
          enum:
            - text
          description: The type of the content part, which is always `text`.
        text:
          type: string
          description: The text content.
  headers:
    x-request-id:
      description: >-
        A unique identifier the server assigns to this request. Include it when
        contacting support. A request ID you send is not reused.
      schema:
        type: string
        format: uuid
      example: 7c9e6679-7425-40de-944b-e07fc1f90ae7
    x-server-request-id:
      description: The same identifier as `x-request-id`.
      schema:
        type: string
        format: uuid
      example: 7c9e6679-7425-40de-944b-e07fc1f90ae7
    x-ratelimit-limit-requests:
      description: The most requests the organization may use per minute.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 600
    x-ratelimit-remaining-requests:
      description: >-
        The requests left in the per-minute capacity, rounded down. Reflects
        this request when it was admitted; a rejected request uses none.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 599
    x-ratelimit-reset-requests:
      description: >-
        The time until the per-minute capacity is fully replenished if no
        further requests are used, as seconds with up to millisecond precision
        and an `s` suffix. Not a retry delay; see `Retry-After`.
      schema:
        type: string
      example: 0.1s
    x-ratelimit-limit-requests-day:
      description: The most requests the organization may use per UTC calendar day.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 100000
    x-ratelimit-remaining-requests-day:
      description: >-
        The requests left in today's UTC calendar-day capacity, rounded down.
        Reflects this request when it was admitted; a rejected request uses
        none.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 99412
    x-ratelimit-reset-requests-day:
      description: >-
        The time until the next UTC midnight, when the daily count resets, as
        seconds with up to millisecond precision and an `s` suffix. Not a retry
        delay; see `Retry-After`.
      schema:
        type: string
      example: 43200s
    x-ratelimit-limit-tokens:
      description: >-
        The most tokens the organization may use per minute. Sent only on chat
        completions.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 2000000
    x-ratelimit-remaining-tokens:
      description: >-
        The tokens left in the per-minute capacity, rounded down. Reflects this
        request when it was admitted; a rejected request uses none. Counts the
        input tokens estimated at admission, not final usage. Sent only on chat
        completions.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 1987650
    x-ratelimit-reset-tokens:
      description: >-
        The time until the per-minute capacity is fully replenished if no
        further tokens are used, as seconds with up to millisecond precision and
        an `s` suffix. Not a retry delay; see `Retry-After`. Sent only on chat
        completions.
      schema:
        type: string
      example: 0.371s
    x-ratelimit-limit-tokens-day:
      description: >-
        The most tokens the organization may use per UTC calendar day. Sent only
        on chat completions.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 200000000
    x-ratelimit-remaining-tokens-day:
      description: >-
        The tokens left in today's UTC calendar-day capacity, rounded down.
        Reflects this request when it was admitted; a rejected request uses
        none. Counts the input tokens estimated at admission, not final usage.
        Sent only on chat completions.
      schema:
        type: integer
        format: int64
        minimum: 0
      example: 187340215
    x-ratelimit-reset-tokens-day:
      description: >-
        The time until the next UTC midnight, when the daily count resets, as
        seconds with up to millisecond precision and an `s` suffix. Not a retry
        delay; see `Retry-After`. Sent only on chat completions.
      schema:
        type: string
      example: 43200s
  responses:
    Unauthorized:
      description: >-
        The credential is missing, malformed, or not valid. Correct the
        credential before retrying.
      headers:
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            Invalid API key:
              summary: Invalid API key
              value:
                error:
                  message: >-
                    Incorrect API key provided. The key may have been revoked or
                    expired.
                  type: authentication_error
                  param: null
                  code: invalid_api_key
            Missing credential:
              summary: Missing credential
              value:
                error:
                  message: >-
                    Provide an API key or Playground credential in the
                    Authorization header.
                  type: authentication_error
                  param: null
                  code: missing_credentials
    PaymentRequired:
      description: >-
        Your organization has no credit left. `insufficient_credits` means it
        has no remaining credits; `prepaid_balance_exhausted` means its prepaid
        balance is used up. Where plan limits do not yet refuse the daily
        allowance, a used-up allowance also returns `insufficient_credits`.
        Retry only once credit is available: the allowance renews daily, and a
        prepaid balance needs a top-up.
      headers:
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            No remaining credits:
              summary: No remaining credits
              value:
                error:
                  message: >-
                    API access is inactive because this organization has no
                    remaining credits.
                  type: permission_error
                  param: null
                  code: insufficient_credits
            Prepaid balance used up:
              summary: Prepaid balance used up
              value:
                error:
                  message: >-
                    This organization's prepaid balance is used up. Top up in
                    Billing to resume API requests.
                  type: permission_error
                  param: null
                  code: prepaid_balance_exhausted
    ChatForbidden:
      description: >-
        The credential is valid, but the request is not allowed: the caller is
        not an active member of the organization, an access restriction or
        suspension applies, the organization's plan does not include this kind
        of credential, or the organization must verify a card before using API
        keys. Retry only after the cause is resolved.
      headers:
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            Card verification required:
              summary: Card verification required
              value:
                error:
                  message: Verify a credit card in Billing before using API keys.
                  type: permission_error
                  param: null
                  code: payment_method_required
            Credential not in plan:
              $ref: '#/components/examples/CredentialNotInPlan'
            Not an organization member:
              $ref: '#/components/examples/NotOrganizationMember'
    ChatModelNotFound:
      description: >-
        The requested model does not exist or is not available to you. Choose a
        model from the models list before retrying.
      headers:
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
        x-ratelimit-limit-requests:
          $ref: '#/components/headers/x-ratelimit-limit-requests'
        x-ratelimit-remaining-requests:
          $ref: '#/components/headers/x-ratelimit-remaining-requests'
        x-ratelimit-reset-requests:
          $ref: '#/components/headers/x-ratelimit-reset-requests'
        x-ratelimit-limit-requests-day:
          $ref: '#/components/headers/x-ratelimit-limit-requests-day'
        x-ratelimit-remaining-requests-day:
          $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
        x-ratelimit-reset-requests-day:
          $ref: '#/components/headers/x-ratelimit-reset-requests-day'
        x-ratelimit-limit-tokens:
          $ref: '#/components/headers/x-ratelimit-limit-tokens'
        x-ratelimit-remaining-tokens:
          $ref: '#/components/headers/x-ratelimit-remaining-tokens'
        x-ratelimit-reset-tokens:
          $ref: '#/components/headers/x-ratelimit-reset-tokens'
        x-ratelimit-limit-tokens-day:
          $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
        x-ratelimit-remaining-tokens-day:
          $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
        x-ratelimit-reset-tokens-day:
          $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            Unknown model:
              $ref: '#/components/examples/UnknownModel'
    ChatBillingConflict:
      description: >-
        Your organization's billing is not ready to serve requests.
        `tier_switch_pending` means its billing tier is changing: wait for the
        number of seconds in `Retry-After`, then retry. `billing_setup_required`
        means its billing setup is incomplete: retry only after finishing setup
        in Billing.
      headers:
        Retry-After:
          description: >-
            The minimum number of seconds to wait before retrying. Sent with
            `tier_switch_pending`.
          schema:
            type: integer
            minimum: 1
        x-should-retry:
          description: >-
            `false` when no retry can succeed until billing setup is finished;
            sent with `billing_setup_required`. Browsers may read it.
          schema:
            type: string
            enum:
              - 'false'
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            Tier switch in progress:
              summary: Tier switch in progress
              value:
                error:
                  message: >-
                    This organization's billing tier is changing. Retry after
                    the indicated delay.
                  type: permission_error
                  param: null
                  code: tier_switch_pending
            Billing setup incomplete:
              summary: Billing setup incomplete
              value:
                error:
                  message: >-
                    This organization's billing setup is incomplete. Finish
                    setup in Billing to resume API requests.
                  type: permission_error
                  param: null
                  code: billing_setup_required
    ChatTooManyRequests:
      description: >-
        A rate limit was exceeded: your organization's requests or tokens per
        period, its plan's daily token allowance, or the requests it or an API
        key may run at once. The code is always `rate_limit_exceeded`. Wait for
        the seconds in `Retry-After`, when present, then retry. A spent daily
        limit or allowance also sends `x-should-retry: false`: it renews at
        00:00 UTC.
      headers:
        Retry-After:
          description: >-
            The minimum number of seconds to wait before retrying. Omitted when
            no retry delay is known.
          schema:
            type: integer
            minimum: 1
        x-should-retry:
          description: >-
            `false` when retrying cannot succeed before `Retry-After`; sent when
            the organization's daily request limit, daily token limit, or daily
            token allowance is spent. Browsers may read it.
          schema:
            type: string
            enum:
              - 'false'
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
        x-ratelimit-limit-requests:
          $ref: '#/components/headers/x-ratelimit-limit-requests'
        x-ratelimit-remaining-requests:
          $ref: '#/components/headers/x-ratelimit-remaining-requests'
        x-ratelimit-reset-requests:
          $ref: '#/components/headers/x-ratelimit-reset-requests'
        x-ratelimit-limit-requests-day:
          $ref: '#/components/headers/x-ratelimit-limit-requests-day'
        x-ratelimit-remaining-requests-day:
          $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
        x-ratelimit-reset-requests-day:
          $ref: '#/components/headers/x-ratelimit-reset-requests-day'
        x-ratelimit-limit-tokens:
          $ref: '#/components/headers/x-ratelimit-limit-tokens'
        x-ratelimit-remaining-tokens:
          $ref: '#/components/headers/x-ratelimit-remaining-tokens'
        x-ratelimit-reset-tokens:
          $ref: '#/components/headers/x-ratelimit-reset-tokens'
        x-ratelimit-limit-tokens-day:
          $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
        x-ratelimit-remaining-tokens-day:
          $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
        x-ratelimit-reset-tokens-day:
          $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            Request rate limit:
              $ref: '#/components/examples/RequestRateLimit'
            Token rate limit:
              $ref: '#/components/examples/TokenRateLimit'
            Daily request limit:
              $ref: '#/components/examples/DailyRequestLimit'
            Daily token limit:
              $ref: '#/components/examples/DailyTokenLimit'
            Daily allowance:
              summary: Daily allowance
              value:
                error:
                  message: >-
                    Your organization has used its daily token allowance. It
                    renews at 00:00 UTC.
                  type: rate_limit_error
                  param: null
                  code: rate_limit_exceeded
            Plan requests in flight:
              summary: Plan requests in flight
              value:
                error:
                  message: >-
                    Your organization has reached its concurrent request limit.
                    Retry after the indicated delay.
                  type: rate_limit_error
                  param: null
                  code: rate_limit_exceeded
            Concurrent request limit:
              summary: Concurrent request limit
              value:
                error:
                  message: >-
                    Your organization or API key has reached its concurrent
                    request limit.
                  type: rate_limit_error
                  param: null
                  code: rate_limit_exceeded
    ChatServiceUnavailable:
      description: >-
        The service is temporarily unavailable. The request may be retried,
        after the number of seconds in `Retry-After` when present.
      headers:
        Retry-After:
          description: >-
            The minimum number of seconds to wait before retrying. Sent when
            inference capacity is temporarily exhausted, billing setup is in
            progress, or a check the request depends on cannot answer.
          schema:
            type: integer
            minimum: 1
        x-request-id:
          $ref: '#/components/headers/x-request-id'
        x-server-request-id:
          $ref: '#/components/headers/x-server-request-id'
        x-ratelimit-limit-requests:
          $ref: '#/components/headers/x-ratelimit-limit-requests'
        x-ratelimit-remaining-requests:
          $ref: '#/components/headers/x-ratelimit-remaining-requests'
        x-ratelimit-reset-requests:
          $ref: '#/components/headers/x-ratelimit-reset-requests'
        x-ratelimit-limit-requests-day:
          $ref: '#/components/headers/x-ratelimit-limit-requests-day'
        x-ratelimit-remaining-requests-day:
          $ref: '#/components/headers/x-ratelimit-remaining-requests-day'
        x-ratelimit-reset-requests-day:
          $ref: '#/components/headers/x-ratelimit-reset-requests-day'
        x-ratelimit-limit-tokens:
          $ref: '#/components/headers/x-ratelimit-limit-tokens'
        x-ratelimit-remaining-tokens:
          $ref: '#/components/headers/x-ratelimit-remaining-tokens'
        x-ratelimit-reset-tokens:
          $ref: '#/components/headers/x-ratelimit-reset-tokens'
        x-ratelimit-limit-tokens-day:
          $ref: '#/components/headers/x-ratelimit-limit-tokens-day'
        x-ratelimit-remaining-tokens-day:
          $ref: '#/components/headers/x-ratelimit-remaining-tokens-day'
        x-ratelimit-reset-tokens-day:
          $ref: '#/components/headers/x-ratelimit-reset-tokens-day'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
          examples:
            Rate limiting unavailable:
              $ref: '#/components/examples/RateLimitingUnavailable'
            Inference capacity exhausted:
              $ref: '#/components/examples/InferenceCapacityExhausted'
            Billing setup in progress:
              summary: Billing setup in progress
              value:
                error:
                  message: >-
                    Billing setup is still in progress. Retry after the
                    indicated delay.
                  type: api_error
                  param: null
                  code: billing_provisioning_pending
            Inference capacity unavailable:
              summary: Inference capacity unavailable
              value:
                error:
                  message: >-
                    Inference capacity is temporarily unavailable. Retry after
                    the indicated delay.
                  type: api_error
                  param: null
                  code: inference_capacity_unavailable
  examples:
    CredentialNotInPlan:
      summary: Credential not in plan
      value:
        error:
          message: >-
            Your organization's plan does not include access with this kind of
            credential.
          type: permission_error
          param: null
          code: plan_access_denied
    NotOrganizationMember:
      summary: Not an organization member
      value:
        error:
          message: You are not an active member of the credential's organization.
          type: permission_error
          param: null
          code: organization_membership_required
    UnknownModel:
      summary: Unknown model
      value:
        error:
          message: The model [reflection-large] you requested is unavailable.
          type: invalid_request_error
          param: model
          code: model_not_found
    RequestRateLimit:
      summary: Request rate limit
      value:
        error:
          message: >-
            Your organization has reached its request limit. Retry after the
            indicated delay.
          type: rate_limit_error
          param: null
          code: rate_limit_exceeded
    TokenRateLimit:
      summary: Token rate limit
      value:
        error:
          message: >-
            Your organization has reached its token limit. Retry after the
            indicated delay.
          type: rate_limit_error
          param: null
          code: rate_limit_exceeded
    DailyRequestLimit:
      summary: Daily request limit
      value:
        error:
          message: >-
            Your organization has reached its daily request limit. It resets at
            00:00 UTC.
          type: rate_limit_error
          param: null
          code: rate_limit_exceeded
    DailyTokenLimit:
      summary: Daily token limit
      value:
        error:
          message: >-
            Your organization has reached its daily token limit. It resets at
            00:00 UTC.
          type: rate_limit_error
          param: null
          code: rate_limit_exceeded
    RateLimitingUnavailable:
      summary: Rate limiting unavailable
      value:
        error:
          message: Rate limiting is temporarily unavailable. Retry the request.
          type: api_error
          param: null
          code: rate_limit_unavailable
    InferenceCapacityExhausted:
      summary: Inference capacity exhausted
      value:
        error:
          message: >-
            Inference capacity is temporarily unavailable. Retry after the
            indicated delay.
          type: api_error
          param: null
          code: infrastructure_rate_limit_exceeded
  securitySchemes:
    apiKey:
      type: http
      scheme: bearer
      bearerFormat: Reflection API key
      description: >-
        A Reflection API key, sent in the `Authorization` header as `Bearer <API
        key>`. Create a key for your project from the [Reflection
        platform](https://platform.reflection.ai/api-keys).
    playgroundCredential:
      type: http
      scheme: bearer
      bearerFormat: Reflection Playground credential
      description: >-
        A short-lived, project-scoped credential that the Reflection Playground
        issues to itself while you use it. It is not meant for direct use; send
        an API key instead.

````