Skip to main content
POST
Openai Responses

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Headers

x-deepinfra-source
string | null
x-deepinfra-service-tier
string | null

Per-request service tier (priority or flex) for clients that cannot set the service_tier body field. The body field wins when both are present; unrecognized values ride the default tier.

xi-api-key
string | null
x-api-key
string | null

Body

application/json
model
string
required
input
required
service_tier
enum<string> | null

The service tier used for processing the request. 'priority' processes the request with higher priority (premium rate); 'flex' processes it at lower priority for a discount, served only when spare capacity exists and may be retried/timed out under load. Both apply only to models that support the respective tier. For compatibility, 'auto' is treated as 'priority' and 'standard_only' as 'default'.

Available options:
default,
priority,
flex
fail_fast
boolean
default:false

If true, the request is rejected immediately with HTTP 429 when the model has no spare capacity, instead of waiting in the queue. Opt-in; the default (false) keeps standard queueing behavior.

models
string[] | null

Ordered list of up to 4 fallback models. The request is attempted on each model in order: when a model rejects it for lack of capacity (HTTP 429 model-busy / flex no-capacity), the next model is tried server-side. The first model that accepts serves the request; the response's model field and billing reflect that model, at that model's pricing. Models before the last are attempted without queueing (as if fail_fast were set); the last model honors the request's own fail_fast value. When models is set, the model field is ignored. Entries must be plain model names (no deploy_id:, custom_hostport, or :revision specifiers); duplicate entries are ignored, keeping the first occurrence.

Required array length: 1 - 4 elements
instructions
string | null
tools
(ResponsesFunctionTool · object | ResponsesWebSearchTool · object | object)[] | null
tool_choice
Available options:
auto,
none,
required
text
ResponsesTextConfig · object | null
reasoning
ResponsesReasoningConfig · object | null
max_output_tokens
integer | null
max_tool_calls
integer | null
temperature
number | null
top_p
number | null
top_logprobs
integer | null
metadata
Metadata · object | null
parallel_tool_calls
boolean | null
stream
boolean
default:false
user
string | null
store
boolean
default:false
previous_response_id
string | null
background
boolean
default:false
include
string[] | null
truncation
enum<string> | null
Available options:
auto,
disabled
prompt
Prompt · object | null
conversation
unknown
min_p
number | null
top_k
integer | null
repetition_penalty
number | null
stop_token_ids
integer[] | null
chat_template_kwargs
Chat Template Kwargs · object | null
continue_final_message
boolean | null
ignore_eos
boolean | null
prompt_cache_key
string | null
prompt_cache_options
PromptCacheOptions · object | null

Response

Successful Response