curl --request POST \
--url https://api.minimax.io/v1/responses \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: <content-type>' \
--data '
{
"model": "MiniMax-M3.1-Flash-Preview",
"reasoning": {
"effort": "max"
},
"input": "Hello!"
}
'{
"id": "abc123",
"object": "response",
"created_at": 1764000000,
"model": "MiniMax-M3.1-Flash-Preview",
"status": "completed",
"output": [
{
"id": "abc123_rs",
"type": "reasoning",
"status": "completed",
"summary": [],
"content": [
{
"type": "reasoning_text",
"text": "The user sent a greeting. Reply with a short friendly greeting and offer help."
}
]
},
{
"id": "abc123_msg",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Hello! I'm MiniMax. How can I help you today?",
"annotations": []
}
]
}
],
"output_text": "Hello! I'm MiniMax. How can I help you today?",
"usage": {
"input_tokens": 8,
"input_tokens_details": {
"cached_tokens": 0
},
"output_tokens": 86,
"output_tokens_details": {
"reasoning_tokens": 76
},
"total_tokens": 94
},
"parallel_tool_calls": true,
"store": false,
"truncation": "disabled"
}Create Response
Call MiniMax models via the OpenAI Responses API compatible main endpoint. Generates model replies, supports streaming and non-streaming.
curl --request POST \
--url https://api.minimax.io/v1/responses \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: <content-type>' \
--data '
{
"model": "MiniMax-M3.1-Flash-Preview",
"reasoning": {
"effort": "max"
},
"input": "Hello!"
}
'{
"id": "abc123",
"object": "response",
"created_at": 1764000000,
"model": "MiniMax-M3.1-Flash-Preview",
"status": "completed",
"output": [
{
"id": "abc123_rs",
"type": "reasoning",
"status": "completed",
"summary": [],
"content": [
{
"type": "reasoning_text",
"text": "The user sent a greeting. Reply with a short friendly greeting and offer help."
}
]
},
{
"id": "abc123_msg",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Hello! I'm MiniMax. How can I help you today?",
"annotations": []
}
]
}
],
"output_text": "Hello! I'm MiniMax. How can I help you today?",
"usage": {
"input_tokens": 8,
"input_tokens_details": {
"cached_tokens": 0
},
"output_tokens": 86,
"output_tokens_details": {
"reasoning_tokens": 76
},
"total_tokens": 94
},
"parallel_tool_calls": true,
"store": false,
"truncation": "disabled"
}Authorizations
HTTP: Bearer Auth
- Security Scheme Type: http
- HTTP Authorization Scheme: Bearer API_key, used to authenticate your account. View it in Account Management > API Keys
Headers
Media type of the request body. Must be set to application/json
application/json Body
Model name to invoke, e.g. MiniMax-M3.1-Flash-Preview
"MiniMax-M3.1-Flash-Preview"
Conversation content. Supports either a simple text or a full conversation history array
Service tier for request admission. Supported values are standard and priority. If omitted, the request uses the standard tier. The priority price is 1.5 times the standard price and ensures priority admission so the request is processed ahead of other requests, leading to faster responses and fewer failures.
standard, priority System instructions
Maximum output token count. Reasoning tokens count toward it too, so a value that is too small yields status: "incomplete" with no message item in output.
Sampling temperature, range (0, 1]
0 <= x <= 1Nucleus sampling, range (0, 1]
0 <= x <= 1Set to true to enable SSE streaming response
Tool list
Show child attributes
Show child attributes
Tool selection strategy: none means no tool will be called; auto lets the model decide whether to call tools
none, auto Request metadata. Both keys and values are strings
Show child attributes
Show child attributes
Prompt cache routing identifier
Output format control
Show child attributes
Show child attributes
Reasoning control. The default depends on the model, so no schema-level default is declared.
MiniMax-M3.1-Flash-Preview: reasoning is always on, including whenreasoningis omitted.effortacceptslow,medium,high,xhighormax, and does tune reasoning depth. When omitted, it defaults tomax.effort: "none"returns HTTP 400.MiniMax-M3: reasoning is off unlesseffortis set to a non-nonevalue, which turns reasoning on but does not tune its depth.- M2.x models: reasoning cannot be disabled;
effort: "none"is accepted but ignored.
Show child attributes
Show child attributes
Response
Successful response
Response ID
"abc123"
Object type, always response
response Response creation time (Unix seconds)
Actual model that processed the request
Response status
completed, incomplete, failed Model output list
Assistant reply
- Message
- Reasoning
- Function Call
Show child attributes
Show child attributes
Convenience field. Concatenation of all text outputs
Show child attributes
Show child attributes
Error info, only returned when status=failed
Show child attributes
Show child attributes
Reason for incompletion, only returned when status=incomplete
Show child attributes
Show child attributes
Whether parallel tool calls are supported
Whether the response is persisted
Context truncation strategy
disabled