Skip to main content
POST
MiniMax-M3.1-Flash-Preview is available only through M Plan and MiniMax Code for now.

Authorizations

Authorization
string
header
required

HTTP: Bearer Auth

  • Security Scheme Type: http
  • HTTP Authorization Scheme: Bearer API_key, used to authenticate your account. View it in Account Management > API Keys

Headers

Content-Type
enum<string>
default:application/json
required

Media type of the request body. Must be set to application/json

Available options:
application/json

Body

application/json
model
string
required

Model name to invoke, e.g. MiniMax-M3.1-Flash-Preview

Example:

"MiniMax-M3.1-Flash-Preview"

input
required

Conversation content. Supports either a simple text or a full conversation history array

service_tier
enum<string>
default:standard

Service tier for request admission. Supported values are standard and priority. If omitted, the request uses the standard tier. The priority price is 1.5 times the standard price and ensures priority admission so the request is processed ahead of other requests, leading to faster responses and fewer failures.

Available options:
standard,
priority
instructions
string

System instructions

max_output_tokens
integer

Maximum output token count. Reasoning tokens count toward it too, so a value that is too small yields status: "incomplete" with no message item in output.

temperature
number<float>
default:1

Sampling temperature, range (0, 1]

Required range: 0 <= x <= 1
top_p
number<float>
default:0.95

Nucleus sampling, range (0, 1]

Required range: 0 <= x <= 1
stream
boolean
default:false

Set to true to enable SSE streaming response

tools
object[]

Tool list

tool_choice
enum<string>

Tool selection strategy: none means no tool will be called; auto lets the model decide whether to call tools

Available options:
none,
auto
metadata
object

Request metadata. Both keys and values are strings

prompt_cache_key
string

Prompt cache routing identifier

text
object

Output format control

reasoning
object

Reasoning control. The default depends on the model, so no schema-level default is declared.

  • MiniMax-M3.1-Flash-Preview: reasoning is always on, including when reasoning is omitted. effort accepts low, medium, high, xhigh or max, and does tune reasoning depth. When omitted, it defaults to max. effort: "none" returns HTTP 400.
  • MiniMax-M3: reasoning is off unless effort is set to a non-none value, which turns reasoning on but does not tune its depth.
  • M2.x models: reasoning cannot be disabled; effort: "none" is accepted but ignored.

Response

200 - application/json

Successful response

id
string
required

Response ID

Example:

"abc123"

object
enum<string>
required

Object type, always response

Available options:
response
created_at
integer
required

Response creation time (Unix seconds)

model
string
required

Actual model that processed the request

status
enum<string>
required

Response status

Available options:
completed,
incomplete,
failed
output
(Message · object | Reasoning · object | Function Call · object)[]
required

Model output list

Assistant reply

output_text
string | null

Convenience field. Concatenation of all text outputs

usage
object
error
object | null

Error info, only returned when status=failed

incomplete_details
object | null

Reason for incompletion, only returned when status=incomplete

parallel_tool_calls
boolean

Whether parallel tool calls are supported

store
boolean

Whether the response is persisted

truncation
enum<string>

Context truncation strategy

Available options:
disabled