Skip to main content
POST
✨ New model — MiniMax-M3.1-Flash-PreviewCore capabilities: 1M long context, multimodal, tunable thinking depth.Available only through M Plan and MiniMax Code for now.
What’s new in MiniMax-M3.1-Flash-Preview:
  1. Image and video understanding — see the example code on the right
  2. Tune thinking depth via output_config.effort, which accepts low, medium, high, xhigh, and max; the default is max
  3. Thinking is on by default with no configuration needed; thinking content is returned separately in thinking content blocks

Authorizations

Authorization
string
header
required

Bearer API Key auth. Send Authorization: Bearer <API_KEY>. If Authorization and x-api-key are both present, Authorization takes precedence.

Headers

Content-Type
enum<string>
default:application/json
required

Media type of the request body, should be set to application/json to ensure JSON format

Available options:
application/json

Body

application/json
model
enum<string>
required

Model ID. MiniMax-M3.1-Flash-Preview and MiniMax-M3 are multimodal models with native support for text, image, and video input, alongside tool use and thinking content blocks. MiniMax-M3.1-Flash-Preview additionally always thinks and supports tunable thinking depth via output_config.effort. The M2.7, M2.5, M2.1, and M2 series support text and tool calls only and do not accept image or video input.

Available options:
MiniMax-M3.1-Flash-Preview,
MiniMax-M3,
MiniMax-M2.7,
MiniMax-M2.7-highspeed,
MiniMax-M2.5,
MiniMax-M2.5-highspeed,
MiniMax-M2.1,
MiniMax-M2.1-highspeed,
MiniMax-M2
messages
object[]
required

Conversation history. MiniMax-M3.1-Flash-Preview and MiniMax-M3 support text, image, video, tool use, tool result, and thinking content blocks. The M2.7, M2.5, M2.1, and M2 series support text and tool-call content blocks only; they do not support image or video input.

service_tier
enum<string>
default:standard

Service tier for request admission. Supported values are standard and priority. If omitted, the request uses the standard tier. The priority price is 1.5 times the standard price and ensures priority admission so the request is processed ahead of other requests, leading to faster responses and fewer failures.

Available options:
standard,
priority
system

Set the role and behavior of the model.

stream
boolean
default:false

Whether to use streaming output, defaults to false. When set to true, the response will be returned in chunks

max_tokens
integer<int64>

Specifies the upper limit for generated content length (in tokens). For MiniMax-M3.1-Flash-Preview and MiniMax-M3 the recommended value is 131072 (128K) and the maximum is 524288 (512K); for other models the recommended value is 65536 (64K) and the maximum is 204800 (200K). Content exceeding the limit will be truncated. If generation stops due to length, try increasing this value

Required range: x >= 1
temperature
number<double>
default:1

Temperature coefficient, affects output randomness. Range [0, 2], default 1. Higher values produce more random output; lower values produce more deterministic output.

Required range: 0 <= x <= 2
top_p
number<double>
default:0.95

Nucleus sampling parameter. Range [0, 1]. Default is 0.95 for MiniMax-M3.1-Flash-Preview and MiniMax-M3, and 0.9 for M2.x models.

Required range: 0 <= x <= 1
tools
object[]

Tool definitions for Anthropic-compatible tool use.

tool_choice
object

Tool selection strategy. Only auto and none are supported.

thinking
object

Controls thinking behavior. The default depends on the model, so no schema-level default is declared.

  • MiniMax-M3.1-Flash-Preview: always thinks. If type is sent it must be adaptive; disabled returns HTTP 400. Use output_config.effort to tune thinking depth.
  • MiniMax-M3: thinking is off when thinking is omitted; send adaptive to enable it and return thinking blocks.
  • M2.x models: thinking cannot be disabled; disabled is accepted but ignored.
output_config
object

Output configuration. Use effort to tune thinking depth for MiniMax-M3.1-Flash-Preview.

metadata
object

Request metadata. user_id is recommended for end-user-level aggregation, rate limiting, and billing analysis.

Response

id
string

Unique ID of this response

type
enum<string>

Object type, fixed as message

Available options:
message
role
enum<string>

Role, fixed as assistant

Available options:
assistant
model
string

Model ID used for this request

content
object[]

List of response content blocks

stop_reason
enum<string>

Reason for stopping generation:

  • end_turn: Model ended naturally
  • max_tokens: Reached max_tokens limit
  • tool_use: Model requested tool use
Available options:
end_turn,
max_tokens,
tool_use
usage
object

Token usage for this request, including prompt cache usage when applicable.