curl --request POST \
--url https://api.minimax.io/anthropic/v1/messages \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: <content-type>' \
--data '
{
"model": "MiniMax-M3.1-Flash-Preview",
"thinking": {
"type": "adaptive"
},
"messages": [
{
"role": "user",
"content": "Which is bigger, 9.11 or 9.9?"
}
],
"max_tokens": 4096,
"output_config": {
"effort": "max"
}
}
'{
"id": "066b367ba949a58a9d6829b4bf8cbcd4",
"type": "message",
"role": "assistant",
"model": "MiniMax-M3.1-Flash-Preview",
"content": [
{
"thinking": "The user is asking which is bigger, 9.11 or 9.9.\n\n9.11 vs 9.9\n\n9.9 is greater than 9.11.\n\nTo compare: 9.11 = 9.11 and 9.9 = 9.90\n\n9.90 > 9.11, so 9.9 is bigger.",
"signature": "6d0315c818f9664ff185dabaa22cd89f2bf28a3a52122095bce23d905471ec5f",
"type": "thinking"
},
{
"text": "**9.9 is bigger than 9.11.**\n\nTo compare decimals, it helps to write them with the same number of decimal places:\n- 9.11 = 9.11\n- 9.9 = 9.90\n\nSince 9.90 > 9.11, **9.9 is larger**.",
"type": "text"
}
],
"usage": {
"input_tokens": 13,
"output_tokens": 149,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 159
},
"stop_reason": "end_turn"
}Messages API
Use the Anthropic API compatible Messages format to call MiniMax models.
curl --request POST \
--url https://api.minimax.io/anthropic/v1/messages \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: <content-type>' \
--data '
{
"model": "MiniMax-M3.1-Flash-Preview",
"thinking": {
"type": "adaptive"
},
"messages": [
{
"role": "user",
"content": "Which is bigger, 9.11 or 9.9?"
}
],
"max_tokens": 4096,
"output_config": {
"effort": "max"
}
}
'{
"id": "066b367ba949a58a9d6829b4bf8cbcd4",
"type": "message",
"role": "assistant",
"model": "MiniMax-M3.1-Flash-Preview",
"content": [
{
"thinking": "The user is asking which is bigger, 9.11 or 9.9.\n\n9.11 vs 9.9\n\n9.9 is greater than 9.11.\n\nTo compare: 9.11 = 9.11 and 9.9 = 9.90\n\n9.90 > 9.11, so 9.9 is bigger.",
"signature": "6d0315c818f9664ff185dabaa22cd89f2bf28a3a52122095bce23d905471ec5f",
"type": "thinking"
},
{
"text": "**9.9 is bigger than 9.11.**\n\nTo compare decimals, it helps to write them with the same number of decimal places:\n- 9.11 = 9.11\n- 9.9 = 9.90\n\nSince 9.90 > 9.11, **9.9 is larger**.",
"type": "text"
}
],
"usage": {
"input_tokens": 13,
"output_tokens": 149,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 159
},
"stop_reason": "end_turn"
}MiniMax-M3.1-Flash-PreviewCore capabilities: 1M long context, multimodal, tunable thinking depth.Available only through M Plan and MiniMax Code for now.MiniMax-M3.1-Flash-Preview:- Image and video understanding — see the example code on the right
- Tune thinking depth via
output_config.effort, which acceptslow,medium,high,xhigh, andmax; the default ismax - Thinking is on by default with no configuration needed; thinking content is returned separately in
thinkingcontent blocks
Authorizations
Bearer API Key auth. Send Authorization: Bearer <API_KEY>. If Authorization and x-api-key are both present, Authorization takes precedence.
Headers
Media type of the request body, should be set to application/json to ensure JSON format
application/json Body
Model ID. MiniMax-M3.1-Flash-Preview and MiniMax-M3 are multimodal models with native support for text, image, and video input, alongside tool use and thinking content blocks. MiniMax-M3.1-Flash-Preview additionally always thinks and supports tunable thinking depth via output_config.effort. The M2.7, M2.5, M2.1, and M2 series support text and tool calls only and do not accept image or video input.
MiniMax-M3.1-Flash-Preview, MiniMax-M3, MiniMax-M2.7, MiniMax-M2.7-highspeed, MiniMax-M2.5, MiniMax-M2.5-highspeed, MiniMax-M2.1, MiniMax-M2.1-highspeed, MiniMax-M2 Conversation history. MiniMax-M3.1-Flash-Preview and MiniMax-M3 support text, image, video, tool use, tool result, and thinking content blocks. The M2.7, M2.5, M2.1, and M2 series support text and tool-call content blocks only; they do not support image or video input.
Show child attributes
Show child attributes
Service tier for request admission. Supported values are standard and priority. If omitted, the request uses the standard tier. The priority price is 1.5 times the standard price and ensures priority admission so the request is processed ahead of other requests, leading to faster responses and fewer failures.
standard, priority Set the role and behavior of the model.
Whether to use streaming output, defaults to false. When set to true, the response will be returned in chunks
Specifies the upper limit for generated content length (in tokens). For MiniMax-M3.1-Flash-Preview and MiniMax-M3 the recommended value is 131072 (128K) and the maximum is 524288 (512K); for other models the recommended value is 65536 (64K) and the maximum is 204800 (200K). Content exceeding the limit will be truncated. If generation stops due to length, try increasing this value
x >= 1Temperature coefficient, affects output randomness. Range [0, 2], default 1. Higher values produce more random output; lower values produce more deterministic output.
0 <= x <= 2Nucleus sampling parameter. Range [0, 1]. Default is 0.95 for MiniMax-M3.1-Flash-Preview and MiniMax-M3, and 0.9 for M2.x models.
0 <= x <= 1Tool definitions for Anthropic-compatible tool use.
Show child attributes
Show child attributes
Tool selection strategy. Only auto and none are supported.
Show child attributes
Show child attributes
Controls thinking behavior. The default depends on the model, so no schema-level default is declared.
MiniMax-M3.1-Flash-Preview: always thinks. Iftypeis sent it must beadaptive;disabledreturns HTTP 400. Useoutput_config.effortto tune thinking depth.MiniMax-M3: thinking is off whenthinkingis omitted; sendadaptiveto enable it and return thinking blocks.- M2.x models: thinking cannot be disabled;
disabledis accepted but ignored.
Show child attributes
Show child attributes
Output configuration. Use effort to tune thinking depth for MiniMax-M3.1-Flash-Preview.
Show child attributes
Show child attributes
Request metadata. user_id is recommended for end-user-level aggregation, rate limiting, and billing analysis.
Show child attributes
Show child attributes
Response
Unique ID of this response
Object type, fixed as message
message Role, fixed as assistant
assistant Model ID used for this request
List of response content blocks
Show child attributes
Show child attributes
Reason for stopping generation:
- end_turn: Model ended naturally
- max_tokens: Reached max_tokens limit
- tool_use: Model requested tool use
end_turn, max_tokens, tool_use Token usage for this request, including prompt cache usage when applicable.
Show child attributes
Show child attributes