minimax-m2.7-highspeed
neev_cloud/minimax-m2.7-highspeed
Context
—
Max output
66K
Input / 1M
$57.12
Output / 1M
$228.49
Modality
Chat
Cutoff
Unknown
About this model
MiniMax-M2.7-highspeed delivers the same performance as MiniMax-M2.7 with significantly faster inference and lower latency.
Best suited for
- Ultra-fast conversational AI
- Low-latency coding assistance
- Real-time AI agents
- Long-context reasoning
- High-throughput production workloads
Capabilities
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
temperatureControls randomness and creativity of responses.
top_pControls diversity by limiting token selection probability.
max_tokensMax Tokens LimitMaximum number of tokens generated in the response.
streamStreams tokens in real-time while generating.
toolsDefines external tools available to the model.
tool_choiceDetermines whether the model can use tools.
response_typeSpecifies the output response format.
parallel_tool_callsEnables multiple tools to run simultaneously.