Alibaba
Qwen3.8-Max
About model
Alibaba’s largest and most capable flagship, a 2.4T sparse MoE (95B active) that is natively multimodal over text, image, and video, built for coding, professional work, and long-horizon agentic execution. 1M context
It takes text, images and video as input and returns text. It works with a context window of 1M tokens and can return up to 125k tokens in a single response.
Calls go through the SCX gateway on an OpenAI compatible endpoint, so moving an existing integration across means changing the base URL and the model name, and nothing else. The default tier allows 15,000 requests per minute, and higher limits are available on request.
1M
Tokens the model can attend to in a single request
78.5
Overall score, ranked 6 on the public leaderboard
125k
Tokens the model can return in one response
Model key capabilities
Reasoning
Works through a problem step by step before answering, which lifts accuracy on multi-step maths, logic and planning tasks.
Vision
Accepts images alongside text in the same request, so documents, screenshots and charts can be read directly.
Tool calling
Returns structured arguments for the tools you define, using the same schema as the OpenAI API, so existing agent frameworks work unchanged.
JSON mode
Constrains output to valid JSON, so responses can be parsed without a repair step.
Streaming
Streams tokens as they are generated, so an interface can show output before the response completes.
Benchmarks
Scores on LiveBench, a contamination-free benchmark refreshed every six months. Higher is better.
- Reasoning
- 88.2
- Coding
- 72.9
- Agentic Coding
- 64.6
- Mathematics
- 91.3
- Data Analysis
- 78.4
- Language
- 79.7
- Instruction Following
- 74.1
Source: LiveBench 2026-06-25 release.
Quick start
curl https://api.scx.ai/v1/chat/completions \
-H "Authorization: Bearer $SCX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-Max",
"messages": [{ "role": "user", "content": "Hello" }]
}'