gemma-4-31B-it
About model
Google Gemma 4 31B dense instruction-tuned multimodal model (text and image in) with a toggleable thinking mode, native tool use, 128k context, and 140+ language coverage
It takes text and images as input and returns text. It works with a context window of 128k tokens and can return up to 8k tokens in a single response. Weights are served at bf16 precision.
Calls go through the SCX gateway on an OpenAI compatible endpoint, so moving an existing integration across means changing the base URL and the model name, and nothing else. It is served entirely from Australian datacentres, so prompts and completions stay onshore and never cross a border in transit or at rest.
128k
Tokens the model can attend to in a single request
8k
Tokens the model can return in one response
Model key capabilities
Reasoning
Works through a problem step by step before answering, which lifts accuracy on multi-step maths, logic and planning tasks.
Vision
Accepts images alongside text in the same request, so documents, screenshots and charts can be read directly.
Tool calling
Returns structured arguments for the tools you define, using the same schema as the OpenAI API, so existing agent frameworks work unchanged.
JSON mode
Constrains output to valid JSON, so responses can be parsed without a repair step.
Streaming
Streams tokens as they are generated, so an interface can show output before the response completes.
Sovereign hosting
Served entirely from Australian datacentres on SCX infrastructure, so request and response data does not leave shore.
Quick start
curl https://api.scx.ai/v1/chat/completions \
-H "Authorization: Bearer $SCX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-31B-it",
"messages": [{ "role": "user", "content": "Hello" }]
}'