scx.ai logo
All pricing

DeepSeek

DeepSeek-V4.1-flash

GlobalChatReasoningVision

About model

DeepSeek’s first causal encoder-decoder model: a 552B multimodal MoE (8B active on prefill, 16B on decode) built for input-heavy agentic work, with text and image input, adjustable reasoning effort, tool use and a 1M context

It takes text and images as input and returns text. It works with a context window of 1M tokens and can return up to 375k tokens in a single response. Weights are served at fp8 precision.

Calls go through the SCX gateway on an OpenAI compatible endpoint, so moving an existing integration across means changing the base URL and the model name, and nothing else.

Try DeepSeek-V4.1-flash today
Context window

1M

Tokens the model can attend to in a single request

Max output

375k

Tokens the model can return in one response

Model key capabilities

  • Reasoning

    Works through a problem step by step before answering, which lifts accuracy on multi-step maths, logic and planning tasks.

  • Vision

    Accepts images alongside text in the same request, so documents, screenshots and charts can be read directly.

  • Tool calling

    Returns structured arguments for the tools you define, using the same schema as the OpenAI API, so existing agent frameworks work unchanged.

  • Streaming

    Streams tokens as they are generated, so an interface can show output before the response completes.

Quick start

curl https://api.scx.ai/v1/chat/completions \
  -H "Authorization: Bearer $SCX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "DeepSeek-V4.1-flash",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'
DeepSeek-V4.1-flash by DeepSeek: Pricing & Specs | SCX.ai