scx.ai logo

Engineering Blog

Same Model, Three Platforms: What Function Calling Benchmarks Reveal

When it comes to function calling and structured output generation, not all platforms deliver equal results — even with identical models. SCX.ai's infrastructure advantages translate to measurably superior performance.

By SCX.ai9 min read
Simple Function Calls98%Multiple Functions95%Multi-turn Accuracy35%SCX.ai • DeepSeek-V3 Function Calling Performance

When it comes to function calling and structured output generation, not all platforms deliver equal results — even with identical models. SCX.ai's infrastructure advantages translate to measurably superior performance across simple tasks, complex scenarios, and production-critical structured output generation.

Why Function Calling Accuracy Matters in Production

Function calling transforms language models from conversational interfaces into actionable systems. When a user asks "What's the weather in Sydney?", function calling allows the model to invoke a weather API, retrieve data, and present the answer — all autonomously. It's the bridge between natural language and structured digital tools.

This is harder than it sounds. The model must select the right function from a set of options, extract parameters from ambiguous input, and format the output precisely. Chain multiple calls together or add multi-turn context, and error rates climb quickly.

Here's what is often overlooked: Function calling performance varies by provider, even for the same model. Each platform implements its own inference stack, tool parsing, and output formatting. The model weights are identical, but the infrastructure around them isn't.

A 90% accuracy rate sounds impressive until you realise that 1 in 10 customer requests fail.

Putting Performance to the Test

To understand how platform infrastructure affects function calling and structured output performance, we conducted comprehensive testing using two industry-standard benchmark suites. We evaluated the same models — DeepSeek-V3 and Llama-4-Maverick-17B — across multiple inference providers: Fireworks, Together AI, and SCX.ai.

The Berkeley Function Calling Leaderboard evaluates models across increasingly complex scenarios:

  • Simple function calls (single function, straightforward parameters)
  • Multiple function selection (choosing from many options)
  • Multi-turn conversations (function calls across dialogue exchanges)

The JSON Schema Bench tests structured output generation — critical for ensuring function parameters are correctly formatted across diverse real-world schemas from GitHub repositories, data visualisation tools, and production applications.

DeepSeek-V3 Results

DeepSeek-V3 Function Calling AccuracySCX.aiFireworksTogether AISimple Function Calls98%94%96%Multiple Function Selection95%94%89%Multi-turn Conversations35%4%Key InsightSCX.ai delivers superior performance across all function calling complexity levels with DeepSeek-V3,achieving up to 31 percentage points higher accuracy in multi-turn scenarios.

When testing DeepSeek-V3, one of the most advanced open-source models available, the platform advantage becomes unmistakable.

Simple Function Calls: The Baseline

Even with straightforward single-function calls, infrastructure optimisation delivers measurable gains. SCX.ai achieved 98% accuracy with DeepSeek-V3 — the highest among all providers. Together AI reached 96%, while Fireworks achieved 94%.

Multiple Function Scenarios: Complexity Amplifies the Gap

The performance advantage grows more pronounced as tasks increase in complexity. When models must select from multiple available functions and coordinate their execution, SCX.ai's infrastructure optimisation shines. With DeepSeek-V3, SCX.ai reached 95% accuracy in multiple function calling scenarios, compared to 94% for Fireworks and 89% for Together AI.

That 6 percentage point lead over Together AI translates to 40% fewer errors — a dramatic improvement in system reliability that directly impacts user satisfaction and operational costs.

Multi-Turn Conversations: The Ultimate Challenge

The most demanding test involves maintaining function calling accuracy across multi-turn conversations, where context must persist and models must reference previous interactions. This remains challenging industry-wide, but the infrastructure gap is stark.

SCX.ai achieved 35% accuracy with DeepSeek-V3 on multi-turn function calling — dramatically outperforming Together AI's 4%. While absolute performance levels indicate this remains a frontier challenge, SCX.ai's 31 percentage point advantage demonstrates how infrastructure optimisation enables capabilities that are otherwise nearly impossible.

Structured Output Excellence: JSON Schema Performance

Function calling reliability fundamentally depends on generating perfectly formatted JSON. A misplaced bracket, incorrect data type, or malformed structure breaks the entire execution chain. The JSON Schema Bench evaluates this critical capability across diverse real-world schemas.

DeepSeek-V3 JSON Schema CoverageSCX.aiFireworksTogether AIGitHub Easy Schemas92%89%80%GitHub Medium Schemas80%74%66%Snowplow Schemas86%81%21%JSON Schema Store30%24%18%Key InsightSCX.ai leads across all JSON schema categories with DeepSeek-V3, demonstrating superior structured output generation for production applications.GitHub Easy+12%vs Together AIGitHub Medium+14%vs Together AISnowplow+65%vs Together AISchema Store+12%vs Together AI

Testing DeepSeek-V3 across four distinct schema categories, SCX.ai led in all four:

GitHub Easy Schemas: SCX.ai achieved 92% coverage, compared to 89% for Fireworks and 80% for Together AI.

GitHub Medium Schemas: SCX.ai reached 80% coverage, outperforming Fireworks (74%) and Together AI (66%) by substantial margins.

Snowplow Schemas (complex data visualisation): SCX.ai delivered 86% coverage — 65 percentage points ahead of Together AI's 21% and 5 points ahead of Fireworks' 81%.

JSON Schema Store (diverse real-world schemas): SCX.ai achieved 30% coverage, the highest among all providers, with Fireworks at 24% and Together AI at 18%.

The pattern is unmistakable: Across every category and complexity level tested, SCX.ai's infrastructure enables measurably superior structured output generation with DeepSeek-V3.

How Infrastructure Affects Results — Same Model, Different Outcomes

A reasonable question emerges from these benchmarks: How can identical model weights produce different accuracy results across platforms?

The answer lies not in the model itself, but in everything surrounding it. When you deploy a model like DeepSeek-V3, the neural network parameters are identical across providers. But the inference stack — the software, hardware, and configuration choices that turn those weights into responses — varies significantly.

Numerical Precision

Different platforms use different precision formats during inference:

  • FP32 (full precision) — highest accuracy, highest compute cost
  • FP16/BF16 (half precision) — balanced performance and accuracy
  • FP8/INT8 (quantised) — fastest and cheapest, but can degrade output quality

When a platform uses aggressive quantisation to reduce costs, the model's ability to produce precise JSON structures and select correct functions degrades. Small numerical errors compound through billions of parameters, manifesting as malformed outputs or incorrect function selections.

SCX.ai's ASIC-based infrastructure maintains higher numerical precision during inference without the power and cost penalties that force GPU-based providers toward aggressive quantisation.

System Prompts and Tool Formatting

Every platform injects its own system prompts and tool definitions when handling function calling requests. These hidden instructions significantly impact accuracy:

Platform A: "You are a helpful assistant. Available functions: {schema}..."
Platform B: "When calling functions, respond ONLY with valid JSON. Tools: {schema}..."
Platform C: "You have access to the following tools. Use them when appropriate..."

Better prompt engineering — refined through extensive testing — produces better function calling accuracy. This is pure software optimisation, not model capability.

Grammar-Constrained Decoding

Some platforms implement grammar-constrained decoding — forcing the model's output to conform to valid JSON syntax at the token level. Others simply hope the model produces valid output and attempt repairs afterward.

Constrained decoding is computationally expensive but dramatically improves structured output reliability. Platforms with more efficient infrastructure can afford to run these constraints without sacrificing latency or throughput.

KV-Cache Management for Multi-Turn

The dramatic gap in multi-turn accuracy (35% vs 4%) primarily reflects memory management differences.

During multi-turn conversations, models must maintain a key-value cache of previous context. When memory pressure forces cache eviction, the model loses access to earlier function calls and their results. The conversation effectively "forgets" what happened before.

ASIC architectures like the SN40L chips in SCX.ai's infrastructure feature unified memory designs optimised for maintaining large context windows without eviction. GPU-based platforms, optimised for training workloads, often struggle with the memory access patterns required for efficient multi-turn inference.

The Compound Effect

These factors don't operate in isolation — they compound:

FactorImpact on Accuracy
Numerical precisionMedium–High
System prompt tuningHigh
Grammar-constrained decodingHigh
KV-cache managementVery High (multi-turn)
Output parsing and retry logicMedium

A platform that maintains high precision, uses well-tuned prompts, implements constrained decoding, and manages context efficiently will outperform one that cuts corners in any of these areas — even with identical model weights.

The hardware enables better software choices. More efficient silicon means platforms can maintain precision, run constrained decoding, and keep larger context windows — without the power costs that force compromises elsewhere.

Why Function Calling Is Your AI's Make-or-Break Feature

Function calling turns language models from purely conversational systems into actionable, agentic systems. As enterprises shift toward agentic AI — where models plan, decide, and execute across tools — function calling becomes the critical bridge that connects natural language reasoning to real-world applications.

Function Calling ApplicationsEnterprise AutomationProcessing invoicesUpdating CRM systemsTriggering workflowsData AnalysisQuerying databasesGenerating reportsExtracting structured infoCustomer ServiceBooking appointmentsChecking order statusProcessing returnsDevelopment ToolsWriting codeExecuting testsDeploying applications

Production Implications

Reduced Failure Rates: Fewer failed calls means less retry logic, less error handling, and fewer frustrated users.

Lower Operational Costs: Failed calls consume tokens, trigger retries, and sometimes require human intervention. Even a few percentage points of accuracy improvement compounds over thousands of daily requests.

Faster Development Cycles: The gap between 89% and 98% accuracy is the difference between debugging edge cases and shipping features.

Predictability: Production systems need consistent results, not just good average performance.

The Sovereign Advantage

Beyond raw performance, SCX.ai delivers these benchmarks with full data sovereignty. Every function call, every JSON schema validation, every multi-turn conversation stays onshore — meeting Australian regulatory requirements while delivering world-class accuracy.

For organisations in regulated industries — finance, healthcare, government — this combination of performance and sovereignty isn't optional. It's essential.

Ready to experience the SCX.ai difference? Explore our platform or contact our team to discuss your production AI requirements.

Related Topics

function callingAI benchmarksDeepSeek-V3Llama-4-Maverickstructured outputJSON schemaagentic AIinference performancesovereign AI
Same Model, Three Platforms: What Function Calling Benchmarks Reveal