Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
Added Jul 23, 2026
Model weightsContext Window
262.1K
Max Output
32.8K
Input Price (Auto)
$0.075/1M
Output Price (Auto)
$0.22/1M
Cache Read (Auto)
$0.015/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
37.8
Coding Index
50.6
Agentic Index
29.3
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
85.5%
Better than 88% of models compared
HLE
Humanity's Last Exam
23.7%
Better than 87% of models compared
AA-LCR
Long context reasoning evaluation
67.0%
Better than 89% of models compared
GDPval-AA
Economically valuable tasks
30.4%
CritPt
Research-level physics reasoning
1.7%
Coding
SciCode
Python programming for scientific computing
41.1%
Better than 82% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
18.2%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
44.1%
Last updated Aug 7, 2026, 12:00 PM
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Ling 3.0 Flash with similar models from the same provider or model family.
Ling 3.0 Flash Thinking
inclusionai/ling-3.0-flash:thinkingLing-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.
Ling 2.6 Flash
inclusionai/ling-2.6-flashLing-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency. It delivers performance comparable to state-of-the-art models at a similar scale while significantly reducing token usage across coding, document processing, and lightweight agent workflows.
Ling 2.6 1T
inclusionai/ling-2.6-1tLing-2.6-1T is an inclusionAI instruction model optimized for large-scale agentic and coding workloads with long-context support and structured output capabilities.
Ring 2.6 1T
inclusionai/ring-2.6-1tRing-2.6-1T is an inclusionAI thinking model for real-world agent workflows, coding agents, tool use, and long-horizon task execution.
Gemini 3.7 Flash
google/gemini-3.7-flashGoogle's frontier-performance Flash model for multimodal and agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Google reports 65.3% on DeepSWE v1.1, up from 49.0% for Gemini 3.6 Flash, and 34% on GDP.pdf, up from 14%.
DeepSeek V4 Flash Latest
deepseek/deepseek-v4-flash-latestCompatibility alias that routes to the newest dated DeepSeek V4 Flash release. Currently uses Fireworks first for DeepSeek V4 Flash 0731, with automatic failover when needed. ⚠️ Privacy and logging guarantees are limited.