Alibaba Prices Qwen3.8-Flash API At $0.16 Per Million Tokens, Cutting Inference Costs For 125B-Parameter Model

The Qwen Team has released Qwen3.8-Flash-Next, an open-weight multimodal mixture-of-experts model that serves as an early architectural preview for the upcoming Qwen4 series. The model balances substantial capacity with exceptional cost efficiency, featuring 125 billion total parameters and an additional 51 billion N-gram embedding parameters while activating only 6 billion parameters per token.
The release follows the precedent established by Qwen3-Next, which introduced hybrid architecture designs later adopted across the Qwen3.5 through Qwen3.8 families. Qwen3.8-Flash-Next will be available through the QwenCloud API at a rate of $0.16 per million input tokens and $0.47 per million output tokens, positioning it as a highly competitive option for high-volume applications, coding assistants, and enterprise agentic workflows.
Benchmark results indicate strong performance across software engineering and autonomous agent tasks. The model achieved 58.7% on DeepSWE 1.1, 62.5% on SWE-bench Pro, and 81.0% on the multilingual variant. In long-horizon office automation measured by CoWorkBench, it scored 73.9%, surpassing both Qwen3.7-Plus and Claude-Opus-4.6.
General reasoning capabilities remain robust, with scores of 91.7% on GPQA Diamond and 91.9% on LiveCodeBench v6. Multimodal competence is equally solid, with 84.5% on AndroidWorld and 76.6% on LVBench for long video understanding. Native context length reaches 262,144 tokens, extensible to one million tokens via YaRN.
Architectural Innovations and Developer Integration
The architecture introduces four systematic upgrades. For attention, the model combines Gated DeltaNet with Qwen Sparse Attention, compressing historical context efficiently while retrieving relevant information through micro-block indexing rather than token-level processing. This design achieves up to 7.6× prefill speedup at one million tokens.
The Gated Residual mechanism expands the residual stream into four parallel branches with dynamic gating, improving cross-layer information flow and training stability while supporting FP8 storage for reduced memory traffic. N-gram Embedding adds capacity through local-context lookups that require minimal computation and can be asynchronously prefetched from host memory. Training employs the Muon optimizer with refined orthogonalization and parameter-splitting strategies, enabling stable convergence at larger batch sizes without traditional warmup procedures.
The model weights are available on HuggingFace and ModelScope, with API access through QwenCloud supporting OpenAI-compatible Chat Completions and Anthropic-compatible protocols. Developers can integrate the model into existing workflows through Claude Code, OpenAI Codex, Qoder CLI, Qwen Code, and OpenClaw, with reasoning effort configurable across low, medium, and xhigh levels. An official production release with built-in tools and default one-million-token context is expected to follow shortly.
Disclaimer
In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.
About The Author
Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.
More articles
Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.



