BTC/USD $77,837.00 ▲ +0.75%
ETH/USD $2,576.90 ▲ +4.86%
BRENT OIL $72.80 ▼ -0.6%
EUR/USD 1.0842 ▲ +0.12%
GOLD $2,310.40 ▼ -0.3%
UPDATED: 19:28 CEST DATA: COINGECKO API
🕒 17:31:26 UTC
[TECH] 4 MIN READ

Alibaba and DeepSeek Accelerate China’s AI Race Toward Lower Inference Costs

BY [ABOUT DESK]
Optimumline
Optimumline EDITORIAL DESK AUTHOR

Podstara editorial desk editor covering cryptocurrency markets, macroeconomics, regional energy infrastructure, and industrial technology.

VIEW AUTHOR ARCHIVE →
PUBLISHED: Aug 7, 2026
FONT:
Alibaba and DeepSeek Accelerate China’s AI Race Toward Lower Inference Costs

BEIJING — Alibaba and DeepSeek are escalating competition in China’s artificial intelligence sector, rolling out new models designed to drastically lower operational costs while maintaining high-level capabilities.

Alibaba recently launched Qwen3.8-Max, its largest model to date, while DeepSeek’s latest V4-Flash model is drawing industry attention for inference pricing that undercuts several competing systems.

Sparse Architecture Drives Efficiency

Qwen3.8-Max features 2.4 trillion total parameters built on a sparse Mixture-of-Experts (MoE) architecture. By activating only a fraction of its network for any given request—roughly 95 billion active parameters per token—Alibaba significantly reduces operating costs and response latency compared to running a full-scale dense model.

DeepSeek employs a similar sparse MoE approach at a smaller scale. According to industry tracker Artificial Analysis, V4-Flash operates with 284 billion total parameters, of which only 13 billion are active during inference. By comparison, Moonshot AI’s Kimi K3 utilizes 2.8 trillion total parameters with approximately 104 billion active.

In addition to text, Qwen3.8-Max processes visual inputs and supports a 1-million-token context window. Alibaba also reported that the model successfully completed a complex, multi-week software engineering project during internal benchmark testing.

The model enters direct competition with Moonshot AI’s Kimi K3, which debuted in July. The two tech giants are competing aggressively on pricing:

  • Qwen3.8-Max: $2.00 per million input tokens / $6.00 per million output tokens.
  • Kimi K3: $3.00 per million input tokens / $15.00 per million output tokens.

Following its release, Qwen3.8-Max climbed to the top position among Chinese text models on the crowdsourced platform Arena.AI, trailing only top-tier Anthropic models overall. It also secured second place on Arena.AI’s visual evaluation leaderboard, behind Anthropic’s Claude Fable 5.

DeepSeek Drives Down Benchmark Costs

Rather than matching the massive scale of Alibaba and Moonshot AI, DeepSeek focused on aggressive pricing with V4-Flash.

V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens. Data from Artificial Analysis shows that utilizing prompt caching reduces the input cost to $0.003 per million tokens—a 98% discount off the standard rate.

These lower rates translate to dramatic cost savings in benchmark evaluations:

  • DeepSeek V4-Flash: ~$0.03 per test run
  • Moonshot Kimi K3: ~$0.86 per test run
  • OpenAI GPT-5.6 Sol: ~$1.86 per test run
  • Anthropic Claude Fable 5: ~$3.15 per test run

V4-Flash earned a score of 40 on Artificial Analysis’ Intelligence Index while generating output at a speed of approximately 118 tokens per second.

Beyond Token Pricing: Total Task Costs

Industry analysts emphasize that advertised API rates tell only part of the financial story. Model architecture, output length, and the total number of interaction turns required for complex tasks heavily influence real-world expenses.

For example, on the AA-Briefcase benchmark for autonomous knowledge work, Moonshot AI’s Kimi K3 averaged $10.57 per task, generating roughly 120,000 output tokens across 83 turns. Despite higher task-level costs due to its thorough output generation, Kimi K3 scored 57 on the Intelligence Index and achieved the second-highest score on the AA-Briefcase evaluation.

This contrast illustrates why enterprise developers increasingly evaluate cost-per-task metrics rather than relying solely on headline API rates.

Open Weights Offer Strategic Flexibility

Deployment flexibility is another key driver shaping the Chinese AI market. Unlike Western developers like OpenAI, Anthropic, and Google—which maintain closed ecosystem access—Alibaba, DeepSeek, and Moonshot AI continue to offer open-weight releases alongside hosted APIs.

DeepSeek V4-Flash is distributed as an open-weight model under an MIT license on Hugging Face, while Kimi K3 is accessible under Moonshot AI’s open license. Alibaba has also committed to releasing open weights for Qwen3.8-Max.

Open-weight access allows enterprise teams to deploy models on private infrastructure or third-party cloud services, eliminating sole dependency on a single managed API.

“Many business workflows do not need the industry’s very best model,” noted Lian Jye Su, Chief Analyst at Omdia. “They need models that are good enough, affordable, transparent, and accessible—and open-weight models help meet that demand.”

ℹ FTC & Amazon Associates Disclosure: Podstara.com is a participant in affiliate advertising programs. Articles within the Shop section may contain affiliate links, which yield a small commission on qualifying purchases at zero extra cost to you. Editorial news reporting remains strictly independent.

RELATED TECH SIGNALS VIEW ALL →

Discover more from

Subscribe now to keep reading and get access to the full archive.

Continue reading