AI Assistants
DeepSeek vs Kimi vs Doubao vs Tongyi: which Chinese LLM should you actually use
Four Chinese LLMs have crossed the frontier line on coding and reasoning, with pricing a tenth of GPT-5 or Claude Opus. DeepSeek, Kimi, Doubao, and Tongyi differ on context window, ecosystem, and whether you can use them outside China. Here's how to pick.
| Dimension | DeepSeek (深度求索) Open weights, strongest price-to-performance | Kimi (月之暗面) 256K context, agentic coding, strong Deep Research | Doubao (字节跳动) Multimodal, adaptive thinking, low latency | Tongyi / Qwen (阿里巴巴) Open ecosystem, 1M-token window on flagship models |
|---|---|---|---|---|
| Developer / company | DeepSeek AI (Hangzhou) | Moonshot AI (北京) | ByteDance / 火山引擎 (Volcano Engine) | Alibaba Cloud / DashScope |
| Flagship model (Sep 2026) | DeepSeek V3.2 Think | Kimi K3 / K2.7 Code | Doubao Seed 2.1 Pro | Qwen 3.7-Plus / 3.8-Max |
| Context window | 128K (164K via third-party hosts) | 256K (1.05M on K3) | 256K (1M on weekly evolving) | 1M (256K on Qwen 3.7-Plus) |
| Coding performance | Strong (37% LiveCodeBench, 42% SWE-bench Verified) | Strong on long-horizon agentic coding (K2.7 Code) | Solid, good front-end and adaptive thinking | Strong, SOTA on SWE-bench among open weights (Qwen3-Coder) |
| Math / reasoning | Strong (V3.2 Thinking mode for hard problems) | Strong (K3 ranks high on public leaderboards) | Solid, with selectable thinking lengths | Strong (Qwen 3.8-Max), strong on AIME-style tasks |
| Chinese language quality | Excellent — idiomatic, mainstream vocabulary | Excellent — long-context Chinese synthesis is the headline | Excellent — strongest at colloquial / internet Chinese | Excellent — strongest on formal / classical Chinese |
| Multimodal input | Text only (image understanding is hosted-only) | Image input (Kimi K2.6+) | Image + video understanding on Seed 1.6+ | Image + video understanding (Qwen 3.7-Plus / Qwen3-VL) |
| API pricing (per 1M tokens) | $0.26 input / $0.89 output (cache hits $0.07) | $0.95 input / $4 output (cache hits $0.16) | $0.25 input / $2 output (CN ¥6 / ¥30) | $0.28–0.40 input / $1.10–1.60 output |
| Free consumer tier | Yes (deepseek.com web + app, low quotas) | Yes — Adagio tier, 6 agent tasks | Yes (in 豆包 client), generous quotas | Yes (chat.qwen.ai + Qwen Work), generous |
| API availability outside China | Yes — OpenAI-compatible at api.deepseek.com | Yes — api.moonshot.ai OpenAI-compatible | Via BytePlus ModelArk (international OpenAI-compatible) | Yes — Singapore, Frankfurt, Virginia, Tokyo regions |
| Open weights | Yes (V3 base, with MIT-ish terms) | Yes (K2 series, modified MIT) | No — closed weights | Yes (Qwen3, Qwen3-Coder — Apache 2.0) |
| Best for | Cost-sensitive API workloads, self-hosted open-weight pipelines | Agentic coding on a budget, long-document Chinese research | Multimodal + low-latency assistants integrated with Douyin / Lark | Open ecosystem, full-stack on Alibaba Cloud, 1M-token window |
Bottom line
If cost matters most and you can live with text-only, DeepSeek V3.2 Think is the rational default — open weights, frontier-class reasoning, and the cheapest cache-hit pricing of the four. If you need long-context Chinese research or agentic coding, Kimi K3 / K2.7 Code is the better pick at a 3-4× price premium. For multimodal and the most generous free consumer tier, Doubao Seed 2.1 Pro is hard to beat. For open ecosystems and 1M-token context on a flagship non-thinking model, Tongyi Qwen3-Plus / Max is the obvious choice. Most production setups stack two: DeepSeek for default cost-sensitive workloads, and one of Kimi / Qwen for the long-context or open-weights layer.
Full pages for each side