Qwen3.8-Max: The 2.4T Parameter Autonomous Giant
Alibaba's Qwen Team just released Qwen3.8-Max — a 2.4-trillion-parameter MoE flagship that doesn't just answer questions, it executes projects. Autonomous coding, multimodal reasoning, and benchmark dominance. Here's what you need to know.
The Big Picture
On August 3, 2026, Alibaba's Qwen Team released Qwen3.8-Max — a 2.4-trillion-parameter Mixture-of-Experts (MoE) flagship model. This isn't an incremental update. It's designed from the ground up for autonomous agency: managing software repositories, navigating complex workflows, and delivering complete projects that previously required human oversight.
Architecture: Massive Scale, Efficient Inference
The 2.4T parameter MoE design allows Qwen3.8-Max to maintain deep specialized knowledge across diverse domains — finance, law, software engineering, design — while activating only a fraction of its parameters per token. This means you get near-dense-model quality at a fraction of the compute cost.
Multimodal capabilities are natively integrated, not bolted on. The model's visual understanding is deeply intertwined with its reasoning engine, enabling end-to-end workflows: planning, execution, and verification across text, images, documents, and video.
The 10-Day Developer: Autonomous Software Engineering
Here's where it gets remarkable. Qwen3.8-Max has demonstrated the ability to autonomously code and deliver complete projects spanning over 10 days of human work — starting from an empty folder and ending with a finished, tested product.
oh-my-cli project (qwen-code-dev-bot/oh-my-cli) was built entirely by the model from scratch — a self-evolving engineering harness with a dispatcher, monitor, and watchdog, including automated CI/CD with Build, Unit Test, E2E, and Desktop Lifecycle validation.
If an abnormal state is detected, the model automatically routes the error back to the relevant issue for a fix. It doesn't just write code — it manages the entire engineering lifecycle.
Benchmark Supremacy
Qwen3.8-Max has been battle-tested in high-stakes competitive environments:
| Benchmark | Score | Notes |
|---|---|---|
| Terminal Bench 2.1 | 86.6 | Coding agent tasks |
| FrontierSWE | 73.5 | Software engineering |
| GPQA Diamond | 92.6 | Graduate-level science |
| PaperBench | 93.0 | Top in class |
| WideSearch | 81.9 | Web reasoning |
| Toolathlon Verified | 72.5 | Tool-use scenarios |
Pricing
| Resource | Price |
|---|---|
| Input | $2.00 / 1M tokens |
| Output | $6.00 / 1M tokens |
| Explicit cache creation | $2.50 / 1M tokens |
| Explicit cache read | $0.17 / 1M tokens |
The aggressive caching pricing makes Qwen3.8-Max ideal for long-context applications and repetitive agentic loops where the same context is reused across iterations.
New users get $4 credit with invite code. Sign up and start building with Qwen3.8-Max today.
Qwen Cloud Coding Plans
Start with $4 free credit, then pick a plan that matches your agent workload. All plans include access to multiple models and harness tools.
- 2,500 credits per 7 days
- 700 credits per 5 hours
- Access multiple models
- Access harness tools
- Run 1–2 agents concurrently
- 10,000 credits per 7 days
- 3,000 credits per 5 hours
- 4× the credits of Lite
- Everything in Lite
- Run 3–4 agents concurrently
- 40,000 credits per 7 days
- 12,000 credits per 5 hours
- 16× the credits of Lite
- Everything in Standard
- Run 6–8 agents concurrently
Use Cases
- Autonomous Software Engineering: Handle tickets, write tests, maintain CI/CD pipelines — the model acts as a "Junior Engineer Plus."
- Legal & Financial: End-to-end document analysis and compliance checking with multimodal understanding.
- Design: Interpret visual briefs and iterate on assets with native multimodal capabilities.
- Advanced RAG & General Agents: Long-context retrieval, web search, and complex tool-use scenarios.
Getting Started
Developers can access Qwen3.8-Max immediately via the Alibaba Cloud API or Qwen Cloud platform. Official SDKs are available for Python and TypeScript. The model supports standard RESTful endpoints for seamless integration into existing workflows.