Qwen 3.8 features and pricing explained: 2.4T MoE model, $2/$6 API rates, 1M context, multimodal support, and how it compares to Kimi K3 and Claude.
If you have been tracking the AI model race in 2026, you have probably seen the name Quen 3.8 features and pricing pop up in search results, benchmark discussions, and API pricing tables. The correct spelling is Qwen 3.8, and it is Alibaba’s most ambitious large language model release to date—a 2.4-trillion-parameter sparse Mixture-of-Experts architecture that claims to trail only Anthropic’s Claude Fable 5 in overall capability.
But here is what most early coverage glosses over: Qwen 3.8 is not just one model. It is a family that includes the flagship Qwen3.8-Max and a promised smaller Qwen3.8-27B variant. It launched with both a standard pay-as-you-go API and a credit-based subscription system that can be confusing to navigate. And while Alibaba has made bold benchmark claims, independent verification is still rolling in.
In this guide, we will break down every confirmed feature, the actual pricing you will pay, how to access the model today, and where the open-weight promise stands. Whether you are a developer choosing an API provider, a startup budgeting inference costs, or a researcher evaluating the next model to self-host, this article gives you the complete picture.
What Is Qwen 3.8?

Qwen 3.8 is the latest flagship model family from Alibaba’s Qwen team, previewed on July 19, 2026, at the World AI Conference in Shanghai and officially launched on August 3, 2026. It represents a significant leap from Qwen 3.7-Max, adding confirmed multimodal inputs, a disclosed parameter count, and—for the first time in the Max line—a promised open-weight release.
The flagship Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with approximately 95 billion active parameters per token. That means only about 4% of the total parameter count activates for any given forward pass, keeping inference costs manageable despite the massive scale.
Alibaba positions Qwen 3.8 for four specific workloads: coding, real-world agent tasks, long-horizon projects, and multimodal agents. The model supports both OpenAI-compatible and Anthropic-compatible API specifications, making it relatively easy to port existing integrations.
Architecture and Technical Specifications
Understanding Qwen 3.8’s architecture is essential for evaluating whether it fits your infrastructure and budget.
| Specification | Qwen3.8-Max |
|---|---|
| Total Parameters | 2.4 trillion |
| Active Parameters | ~95 billion per token |
| Architecture | Sparse Mixture-of-Experts (MoE) |
| Context Window | 1 million tokens |
| Max Output | 128,000 tokens |
| Modalities | Text, images, video, documents (confirmed); speech and image generation reported but unsettled |
| API Launch | August 3, 2026 |
| Open Weights | Promised; not yet released |
The 95 billion active parameter count is the number that matters for serving cost and latency. For comparison, Moonshot’s Kimi K3—announced the same week—has 2.8 trillion total parameters with 104 billion active per token. Qwen 3.8 is slightly smaller in both total and active parameters, which could translate to marginally lower inference costs.
Multimodal Capabilities

Qwen 3.8 is Alibaba’s first multimodal model above one trillion parameters. Text plus visual inputs are confirmed. Early reports also mention video, document, speech, and image generation capabilities, but Alibaba has not published a complete spec sheet, so treat the full modality list as provisional.
What this means in practice:
- Text: Standard chat, reasoning, coding, and long-document analysis
- Images: Visual understanding, chart interpretation, and image-based reasoning
- Video: Reported but not fully documented; likely input analysis rather than generation
- Documents: PDF and structured document parsing for enterprise workflows
For developers building agents that need to process screenshots, invoices, or UI elements alongside text, the multimodal support is a genuine step forward from the text-first Qwen 3.7 line.
Benchmark Performance: What the Numbers Actually Show

Alibaba’s headline claim is that Qwen 3.8 is “second only to Fable 5.” Here is what has been published versus what remains unverified.
Published Claims
| Benchmark | Qwen3.8-Max Score | Comparison |
|---|---|---|
| OSWorld-Verified | 86.1 | Ahead of GPT-5.6 Sol Max (83.2), Claude Fable 5 (85.0), Gemini 3.1 Pro (76.2) |
| PaperBench | 93.0 | Highest reported score |
The Caveat
These are Alibaba’s internal numbers. No independent benchmark table from Artificial Analysis, LMArena, or community evaluators had been published as of early August 2026. The “second only to Fable 5” claim is marketing language until third parties reproduce the results.
How good is Qwen 3.8? Alibaba claims Qwen 3.8-Max scores 86.1 on OSWorld-Verified and 93.0 on PaperBench, which would place it ahead of GPT-5.6 Sol Max and Gemini 3.1 Pro on agentic desktop tasks. However, independent benchmarks had not been published as of August 2026.
Qwen 3.8 Pricing: API vs Token Plan

Qwen 3.8 has two distinct pricing structures, and understanding the difference is critical for budgeting.
Standard API Pricing (Pay-As-You-Go)
Launched August 3, 2026, this is the production-ready route for applications and teams:
| Metric | Price |
|---|---|
| Input tokens | $2.00 per million |
| Output tokens | $6.00 per million |
| Cached input | $0.25 per million |
At $8 combined per million tokens (input plus output), Qwen 3.8 undercuts Claude Opus 5’s $30 and GPT-5.6 Sol Standard’s $35 by more than half. It matches Grok 4.5, which also prices at $2/$6.
Token Plan Subscription (Individual Use)
For individuals and hobbyists, Alibaba offers a credit-based subscription with promotional pricing:
| Plan | Promotional Price | Regular Price | 5-Hour Limit | 7-Day Limit |
|---|---|---|---|---|
| Lite | $6/month | $8/month | 700 Credits | 2,500 Credits |
| Standard | $18/month | $25/month | 3,000 Credits | 10,000 Credits |
| Pro | $68/month | $80/month | 12,000 Credits | 40,000 Credits |
Important restrictions on the Token Plan:
- Sliding windows: Reaching either the 5-hour or 7-day credit limit pauses access until older usage falls out of the window
- No automatic overage billing: When credits are exhausted, access stops
- Interactive use only: Terms prohibit automated scripts, application backends, batch processing, and scheduled automation
- Off-peak discount: Additional 80% reduction from 22:00 to 08:00 UTC+8
Token Plan Team Pricing
| Plan | Price Per Seat | Monthly Quota |
|---|---|---|
| Standard | $20/month | 25,000 Credits |
| Pro | $75/month | 100,000 Credits |
| Max | $200/month | 250,000 Credits |
| Shared Package | $700/package | 625,000 Credits (shared across seats) |
Team plans do not use the sliding window system. Conversation data in Team Edition is not used for model training.
How much does Qwen 3.8 cost? The standard API costs $2 per million input tokens and $6 per million output tokens. The Token Plan subscription starts at $6 per month (Lite) for individual interactive use. Team plans start at $20 per seat per month.
How to Access Qwen 3.8 Today
As of August 2026, there are three ways to use Qwen 3.8:
1. Standard API. The production route at $2/$6 per million tokens. Supports OpenAI-compatible and Anthropic-compatible endpoints. Suitable for applications, agents, and automated workflows.
2. Token Plan Subscription. The credit-based route for individuals and teams. Best for evaluation, interactive coding, and light personal use. Not suitable for production backends due to usage restrictions.
3. Qoder and QoderWork. Alibaba’s coding and workplace productivity tools that carry the model. If your use case is agentic coding, this is the fastest way to test the model’s capabilities.
What does not exist yet: Open weights on Hugging Face or ModelScope, third-party inference providers like Together AI or Fireworks AI, and independent benchmark verification.
Qwen 3.8 vs Competitors

| Dimension | Qwen3.8-Max | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Total Parameters | 2.4T | 2.8T | Undisclosed | Undisclosed |
| Active Parameters | ~95B | ~104B | Undisclosed | Undisclosed |
| Context Window | 1M | 1M | 1M | 1.05M |
| API Input Cost | $2/M | $3/M | $10/M | $5/M |
| API Output Cost | $6/M | $15/M | $50/M | $30/M |
| Open Weights | Promised | Released July 27 | No | No |
| OSWorld-Verified | 86.1 (claimed) | Not disclosed | 85.0 | 83.2 |
| License | TBD | Kimi K3 License | Proprietary | Proprietary |
On price alone, Qwen 3.8 is aggressively positioned. At $2/$6 per million tokens, it is one-third the cost of Claude Fable 5 and roughly one-quarter the cost of GPT-5.6 Sol Standard. For cost-sensitive agentic workloads, that price gap is meaningful. However, Fable 5 still leads on SWE-bench Pro for pure software engineering tasks, so the right choice depends on whether your priority is cost per task or absolute coding performance.
The Open-Weight Promise: What We Know
For the first time in the Max line, Alibaba has committed to releasing open weights for Qwen 3.8. Both Qwen3.8-Max and a smaller Qwen3.8-27B variant are promised “within about a week” of the August 3 launch.
What remains unknown:
- Exact release date: “Soon” and “within about a week” are not firm dates
- License terms: Whether it will be Apache 2.0, a custom license, or something more restrictive
- Serving requirements: The full 2.4T model will require multi-node GPU clusters; the 27B variant is what most teams will actually self-host
- Distillation quality: Whether the smaller 27B model retains enough of the flagship’s capability to be useful
The r/LocalLLaMA community has shown more excitement for the 27B variant than the full Max model, with several commenters focused on squeezing useful intelligence onto a single consumer GPU.
Real-World Use Cases for Qwen 3.8

Autonomous Software Engineering
Qwen 3.8 is positioned as a coding model. Alibaba showcased a 16-day autonomous coding run and a 5-day research reproduction as vendor demonstrations. Until independent verification arrives, treat these as promising signals rather than proven capabilities.
Desktop Agent Workflows
The 86.1 claimed score on OSWorld-Verified—an agentic desktop environment benchmark—suggests strong performance on tasks that require operating real software interfaces. If verified, this makes Qwen 3.8 a serious candidate for building agents that interact with operating systems, browsers, and enterprise applications.
Long-Document Analysis
With a 1-million-token context window, Qwen 3.8 can process entire codebases, legal documents, or research papers in a single pass. The 128,000-token output limit also allows for substantial generated summaries or reports.
Multimodal Enterprise Applications
The confirmed text-plus-image support opens use cases in invoice processing, visual QA, UI automation, and document understanding. For teams already in the Alibaba Cloud ecosystem, this is a natural extension.
Limitations and Concerns
Preview Uncertainty
The model launched on August 3, but the endpoint may still change. Alibaba’s own documentation notes that the preview can receive continuous upgrades and may eventually be replaced. Any evaluation you run today is testing a moving target.
Speed Issues
Early testers consistently report that Qwen3.8-Max-Preview is slow, sometimes dramatically so. If this persists as Alibaba scales capacity, it limits the model’s usefulness for anything requiring fast iteration.
No Independent Benchmarks
Alibaba’s claims are impressive, but without third-party verification from sources like Artificial Analysis or LMArena, they remain vendor marketing. The smart move is to run your own evaluations on your specific use cases before committing.
Geographic and Compliance Considerations
Qwen Cloud currently operates from the Singapore region with Global deployment. Prompts and outputs involve cross-border data transfer, which is a procurement and compliance consideration for enterprises handling sensitive data.
Who Should Use Qwen 3.8?

Qwen 3.8 is worth evaluating if:
- You need a frontier-capable model at roughly one-third the API cost of Claude or GPT
- You are building agentic systems that operate desktop environments or software interfaces
- You want a multimodal model with a 1M context window for long-document or codebase analysis
- You are waiting for open weights to self-host rather than rely on proprietary APIs
- You are already embedded in the Alibaba Cloud or Qwen ecosystem
Qwen 3.8 is not the right choice if:
- You need proven, independently verified benchmarks before production deployment
- Your application requires fast, low-latency responses (early reports indicate slowness)
- You need unrestricted production API access (Token Plan has usage restrictions)
- Your compliance requirements prohibit cross-border data transfer
- You need the absolute best software engineering performance (Fable 5 still leads SWE-bench Pro)
Frequently Asked Questions
What is Qwen 3.8?
Qwen 3.8 is Alibaba’s flagship AI model family launched on August 3, 2026. The flagship Qwen3.8-Max has 2.4 trillion parameters in a sparse MoE architecture with approximately 95 billion active parameters per token, a 1-million-token context window, and multimodal support for text, images, video, and documents.
How much does Qwen 3.8 cost?
The standard API costs $2 per million input tokens and $6 per million output tokens, with cached input at $0.25 per million. The Token Plan subscription starts at $6 per month (Lite) for individuals and $20 per seat per month (Standard) for teams.
Is Qwen 3.8 open source?
Not yet. Alibaba has promised to release open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B variant, but as of early August 2026, no weights had been published to Hugging Face or ModelScope, and no license terms had been named.
When was Qwen 3.8 released?
Alibaba previewed Qwen 3.8-Max on July 19, 2026, at the World AI Conference in Shanghai. The official launch with standard API access and published pricing occurred on August 3, 2026.
Is Qwen 3.8 better than GPT-5.6 and Claude Fable 5?
Alibaba claims Qwen 3.8 is “second only to Fable 5,” with an 86.1 score on OSWorld-Verified that would place it ahead of GPT-5.6 Sol Max (83.2) and Claude Fable 5 (85.0) on agentic desktop tasks. However, no independent benchmarks had been published as of August 2026, and Fable 5 still leads on SWE-bench Pro for software engineering.
What is the difference between Qwen 3.8 and Qwen 3.7?
Qwen 3.8 adds confirmed multimodal support, a disclosed 2.4T parameter count with ~95B active parameters, and a promised open-weight release. Qwen 3.7-Max was text-first with undisclosed architecture details and no open-weight commitment. Both share the 1M context window and agentic positioning.
Can I self-host Qwen 3.8?
Not yet. The full 2.4T model will require multi-node GPU clusters. The smaller Qwen3.8-27B variant is designed for more accessible self-hosting and is promised alongside the Max weights. Until the weights are released, only Alibaba’s hosted API is available.
What is the Qwen 3.8 context window?
Qwen 3.8-Max has a 1-million-token context window and a maximum output of 128,000 tokens (some metadata lists 131,072). This allows processing of entire codebases, long legal documents, or extensive research papers in a single pass.
Is there a free tier for Qwen 3.8?
The Token Plan Lite tier at $6 per month (promotional price) is the closest thing to an entry-level option. Alibaba Cloud Model Studio also offers a free tier with trial credits for testing. There is no permanently free API tier with generous quotas.
How do I access the Qwen 3.8 API?
You can access Qwen 3.8-Max through Alibaba’s QwenCloud platform using the model ID qwen3.8-max. The API supports both OpenAI-compatible chat completions and Anthropic-compatible interfaces, allowing you to point existing tools like Claude Code, Cursor, or Codex at it with minimal configuration changes.
Last Conclusion: Should You Build on Qwen 3.8 in 2026?
The search for Quen 3.8 features and pricing leads to one of the most aggressively positioned models of 2026. At $2 per million input tokens and $6 per million output tokens, Qwen 3.8 undercuts every major Western proprietary model by a wide margin while claiming frontier-level performance on agentic benchmarks. The 1-million-token context window, multimodal support, and promised open-weight release make it a genuine contender for teams who want capability without vendor lock-in.
But the caveats are real. The model is new. Independent benchmarks are pending. Early users report speed issues. The open-weight promise has not yet materialized. And the Token Plan’s usage restrictions mean it is not a drop-in replacement for unrestricted production API access.
My recommendation: If you are evaluating frontier models for agentic coding, desktop automation, or long-context analysis, Qwen 3.8 deserves a place in your comparison set. Start with the standard API on a small test workload. Run your own benchmarks against Claude Fable 5 and Kimi K3 on your specific tasks. Track latency, cost per task, and output quality. If the numbers hold up in your environment, the price advantage is substantial enough to justify a deeper investment.
Qwen3.8-Max is part of a rapidly evolving AI landscape where users have more powerful options than ever before. While Qwen offers an interesting alternative, ChatGPT and Claude remain two of the most widely used AI assistants for writing, research, coding, and everyday productivity. If you’re trying to decide between these platforms, you can read our detailed Claude Pro vs ChatGPT Plus comparison to see how their features, pricing, and capabilities differ. Another major AI model worth exploring is Kimi, which has gained attention in the Chinese AI ecosystem and offers a different approach to AI assistance; our Kimi AI vs ChatGPT comparison explores how the two platforms stack up in terms of features and overall usability. Looking at Qwen alongside these alternatives can help you understand which AI model or subscription is best suited to your specific needs.
Related Reads:
What Is Lovable AI? The Complete Guide to the AI App Builder Everyone’s Talking About
Is Lovable AI Free? Here’s What You Actually Get
Claude Pro vs ChatGPT Plus: The Real Difference Between These $20 AI Plans
How to Fix AI Image Generator Deformities in Fingers, Toes, and Other Anomalies
What Is Meta AI Muse Spark? The Complete 2026 Guide to Meta’s Flagship AI Model
Kimi K3 Explained: Inside Moonshot AI’s Record-Breaking Open Mode
Best AI App Builder in 2026: 10 Top Tools to Build an App Without Coding
Seedance 2: Inside ByteDance’s Most Realistic AI Video Model Yet (2026)
Is Kling 3.0 Unlimited? The Honest Truth About Free Credits, Paid Plans, and Hidden Costs in 2026

