Moonshot AI Kimi K2 Explained: The Open-Source Model Series That Changed AI in 2026

Moonshot AI’s Kimi K2 series redefined open-source AI in 2026. Explore K2.5, K2.6 specs, Kimi Work, Kimi Code, Agent Swarm, pricing, and what replaced it.

If you have been following AI model releases in 2026, you have probably seen the name moonshot ai kimi k2 appear in benchmark leaderboards, developer forums, and funding headlines. But here is the reality most guides gloss over: the Kimi K2 series was officially discontinued in May 2026. So why are people still talking about it? Because the K2 line—spanning K2, K2.5, and K2.6—was the foundation that propelled Moonshot AI from a promising Chinese startup to a $35 billion juggernaut rivaling OpenAI and Anthropic.

In this guide, we will trace the full K2 lineage from its mid-2025 launch through its final K2.6 release, explain how it powered Moonshot’s agent ecosystem including Kimi Work and Kimi Code, and clarify what replaced it when the series was sunset. Whether you are a developer evaluating open-weight models, a founder choosing AI infrastructure, or simply curious about one of the most consequential model families of 2026, this article gives you the complete picture.


What Is Moonshot AI?

moonshot ai

Moonshot AI is a Beijing-based artificial intelligence lab founded in 2023 by Yang Zhilin, a former Meta AI and Google Brain researcher. The company focuses on large language models, AI agents, and enterprise AI infrastructure, with its Kimi chatbot platform becoming one of China’s highest-profile consumer AI products.

The company’s growth has been staggering. In April 2026, Moonshot’s annual recurring revenue surpassed $200 million, up from roughly $100 million just six weeks earlier. By June 2026, ARR hit $300 million.

In July 2026, Moonshot closed a $3.5 billion funding round at a $35 billion valuation—far exceeding its $1 to $2 billion target—with the National Artificial Intelligence Industry Investment Fund of China among the lead investors. The company is now reportedly preparing for a Hong Kong IPO.


The Kimi K2 Series: From K2 to K2.6

kimi k2.6

The moonshot ai kimi k2 family was a series of open-weight, Mixture-of-Experts (MoE) models that established Moonshot as a serious contender in the global AI race. Here is how the lineage evolved:

Kimi K2 (Mid-2025)

The original Kimi K2 launched around mid-2025 with a 1-trillion-parameter MoE architecture, activating approximately 32 billion parameters per token. It featured a 256,000-token context window, native multimodal support through the MoonViT vision encoder, and was released under open licenses that allowed commercial use.

Kimi K2.5 (January 2026)

Kimi K2.5 launched on January 27, 2026, building on the same 1T MoE architecture but trained on significantly more data—approximately 32 trillion tokens across the full pipeline, including mixed visual and text tokens. It introduced the Agent Swarm research preview, supporting up to 100 parallel sub-agents.

Kimi K2.6 (April 2026)

Kimi K2.6 arrived on April 20, 2026, and represented the pinnacle of the K2 line. It maintained the 1T MoE / 32B active parameter architecture and 256K context window but delivered measurable improvements in long-horizon coding, agentic execution, and swarm orchestration.

Key K2.6 advancements included:

  • Agent Swarm scaling from 100 to 300 sub-agents and from 1,500 to 4,000 coordinated steps
  • Long-horizon coding capabilities supporting 12+ hour autonomous runs with thousands of tool calls
  • Coding-driven design that transforms prompts into production-ready front-end interfaces
  • A Modified MIT license for full commercial use

The End of the K2 Line (May 2026)

On May 25, 2026, Moonshot officially discontinued the entire kimi-k2 series. The API platform documentation states: “The kimi-k2 series models were officially discontinued on May 25, 2026 and are no longer maintained or supported.”

This discontinuation was not a failure—it was a transition. Moonshot had already begun shifting focus to its next-generation architecture, culminating in the Kimi K3 launch in July 2026.


Kimi K2.6 Architecture and Technical Specifications

Understanding why the moonshot ai kimi k2 series mattered requires looking under the hood.

SpecificationKimi K2.6
Total Parameters1 trillion
Active Parameters32 billion per token
ArchitectureMixture-of-Experts (MoE)
Experts384 total, 8 selected per token + 1 shared
Layers61
Context Window262,144 tokens (256K)
AttentionMulti-Head Latent Attention (MLA)
Vision EncoderMoonViT (400M parameters)
ActivationSwiGLU
Vocabulary160K
LicenseModified MIT

The MoE design was central to K2.6’s efficiency. By activating only 32B of its 1T parameters per token, the model delivered frontier-scale capability while keeping inference costs closer to a dense 32B model.

The 400M-parameter MoonViT vision encoder handled native image and video input—not through adapters, but as a core part of the model architecture. For video, consecutive frames were grouped in fours and temporally pooled, achieving 4x compression.


Benchmark Performance: How Kimi K2.6 Stacked Up

kimi k2.6 Benchmark Performance

K2.6’s benchmark story was nuanced: it dominated in agentic and coding tasks while trailing slightly on pure reasoning without tools.

Where K2.6 Led

  • SWE-Bench Pro: 58.6%—ahead of GPT-5.4 (57.7%), Claude Opus 4.6 (53.4%), and Gemini 3.1 Pro (54.2%)
  • DeepSearchQA: 83.0% accuracy and 92.5 F1—substantially ahead of GPT-5.4 (63.7% / 78.6 F1)
  • HLE-Full with tools: 54.0%—edging out Claude (53.0%) and GPT-5.4 (52.1%)
  • LiveCodeBench v6: 89.6%—competitive with Gemini (91.7%) and ahead of Claude (88.8%)

Where K2.6 Trailed

  • AIME 2026 (pure math): 96.4% vs GPT-5.4’s 99.2%
  • GPQA-Diamond: 90.5% vs Gemini’s 94.3%
  • HLE-Full without tools: 34.7% vs Claude’s 40.0% and Gemini’s 44.4%

Featured snippet answer: Is Kimi K2.6 better than GPT-5.4 and Claude? It depends on the task. K2.6 led on software engineering (SWE-Bench Pro), deep search, and agentic workflows with tools. GPT-5.4 and Gemini led on pure math reasoning and knowledge retrieval without external tools.


The Agent Ecosystem Built on Kimi K2

What made the moonshot ai kimi k2 series truly significant was not just the model weights—it was the agent ecosystem Moonshot built around it.

Kimi Work: The Desktop Knowledge Agent

Kimi Work is Moonshot’s desktop agent for knowledge work, launched in June 2026. It is a local application for macOS (Apple Silicon only) and Windows that mounts folders, drives real browsers through WebBridge, executes scheduled Cron jobs, and coordinates Agent Swarm runs into PowerPoint and Excel outputs.

Unlike web chat, which answers “What does this mean?” Kimi Work answers “Do this every night and put the result in my folder.” It includes native finance data for A-shares, Hong Kong, and US equities, and can output finished documents without the copy-paste loop common to other AI tools.

The app is free to download, but agent tasks consume membership credits. On the free Adagio tier, users get approximately six agent task equivalents per month.

Kimi Code: AI Programming Assistant

kimi code

Kimi Code is Moonshot’s coding assistant, available as a VS Code extension and command-line interface. Built for long-context workflows, it autonomously explores codebases, reads and writes code, runs terminal commands with permission, and supports MCP servers for extended capabilities.

Key features include thinking controls, native VS Code diff viewer integration, slash commands like /init and /compact, and the ability to work with the full 256K context window. The VS Code extension requires VS Code 1.100.0 or later.

Kimi Web Bridge: Local Browser Automation

limi web bridge

Kimi Web Bridge is a Chrome/Edge extension that lets AI agents control your real browser using the Chrome DevTools Protocol (CDP). Unlike cloud-based automation, WebBridge runs locally—your login sessions, cookies, and page content never leave your machine.

This enables agents to navigate authenticated sites, fill forms, extract tables, and perform multi-site research using your existing browser sessions.

Kimi Claw: 24/7 Cloud Agent

Kimi Claw is Moonshot’s browser-native, always-on AI agent built on the OpenClaw framework. Launched in beta in February 2026, it runs 24/7 in a browser tab with persistent memory, proactive scheduling, and access to 5,000+ community skills through ClawHub. It includes 40GB of cloud storage and supports Telegram integration.

Unlike Kimi Work, which stops when your laptop sleeps, Kimi Claw runs continuously in the cloud. It is available on Allegretto tier and above.

Agent Swarm: Parallel Multi-Agent Orchestration

Agent Swarm dynamically decomposes tasks into heterogeneous subtasks executed concurrently by domain-specialized agents. K2.6 scaled this to 300 sub-agents executing 4,000 coordinated steps simultaneously—up from K2.5’s 100 sub-agents and 1,500 steps.

The swarm can deliver end-to-end outputs spanning documents, websites, slides, and spreadsheets in a single autonomous run. It also converts high-quality files into reusable Skills that preserve structural and stylistic DNA.


Pricing and Subscription Tiers

kimi pricing

Moonshot names its subscription tiers after musical tempos. As of 2026, the lineup is:

PlanMonthly PriceKey Features
AdagioFreeBasic chat, ~6 agent credits, 1 concurrent task
Moderato$1960 agent credits, Kimi Code 1x, 25 swarm runs
Allegretto$39150 agent credits, Kimi Code 5x, 50 swarm runs, Kimi Claw
Allegro$99360 agent credits, Kimi Code 15x, 120 swarm runs
Vivace$199720 agent credits, Kimi Code 30x, 240 swarm runs

All paid plans include Agent Swarm access, with higher tiers increasing concurrency and credit multipliers. Annual billing saves roughly 15–25%.

Important: Subscription plans and API access are completely separate systems. Paying for Moderato does not discount API tokens, and API top-ups do not unlock subscription features.

API pricing for K2.6 was approximately $0.95 per million input tokens and $4.00 per million output tokens, with cached input dropping to roughly $0.10–$0.16.


What Replaced Kimi K2? Introducing Kimi K3

kimi k3

On July 16, 2026, Moonshot launched Kimi K3—the successor that made the K2 series obsolete. K3 is a 2.8-trillion-parameter MoE model with a 1-million-token context window, native vision, and new architectural innovations including Kimi Delta Attention (KDA) and Attention Residuals (AttnRes).

K3 vs K2.6: Key Differences

FeatureKimi K2.6Kimi K3
Total Parameters1T2.8T
Active Parameters32B~50B equivalent (16/896 experts)
Context Window256K1M tokens
ArchitectureMLA + MoEKDA + AttnRes + Stable LatentMoE
LicenseModified MITKimi K3 License
API Input Cost~$0.95/M$3/M ($0.30 cached)
API Output Cost~$4/M$15/M

K3 ranks #5 of 215 models on the BenchLM Intelligence Index with a score of 79.8, placing in the 98th percentile for agentic tasks and 97th for coding.

On the Arena WebDev benchmark, K3 ranked #1 in frontend coding, surpassing Claude Fable 5.

The full K3 weights were released on July 27, 2026, making it the largest open-weight model ever released. However, at 1.56 TB in MXFP4 precision, self-hosting requires multi-node GPU clusters with 64+ accelerators—not a single workstation.


Real-World Use Cases for the K2 Architecture

Even with K2 deprecated, understanding its capabilities helps evaluate the K3 ecosystem that inherited them:

Autonomous coding agents: K2.6 sustained 12+ hour runs, overhauling legacy codebases like the 8-year-old exchange-core financial matching engine, improving throughput by 185% across 1,000+ tool calls.

Deep research: K2.6’s DeepSearchQA performance made it ideal for multi-source research synthesis, outperforming GPT-5.4 by 14 F1 points.

Knowledge work automation: Through Kimi Work, K2.6-powered swarms generated PowerPoint decks, Excel spreadsheets, and briefing documents from natural language prompts.

Multi-modal workflows: Native vision support enabled chart interpretation, visual reasoning, and video understanding without separate vision models.


Who Should Still Care About Kimi K2?

If K2 is discontinued, why does it still matter? Three reasons:

1. Understanding the ecosystem: Kimi Work, Kimi Code, and WebBridge were all architected around K2’s capabilities. Knowing what K2 did explains what K3 improves.

2. Self-hosted deployments: Teams already running K2.6 on their own infrastructure may not need to rush to K3. The Modified MIT license means K2.6 remains fully usable.

3. Cost efficiency: K2.6 API pricing undercuts K3 significantly. For workloads where K2.6’s 256K context and capabilities suffice, it may still be the more economical choice through third-party providers.

Internal linking opportunity: Link to a dedicated guide on migrating from K2.6 to K3 for existing Moonshot users.


Moonshot AI Kimi K2 vs Competitors

DimensionKimi K2.6Claude Opus 4.6GPT-5.4
Open weightsYes (Modified MIT)NoNo
SWE-Bench Pro58.6%53.4%57.7%
Context window256K1M1M
Agentic tasksLeads with toolsStrong reasoningStrong math
API cost (output)~$4/M~$25/M~$15–30/M
Self-hostableYes (8x H100)NoNo

K2.6’s core advantage was not universal dominance—it was specialization. For coding agents, deep search, and tool-augmented workflows, it punched above its weight. For pure reasoning without tools, closed models maintained an edge.

Kimi AI isn’t the only powerful AI assistant available today. For a deeper look at OpenAI’s ecosystem, read our comprehensive ChatGPT guide, where we compare models, pricing plans, and premium features to help you choose the right AI tool.


Frequently Asked Questions

What happened to Kimi K2?

The entire Kimi K2 series (K2, K2.5, K2.6) was officially discontinued on May 25, 2026. Moonshift shifted focus to the K3 architecture, which launched in July 2026. Existing K2.6 weights remain available under the Modified MIT license for self-hosting.

Is Kimi K2.6 still usable?

Yes. While discontinued from official API support, K2.6 weights remain on Hugging Face and can be self-hosted. Third-party inference providers may still offer K2.6 API access. The Modified MIT license permits commercial use indefinitely.

What is the difference between Kimi K2.5 and K2.6?

K2.6 improved on K2.5 with better long-horizon coding (+15% on internal benchmarks), enhanced instruction following, more thorough reasoning, and Agent Swarm scaling from 100 to 300 sub-agents. Both share the same 1T MoE architecture and 256K context.

What is Kimi Work?

Kimi Work is Moonshot’s desktop agent for macOS (Apple Silicon) and Windows. It mounts local folders, automates browsers via WebBridge, runs scheduled Cron jobs, and coordinates Agent Swarm outputs into documents and spreadsheets.

What is Kimi Code?

Kimi Code is Moonshot’s AI coding assistant, available as a VS Code extension and CLI. It explores codebases autonomously, writes and edits code, runs terminal commands, and supports MCP servers for extended tool access.

How much does Kimi K2.6 cost?

K2.6 API pricing was approximately $0.95 per million input tokens and $4 per million output tokens. Subscription tiers for the consumer app range from free (Adagio) to $199/month (Vivace).

Is Kimi K2.6 open source?

K2.6 was released under a Modified MIT license, allowing commercial use, modification, and self-hosting. The weights are available on Hugging Face. Note that the K3 successor uses a different Kimi K3 License with attribution requirements for very large-scale use.

Can I run Kimi K2.6 locally?

Yes, but it requires serious hardware. K2.6 needs a minimum of 8× NVIDIA H100-80GB GPUs. The model weights are approximately 595 GB in native INT4 precision.

What replaced Kimi K2?

Kimi K3, launched July 16, 2026, is the direct successor. It features 2.8T parameters, a 1M context window, KDA + AttnRes architecture, and open weights released July 27, 2026.

Is Moonshot AI owned by Google?

No. Moonshot AI is an independent, privately held company. It is a Google Cloud technology partner and received investment from Google Ventures, but it is not a Google subsidiary. (Note: This citation refers to FlutterFlow’s structure; Moonshot is similarly independent.)


My Final Conclusion: The Legacy of Moonshot AI Kimi K2

The moonshot ai kimi k2 series was more than a line of models—it was a proof of concept that open-weight AI could compete with closed frontier systems on the tasks that matter most to developers. K2.6’s dominance on SWE-Bench Pro, its 12-hour autonomous coding runs, and its 300-agent swarm orchestration showed what was possible when architectural efficiency met open licensing.

Its discontinuation in May 2026 was not an ending but a graduation. The K3 architecture that replaced it carries forward every K2 innovation while scaling parameters, context, and efficiency to a new order of magnitude.

My recommendation: If you are starting fresh with Moonshot AI today, focus on K3. It is the current flagship, actively supported, and available across all Kimi surfaces. If you are already running K2.6 in production, evaluate whether K3’s 1M context and architectural improvements justify the migration cost. For many coding and agentic workloads, K2.6 remains perfectly capable—and at a fraction of K3’s API pricing.

Ready to explore the Kimi ecosystem? Start with the free Adagio tier at kimi.com. Test a coding task with Kimi Code, run a research sweep with Kimi Work, or simply chat with K3 to see how far Moonshot has come since the K2 days. The best way to understand what these models can do is to let them do your work.

Related Reads:

What Is Lovable AI? The Complete Guide to the AI App Builder Everyone’s Talking About

Is Lovable AI Free? Here’s What You Actually Get

Claude Pro vs ChatGPT Plus: The Real Difference Between These $20 AI Plans

How to Fix AI Image Generator Deformities in Fingers, Toes, and Other Anomalies

What Is Meta AI Muse Spark? The Complete 2026 Guide to Meta’s Flagship AI Model

Kimi K3 Explained: Inside Moonshot AI’s Record-Breaking Open Mode

Best AI App Builder in 2026: 10 Top Tools to Build an App Without Coding

Bolt AI App Builder: The Complete Guide to Building Apps in Your Web Browser

Higgsfield AI Pricing in 2026: Every Plan, Credit Cost, and What You Actually Get

Base 44 Pricing in 2026: Every Plan, Credit Cost, and Hidden Fee Explained

Base44 Reviews 2026: What Real Users Actually Think Before You Build

FlutterFlow AI App Builder: The Complete 2026 Guide to Building Real Native Apps

Kimi AI App: The Complete Guide to Download, Features, and How It WorksIs

Kimi AI Free Enough in 2026? A Kimi AI Free Plan (2026): Features, Daily Limits & Hidden Restrictions

Kimi AI vs ChatGPT: Which AI Assistant Actually Wins in 2026?