Moonshot AI has officially unveiled Kimi K3, its flagship model boasting a staggering 2.8 trillion total parameters. Positioned as the world’s first open 3T-class model, K3 enters the arena with full weights scheduled under a Modified MIT license. Rather than optimizing strictly for quick chat responses, Moonshot engineered K3 specifically for heavy-lifting developer tasks, multi-hour engineering runs, and deep agentic workflows.
Key Technical Breakthroughs
- Sparse Mixture of Experts (MoE): Kimi K3 utilizes a Stable LatentMoE framework with 896 total experts, dynamically activating 16 per token to maximize compute efficiency.
- Kimi Delta Attention (KDA): Replaces standard attention blocks with linear attention hybrids, enabling up to 6.3x faster decoding across ultra-long contexts.
- 1-Million Token Context Window: Full support for feeding entire software codebases or massive technical documents into a single prompt.
- Native Multimodality: Processes images and video directly within the core network, allowing real-time frontend debugging via browser screenshots.
- Always-On Chain-of-Thought: Defaults to a “Max” reasoning effort designed for multi-step reasoning, tool execution, and self-correction.
Benchmark Comparison
| Feature / Model | Kimi K3 | DeepSeek v4 Pro | Claude Fable 5 |
| Total Parameters | 2.8 Trillion | 1.6 Trillion | Undisclosed |
| Active Experts | 16 / 896 | Variable | Undisclosed |
| Context Window | 1,000,000 Tokens | 128,000 Tokens | 200,000+ Tokens |
| Open Weights | Yes (July 2026) | Yes | No |
| Primary Focus | Long-Horizon Coding | General Reasoning | Multi-Modal Synthesis |
Why It Matters: Open-weight models of this scale offer enterprises complete data sovereignty and domain-specific fine-tuning capabilities without relying exclusively on closed commercial APIs.
Kimi K3 marks a clear evolution in the AI landscape: the transition from simple text generation to autonomous execution. By handling long-running terminal sessions, running local code environments, and iterating on system errors autonomously for up to 48 hours in verified benchmarks, K3 establishes open models as direct competitors to top-tier proprietary APIs.















