Source-led article

Kimi’s K3 AI Model Approaches Top-Tier Performance, Signals Shift in Chinese AI Pricing

AI News India//4 min read
A performance graph comparing Kimi K3 against other leading AI models like Claude Fable 5 and GPT 5.6 Sol, illustrating benchmark scores.
A performance graph comparing Kimi K3 against other leading AI models like Claude Fable 5 and GPT 5.6 Sol, illustrating benchmark scores.
Featured image from the source article

Kimi, a prominent AI developer, has introduced its new flagship AI model, K3, a multimodal, open-weight system boasting 2.8 trillion parameters and an extensive context window of one million tokens. This launch marks a significant moment, as K3’s performance in company benchmarks places it in close contention with leading proprietary models such as Claude Fable 5 and OpenAI’s GPT 5.6 Sol. Concurrently, the pricing structure for K3 suggests a notable shift in the landscape of Chinese AI offerings, moving away from the previously ultra-low cost models.

K3’s technical specifications include its ability to natively process images and video, positioning it as a versatile tool for complex AI applications. Kimi states that K3 is the first open model to reach approximately 3 trillion parameters. The full model weights are anticipated to be released by July 27, making it accessible for broader development and research.

Benchmark Performance and Capabilities

Internally, Kimi’s benchmarks indicate that K3, while still trailing Claude Fable 5 and GPT 5.6 Sol in some areas, consistently outperforms other major systems, including Claude Opus 4.8 and Chinese rival GLM-5.2, often by substantial margins. These results were obtained under conditions of maximum or high “thinking intensity.” Across 35 tests, K3 secured the top position in about seven instances and consistently ranked in the top three for most others.

Independent evaluations from Artificial Analysis largely corroborate Kimi’s claims. On their Intelligence Index, K3 scores 57, putting it on par with Opus 4.8 and GPT-5.5, though still behind Fable 5 and GPT-5.6 Sol. For agentic tasks, K3 achieved an Elo rating of 1,668 on GDPval v2, a significant improvement over its predecessor K2.6 and surpassing GLM-5.2, GPT-5.5, and Claude Opus 4.8. K3 also leads on AutomationBench-AA, an evaluation for agentic SaaS workflow, with a score of 53 percent. Furthermore, on AA-Briefcase, a private long-horizon knowledge work evaluation, K3 reached an Elo of 1,547, with its rubric scoring and analytical quality approaching Fable 5’s level.

Key Facts

Feature Detail
Model Name Kimi K3
Parameters 8 trillion (multimodal, open-weight)
Context Window 1 million tokens
Benchmark Performance Nears Claude Fable 5, GPT 5.6 Sol; beats Opus 4.8, GLM 5.2
Primary Use Case Long-running software development, knowledge work, complex reasoning

Advanced Architecture and Use Cases

K3 employs a mixture-of-experts (MoE) architecture, activating only 16 of its 896 experts at any given time. This is paired with a new attention architecture called Kimi Delta Attention, which Kimi claims enables up to 6.3 times faster decoding for million-token contexts. “Attention residuals” are also reported to boost training efficiency by approximately 25 percent with minimal computational overhead.

Kimi highlights K3’s primary use case as long-running software development with minimal human oversight. The model is designed to analyze large codebases, coordinate terminal tools, and maintain focus across multiple work steps. A key feature is its “Vision in the Loop” system, where K3 examines screen captures, modifies code, and then checks the visible output. This closed-loop system is envisioned as a foundation for game development, UI design, and CAD. Demos have included a procedurally generated 3D open-world game built entirely in the browser and an interactive black hole visualization.

Pricing Shift and Market Impact

The pricing for K3 indicates a significant departure from previous low-cost Chinese AI models. According to the Kimi API documentation, one million input tokens cost $0.30 with a cache hit and $3.00 without, while one million output tokens, including reasoning, cost $15.00. These prices apply regardless of context length. While considerably higher than its predecessor K2.6 (which cost $0.16 per million tokens with cache hit and $4.00 for output), K3 remains more affordable than many top Western models. For example, Anthropic’s Sonnet 5 costs $3 per million input tokens and $15 for output but delivers lower performance.

Artificial Analysis estimates K3’s average cost per task on the Intelligence Index at $0.94, comparable to GPT-5.6 Sol ($1.04) and about half the price of Opus 4.8 ($1.80). However, it is still more expensive than open-weight peers like GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04). Despite higher token prices, K3’s efficiency – needing fewer tokens per task – might balance overall costs. This pricing strategy suggests that Chinese providers are increasingly valuing their frontier models, moving towards a more competitive, mid-range pricing structure rather than undercutting the market.

Availability and Future Plans

K3 is currently accessible via Kimi.com, its mobile app (iOS, Android, HarmonyOS), and the Kimi Work desktop client (version 3.1.0 and later). Developers can also find it on OpenRouter under the identifier “moonshotai/kimi-k3.” The full open weights are expected by the end of July. For businesses, Kimi offers a separate version with member management and account splitting. A planned platform, Kimi Hosted Agent, will provide isolated environments for long-running tasks, with a waitlist available for interested users.

This development is particularly relevant for the Indian AI and tech community, as it introduces a powerful new open-weight model that could drive innovation in AI development and application. The shift in pricing also has implications for Indian businesses and startups considering various AI solutions, offering a new competitive option in the global AI market.

Source: The Decoder, “Kimi’s open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI” https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/