On the evening of July 27, Moonshot AI released the complete model weights for Kimi K3 on Hugging Face, simultaneously publishing the technical report and three training infrastructure technologies. This marks the world's first open-source model to step into the 3-trillion parameter tier, pushing the open-source large model arms race into a new phase.
We immediately conducted a systematic empirical test of Kimi K3. Combining this with the official technical report, we provide a comprehensive interpretation from three dimensions: architecture design, capability performance, and engineering deployment.
1. Model overview
1.1 Basic parameters
| Project | Parameter |
|---|---|
| Total parameters | 2.8 trillion |
| Active parameters | 104 billion (per token) |
| Architecture | Stable LatentMoE (Mixture of Experts) |
| Expert module | 896 routing experts, 16 activated per token |
| Context window | 1 million tokens |
| Vision capabilities | Native visual understanding (MoonViT-V2 encoder) |
| Open source license | Modified MIT License |
1.2 Release background
Kimi K3 was first unveiled on July 16, just before WAIC 2026, with the weights officially opened on July 27. The parameter scale is roughly three times that of the previous generation, Kimi K2.5.
K3 doesn't just blindly stack parameters. Under the premise of limited compute resources, through a combination of technologies like Kimi Delta Attention, Attention Residuals, and MoonEP, the overall compute scaling efficiency improved by approximately 2.5 times compared to K2.
The expansion path from K2 to K3 is very clear: the main hidden dimension remains unchanged at 7168, layers increased from 61 to 93, routing experts expanded from 384 to 896, and the training context window was stretched from 128K to 1M. The newly added parameters were primarily allocated to more layers and a larger expert pool.
2. Architecture analysis
K3's architecture is not a standard Transformer pipeline, but rather three orthogonal mixed paths: the sequence direction is mixed by KDA + MLA, the depth direction uses AttnRes to select representations from previous layers, and the width direction uses Stable LatentMoE to allocate tokens among experts.
2.1 Stable LatentMoE: sparse activation of 16 out of 896
The core idea behind the MoE (Mixture of Experts) architecture is that the model contains many "experts," but each token only calls upon a small fraction of them, trading sparse activation for a balance between massive capacity and low computational overhead.
Kimi K3 pushes this concept to the extreme: 896 routing experts, with only 16 activated per token. The sparsity (activation ratio) is about 1.8%, meaning that over 98% of the model's parameters remain dormant during a single forward pass.
This brings two direct benefits:
- Controllable inference costs. A total of 2.8 trillion parameters sounds intimidating, but the actual parameters participating in the computation per token is only 104 billion. The inference cost is far lower than that of a dense model with equivalent capabilities.
- Specialized division of labor. A massive number of experts allows the model to cultivate dedicated "sub-networks" for different domains (code, math, Chinese, vision) without them interfering with one another.
An old problem with MoE is routing instability. During training, expert load becomes uneven, easily leading to a collapse where a few experts are overworked while the majority sit idle. K3's Stable LatentMoE, combined with the MoonEP communication library, solves exactly this issue, achieving what official reports call "perfect load balancing."
Routing experts do not directly process the full 7168-dimensional hidden state. Instead, they project it down to a 3584-dimensional latent space. The 16 selected experts compute in this half-width space, aggregate the results, pass them through RMSNorm, and project them back to 7168 dimensions. Two shared experts maintain a full-width path. Quantile Balancing handles equilibrium at the routing layer, while MoonEP dynamically replicates "hot" experts at the distributed execution layer. Together, these two layers resolve the load balancing challenges at the 896-expert scale.
2.2 Kimi Delta Attention: the engineering foundation for 1M context
A 1-million token context is not just a marketing number; it is backed by a fundamental rebuild at the attention mechanism level.
Standard Transformer attention has O(n²) complexity, making the compute and VRAM overhead for a million-level context unacceptable. Kimi K3 adopts a proprietary Kimi Delta Attention (KDA) mixed linear attention mechanism, paired with Attention Residuals technology, to compress the computational cost of long contexts down to an engineerable level.
K3 is not a pure linear attention model. It repeats a basic block of 3 layers of KDA and 1 layer of Gated MLA at a 3:1 ratio, and places a final layer of MLA to ensure a global attention pass at the end of the backbone. KDA uses a fixed-size recurrent state to compress history, while MLA periodically restores global retrieval capabilities.
The accompanying high-performance FlashKDA operators have also been open-sourced. According to official data, on an NVIDIA H200, prefill speeds improved by 1.72x to 2.22x compared to the baseline.
A crucial detail that makes KDA engineerable: K3 restricts the step-wise log-decay to (-5, 0), ensuring the corresponding retention factor is always greater than e^-5. The cumulative log-decay for a 16-token tile falls within (-80, 0), keeping the inverse scaling entirely within the BF16 dynamic range. As a result, both diagonal and non-diagonal causal tiles can be computed using dense Tensor Core matrix multiplication, eliminating the need for special diagonal paths. This modification doesn't alter KDA's definition as a recurrent state, but it directly dictates which hardware paths the kernel can utilize.
For developers, the practical significance of a million-token context means analyzing entire books, understanding massive code repositories, and maintaining agent memory across long, multi-turn conversations, all without the need for complex RAG (Retrieval-Augmented Generation) chunking engineering.
Accompanying KDA is Attention Residuals (AttnRes). Standard Transformer residual connections use fixed-weight additions; layer 90 only sees the accumulated result of all previous layers mashed into the same hidden state. AttnRes pivots the attention concept from the token axis to the layer axis, allowing each layer to selectively read the representations of previous layers. K3 uses Block AttnRes to divide its 93 layers into 8 blocks, dropping cross-block state memory and communication from O(Ld) down to O(Nd), making it practically viable.
Inference-side cache design was also altered. KDA saves a fixed-size recurrent state, while MLA saves a per-token KV Cache, two states with different lifecycles. K3 unifies their management in a single paged pool. A prefix can only be reused when both the MLA KV and the KDA state at the corresponding position result in a hit. Whether a million-token context actually works in production ultimately depends on whether states can be saved, transferred, hit, and scheduled, not just changing a context_length variable in a config file.
2.3 MoonViT-V2: native vision, not a bolt-on
Many multimodal models achieve visual capabilities by bolting on an off-the-shelf vision encoder (like SigLIP) to a language model and aligning them via contrastive learning. This training pipeline is cumbersome, and there is always a barrier between the vision and language components.
Kimi K3's vision encoder, MoonViT-V2, was built from scratch and trained directly using next-token prediction. It did not use models like SigLIP for initialization, and it shares the same objective function as the language model, entirely bypassing the contrastive pre-training phase. Comparative experiments in the technical report show that joint training initialized with pre-trained vision encoders causes higher and more frequent gradient spikes, whereas MoonViT-V2, trained from scratch, converges much more stably and ultimately matches the baseline performance of SigLIP initialization.
The advantage of native vision is that visual information enters the model at a much deeper level. It isn't translated into text before being understood; it is fused directly at the representation level.
3. Empirical capability testing
Scoring well on a technical report is one thing; actual API calls are another. We ran an empirical test on Kimi K3 with 21 evaluation questions (sourced from public test sets like LatePost Eval and QbitAI). Here are a few representative results.
3.1 Complex logic: the most complete reasoning chain
We used the "Prisoner Hat" problem to test multi-layered nested reasoning ("I know that you know that I don't know"):
A can see B and C, B can see C, and C can see no one. There are 3 red hats and 2 blue hats. A says he doesn't know, B says he doesn't know. Can C determine the color of his own hat?
Kimi K3 advanced its answer in two layers: first, it analyzed A's statement to rule out the "double blue" possibility for B and C. Then, it analyzed B's statement, knowing that B also heard A's information but still didn't know, to deduce that C must be wearing a red hat. The logic chain was flawless, with no skipped steps.
On the "Planet version of the farmer crossing the river" (an anti-memorization variant), K3 used a table to list the operations and the state of both riverbanks at each step. It also proactively added a boundary condition analysis: "If an additional rule states that Earth must take a planet across every time and the ship can never be empty, then this problem is unsolvable." Among the 7 models tested, K3 was the only one to actively consider boundary conditions in this manner.
3.2 Fact-checking: perfect score
On three fact-checking questions (correcting false premises, identifying fabricated papers, and cross-validating contradictory sources), Kimi K3 answered all correctly. When asked to summarize a non-existent Nature paper, it directly pointed out that the paper did not exist, rather than hallucinating content to appease the prompt.
3.3 API call notes
Two crucial details observed during testing:
- temperature must be set to 1. K3 has a strict constraint on sampling temperature; setting it to any other value will directly throw an error. This differs from the vast majority of models and is an easy pitfall when migrating code.
- Ensure sufficient max_tokens. As a deep reasoning model, K3 consumes a significant number of tokens during the reasoning phase of complex tasks. We recommend setting
max_tokensto at least 4096 for complex tasks, otherwise you may encounter situations where the reasoning is finished but the final answer gets cut off.
Real-world latency data: simple Q&A returns in a few seconds; complex logic questions take 35-134 seconds. Deep reasoning trades speed for inference quality, making it unsuitable for latency-sensitive, real-time scenarios.
4. Engineering capabilities: more than just a model
The most substantial part of this open-source release is actually the two sets of engineering experiments disclosed in the technical report.
4.1 Autonomous GPU kernel optimization
The technical report reveals that Kimi K3 independently completed the optimization of four representative GPU kernels within 24 hours: AttnRes latency dropped from 283.6ms to 114.4ms, DSA and KDA runtime decreased by 55.1% and 73.6% respectively, and MLA approached peak TFLOPS. This performance matches Claude Fable 5 and surpasses GPT-5.6 Sol.
K3 also autonomously developed a compact GPU compiler named MiniTriton, completing the full pipeline from Python frontend to PTX code generation entirely on its own.
An even more critical signal: the testing environment covered both NVIDIA H200 and GPUs from other vendors. This means K3's deployment capabilities are not locked into the NVIDIA ecosystem; adapting it to domestic (Chinese) compute cards is a realistic option.
4.2 Designing a chip in 48 hours
The technical report also detailed a proof-of-concept: over a 48-hour autonomous run, using open-source EDA tools and the Nangate 45nm process library, K3 completed the design, optimization, and verification of an inference chip prototype for a micro-model utilizing mixed KDA, MLA attention, and MoE routing. The chip area is 4 mm², integrating 1.46 million standard cells and 0.277MB of SRAM, achieving 100MHz timing closure and a decoding throughput of over 8700 tokens/second.
A chip designed by a model, to serve a model. From algorithms to hardware, this full-stack autonomous engineering capability is a signal far more worthy of attention than any benchmark score.
5. Quick start
Kimi K3 is now live on TokenGo and is compatible with the OpenAI interface, call it with our API now:
from openai import OpenAI
client = OpenAI(
api_key="your_api_key",
base_url="https://api.tokengo.com/v1"
)
response = client.chat.completions.create(
model="moonshotai/kimi-k3",
messages=[{"role": "user", "content": "Hello!"}],
temperature=1.0 # Note: K3 must be exactly 1.0
)
print(response.choices[0].message.content)With our serverless inference, there's no need to deploy the 2.8 trillion parameter weights yourself, nor do you need to manage low-level optimizations like MoonEP or FlashKDA. Our plug and play platform lets you call it as is in seconds.
6. Conclusion
The core takeaways from the open-source release of Kimi K3 can be summarized into three points:
- Architecture: The 16/896 Stable LatentMoE pushes sparse activation to new heights. Out of 2.8 trillion total parameters, only 104 billion are active, achieving both massive capacity and controllable inference costs. KDA linear attention acts as the backbone for a true one-million-token context.
- Capabilities: Based on our empirical testing, K3's strengths lie in the completeness of its complex logic reasoning and the reliability of its fact-checking. Its reasoning chains are meticulous, and it proactively considers boundary conditions. The trade-off is a slower response time compared to non-reasoning models, making it ideal for deep tasks rather than high-frequency rapid responses.
- Ecosystem: The triad of model weights, technical report, and training infrastructure have all been opened. The experiments in GPU kernel optimization and chip design demonstrate engineering potential that extends far beyond the model itself. Its deployment capabilities are not bound to a single chip ecosystem, a point of particular importance to domestic developers.
Open-source large models have officially entered the 3-trillion parameter era. The next thing to watch: who can deliver these capabilities into the hands of developers with the lowest possible barrier to entry.