Original title: đ° Tech Blog | đ Full Report
Article
Moonshot AI presented Kimi K3 as an open-weights 2.8T-parameter multimodal model with a 1-million-token context window, built with Kimi Delta Attention, Attention Residuals, and a 896-expert stable latent MoE design. The release details native text-image-video capabilities, API-compatible deployment through Kimiâs platform, and support across Transformers, vLLM, SGLang, Docker, and local runtime workflows for coding and tool-augmented agents. The report emphasizes that Kimi K3 expects full preserved reasoning content to be retained across turns, exposes a reasoning effort setting, and provides usage guidance for vision and structured outputs. It also includes architecture and benchmark tables spanning reasoning, coding, agentic, and vision suites with per-benchmark notes on data windows, harnesses, date stamps, and caveats such as tool augmentation and fallback behavior. The model weights and code are distributed under a Kimi K3 License that imposes extra commercial terms for very large operators, and the documentation states quantization-aware training with MXFP4 weights and MXFP8 activations. The release highlights both flagship performance claims and practical release mechanics, and links to checkpoints and leaderboards. Observers also note reported storage size near 1.6TB and mixed endpoint availability, which shifts focus from capability claims to deployment feasibility.
The comments welcome the release as a major open-model milestone but quickly pivot to economics and feasibility. Readers estimate hosting is likely costly because of large VRAM and throughput needs, and they monitor provider pricing for input, cached input, and output tokens to infer whether AI APIs are sustainably above marginal cost. Many report 404 errors and link instability around Hugging Face and partner endpoints, treating reliability as a real concern during launch. Several participants compare current prices to GLM 5.2 and suggest competition is still driving aggressive discounting, while others question whether providers can eventually undercut full infrastructure costs. A strong thread asks whether the model can be meaningfully distilled or compressed for local or smaller deployments, especially via MoE-aware methods, and whether tools like unsloth or GB300-era quantization could help. Security and governance concerns appear through questions about censorship controls, political bias, and rumor-driven speculation about rapid commercialization or future policy pressure. Overall sentiment balances excitement over frontier capability with skepticism that most users will ever run it locally.