Moonshot
Kimi K3
Specs
- Input price
- $3/M
- Cached input price
- $0.30/M
- Output price
- $15/M
- Context window
- 1M tokens
Moonshot AI's 2.8-trillion-parameter flagship - the #1 open-weight model on the Artificial Analysis Intelligence Index (score 57), with a 1M-token context window, native vision, and always-on reasoning.
Capabilities
- 2.8T parameter sparse MoE with 896 experts (16 active per token)
- Kimi Delta Attention with Attention Residuals for efficient 1M-token context
- Native vision - reads images, scanned records, forms, and charts
- Always-on reasoning tuned via reasoning_effort (Low, High, or Max)
Best for
- Long-horizon autonomous agentic workflows
- Repository-scale software engineering and deep debugging
- Document-heavy clinical workflows with native image input
Limitations to keep in mind
- Modest throughput (~34 tokens/second) - pair with a fast model for interactive tiers
- Verbose generation; thinks out loud and consumes more output tokens per task
- Text and image only - no native video input
HIPAA-compliant hosting
Kimi K3 is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.
Pricing in context
Open source models like Kimi K3 can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.
Design partner program
Run Kimi K3 under HIPAA
Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.