Qwen
Qwen3.8-Flash-Next
Specs
- Input price
- $0.15/M
- Cached input price
- $0.016/M
- Output price
- $0.47/M
- Context window
- 262.144K tokens
Qwen's public first look at the Qwen4 architecture - 125B MoE at 6B active with Qwen Sparse Attention and n-gram embeddings, at preview economics.
Capabilities
- 125B MoE activating only 6B parameters per token
- Qwen Sparse Attention at the micro-block level for cheap long-context turns
- N-gram embedding pool (51B params) amenable to compute offload
- Vision-language input with reasoning_effort and preserve_thinking control
Best for
- Ahead-of-curve evaluation of where Qwen4 is heading
- Efficiency-first agentic work (DeepSWE 58.7 at a quarter of 27B's size)
- Mobile and web agents: AndroidWorld 84.5, Vision2Web 64.0
Limitations to keep in mind
- Experimental research preview - architecture and weights may change before Qwen4
- qwen-community-1.0 license: hosted third-party API access needs a separate Qwen license
- Vendor-reported benchmarks on a new harness - treat deltas as directional
HIPAA-compliant hosting
Qwen3.8-Flash-Next is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.
Pricing in context
Open source models like Qwen3.8-Flash-Next can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.
Design partner program
Run Qwen3.8-Flash-Next under HIPAA
Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.