Alibaba Launches Qwen3.7-Flash, a Lightweight, Low-Cost Vision-Language Model with Stronger Agent Skills
Alibaba has released βQwen3.7-Flash,β positioned as the lightweight, low-cost tier of the Qwen3.7 series. The series already includes Qwen3.7-Max (May 2026) and Qwen3.7-Plus (June 2026), with Flash arriving last as the efficiency-focused option.
Details
- Native vision-language model: Accepts interleaved text and image input and returns text, as a native multimodal design representing a comprehensive upgrade over the prior Qwen3.6-Flash
- Stronger multimodal foundations: Improved universal object recognition, along with further gains in real-world perception and spatial intelligence
- Upgraded agent capabilities: Significantly enhanced multimodal agent performance for scenarios like Search Agent and CI Agent, with more stable end-to-end task execution
- Optimized multimodal coding: Improved accuracy on multimodal coding tasks
- Context length: Roughly a 1-million-token context window with up to 65,536 tokens of output, enabling long multi-image sequences, long documents, and extended agent trajectories in a single request
- Pricing: Set low at $0.03 per 1M input tokens and $0.13 per 1M output tokens
- Series positioning: The third model in the 3.7 generation, following Qwen3.7-Max (May 2026) and Qwen3.7-Plus (June 2026); its internal snapshot name is βqwen3.7-flash-2026-07-15β
How to try it
- Available via API through Alibaba Cloudβs Model Studio (DashScope), and also offered through third parties such as OpenRouter
- No announcement yet on availability in a consumer-facing chat UI or on any open-weight release; further updates are expected