AI tools ·
Alibaba’s Qwen3.8-Omni-Flash slashes AI audio costs by 98%
What’s the newest cheap omnimodal AI model you can try right now? Alibaba released Qwen3.8-Omni-Flash on September 18, 2026 — a native omnimodal model that handles text, image, audio, and video inputs, with a 1-million-token context window and audio input pricing cut 98% versus the prior generation. It’s available now in the Qwen chat interface and mobile app (pick it from the model picker) and through an OpenAI-compatible API.
International pricing on QwenCloud: $0.15 per 1 million input tokens and $0.47 per 1 million output tokens ($0.016 for cached input). Developers connect with standard OpenAI-compatible client libraries using the DashScope base URL and a DASHSCOPE_API_KEY. Alibaba’s own benchmarks put the model close to Gemini 3.8 Flash: OmniVideoBench rose from 63.4 to 67.8 while token consumption fell 45.7% (Neowin) — though those figures are Alibaba-reported, not independently verified.
The launch also included a realtime variant, Qwen3.8-Omni-Flash-Realtime, plus two open-source agent projects: Qwen-MM-Plugins and Qwen-Live Harness. API regions include Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia.
Why it matters
Audio-heavy AI work — transcription, meeting analysis, voice agents, video understanding — has been one of the priciest AI workloads to run. Cutting audio input costs 98% at near-frontier quality changes the math for startups and developers, especially those priced out of western frontier APIs.
Sources
- Neowin — Alibaba’s Qwen3.8-Omni-Flash undercuts Gemini on audio
- MarkTechPost — Alibaba Qwen releases Qwen3.8-Omni-Flash
- RuntimeWire — Alibaba Qwen3.8-Omni-Flash audio, video, agents
Spot an error? We correct quickly and note it. Contact us via the contact page.