AI tools ·

Alibaba’s Qwen3.8-Omni-Flash slashes AI audio costs by 98%

What’s the newest cheap omnimodal AI model you can try right now? Alibaba released Qwen3.8-Omni-Flash on September 18, 2026 — a native omnimodal model that handles text, image, audio, and video inputs, with a 1-million-token context window and audio input pricing cut 98% versus the prior generation. It’s available now in the Qwen chat interface and mobile app (pick it from the model picker) and through an OpenAI-compatible API.

International pricing on QwenCloud: $0.15 per 1 million input tokens and $0.47 per 1 million output tokens ($0.016 for cached input). Developers connect with standard OpenAI-compatible client libraries using the DashScope base URL and a DASHSCOPE_API_KEY. Alibaba’s own benchmarks put the model close to Gemini 3.8 Flash: OmniVideoBench rose from 63.4 to 67.8 while token consumption fell 45.7% (Neowin) — though those figures are Alibaba-reported, not independently verified.

The launch also included a realtime variant, Qwen3.8-Omni-Flash-Realtime, plus two open-source agent projects: Qwen-MM-Plugins and Qwen-Live Harness. API regions include Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia.

Why it matters

Audio-heavy AI work — transcription, meeting analysis, voice agents, video understanding — has been one of the priciest AI workloads to run. Cutting audio input costs 98% at near-frontier quality changes the math for startups and developers, especially those priced out of western frontier APIs.

Key facts

  • Released September 18, 2026; native omnimodal (text, image, audio, video inputs) with a 1M-token context window
  • Audio input pricing cut 98% vs prior generation; $0.15/1M input tokens, $0.47/1M output tokens, $0.016 cached input
  • Available now in Qwen chat (web) and mobile app — must select manually from the model picker
  • OpenAI-compatible API via DashScope base URL + DASHSCOPE_API_KEY; regions include Beijing, Singapore, Frankfurt, Virginia
  • Alibaba-reported benchmarks close to Gemini 3.8 Flash (OmniVideoBench 63.4→67.8, 45.7% fewer tokens)
  • Also launched: Qwen3.8-Omni-Flash-Realtime plus open-source Qwen-MM-Plugins and Qwen-Live Harness

Sources

Spot an error? We correct quickly and note it. Contact us via the contact page.

← More AI news