Back
AI News

Qwen3.8-Omni-Flash: million-token context and low list prices

Alibaba's Qwen team is positioning its new native omni-modal model for audio and video agents: by the company's own account, its average score is more than 25% better than its predecessor, while audio input has become dramatically cheaper.

AIMag.no
AIMag.no
September 18, 2026 · 6 min
Illustration: a long paper strip covered in dense markings unspools from a small compact cassette player on a gray floor, symbolizing million-token context and low prices.

Qwen3.8-Omni-Flash: million-token context and low list prices

Alibaba's Qwen team is positioning its new native omni-modal model for audio and video agents: by the company's own account, its average score is more than 25% better than its predecessor, while audio input has become dramatically cheaper.

Alibaba's Qwen team is positioning its new native omni-modal model for audio and video agents: by the company's own account, its average score is more than 25% better than its predecessor, while audio input has become dramatically cheaper. But all of the performance figures come from Qwen's own tests.

What actually launched

On September 18, 2026, the Qwen team launched Qwen3.8-Omni-Flash, which the company describes as its "next-generation native omni-modal model" (Qwen blog, source 6bbb2cef). The model accepts text, images, audio, and video as input, and it is marketed with a context window of one million tokens — that is, capacity to work with very long audio recordings, video streams, and documents in a single continuous session.

The details behind the marketing language are partly verifiable. Techgenyz has reviewed Alibaba's own documentation in Alibaba Cloud Model Studio, where the model is listed with a one-million-token window — but with concrete per-mode ceilings: 991,808 tokens of input in non-thinking mode and 983,616 tokens in thinking mode (Techgenyz, source 1c30cb39). In other words: "one million" holds up as marketing, but the practical ceiling sits just below it.

TechNode reports that the model is available via the Qwen AI platform and confirms the launch date (TechNode, source b625faff).

What it costs — documented prices and steep cuts

Pricing is one of the most concrete elements of the launch. Alibaba Cloud's Singapore-region pricing page lists the following input prices (Techgenyz, source 1c30cb39):

  • Text: $0.43 per million tokens
  • Audio: $3.81 per million tokens
  • Image and video: $0.78 per million tokens

TechNode reports that the API input price has been cut to as low as RMB 0.8 per million tokens — but with an important caveat: TechNode cites the Chinese outlet IT Home, and the original source is not available in our source material, so the figure rests on secondary reporting (source b625faff).

The percentage cuts are a different matter. Qwen itself claims that the API price per hour of audio input has fallen by more than 98 percent, and the price per hour of audio and video input by more than 93 percent (source 6bbb2cef). These figures are the company's own, not independently verified.

The company's performance claims — with the doubt visible

Qwen's blog is the source of all the performance figures in the launch, and they should be read as the company's own results. The key ones:

  • An average score improvement of more than 25 percent compared with the predecessor Qwen3.5-Omni-Plus, measured across 29 evaluations (source 6bbb2cef).
  • According to Qwen's blog, the model scores 36.5 on WildClawBench-MM and 22.3 on AgenticVBench — benchmarks aimed at audio and video agents, coding, and long-running tasks. The blog presents this as an improvement but does not state the baseline it is measured against; what we can quote directly is the scores.
  • A score of 69.6 on UniClawBench (source 6bbb2cef).

There is also a numerical disagreement between the sources worth noting: Qwen's blog says "29 evaluations" and "more than 25 percent," while both TechNode and Techgenyz report "30 evaluations" and "more than 26 percent." We report Qwen's blog as the primary source and have not been able to resolve the discrepancy.

Qwen also claims that the model, "by scaling data, context, and agent environments," achieves "audio and video performance close to Gemini 3.8 Flash and general audio performance that exceeds Gemini 3.8 Flash" (source 6bbb2cef). The comparison rests on a pricing model specified by Qwen itself — based on estimates of input cost at 30 times two minutes — that has not been independently audited. None of the benchmark names (WildClawBench-MM, AgenticVBench, UniClawBench, LongAudioSpan, OmniVideoBench) can be checked against public leaderboards based on the available documentation.

The agent framework: Qwen-MM-Plugins and Qwen-Live Harness

What distinguishes this launch from an ordinary model update is the tooling around it. Alibaba has released Qwen-MM-Plugins and open-sourced Qwen-Live Harness, which according to TechNode and the Qwen blog are aimed at long-running and real-time workflows (sources b625faff and 6bbb2cef).

The model is said by the company to support four concrete use cases: long-video analysis, meeting summaries, video research, and multimodal tool use (source b625faff). It is this combination — one million tokens of context, multimodal input, and agent tooling — that constitutes the company's real sales pitch: that the model should be able to power agents working over long audio and video streams, not just answer one query at a time.

It is worth noting that the "Flash" name suggests positioning as fast and cost-effective rather than as a flagship — and the price list points in the same direction, with a text input price of $0.43 per million tokens.

What remains to be proven

There are three open questions at this point.

First: all the benchmark results come from Qwen's own tests. No independent third-party evaluation or listing on public leaderboards exists in the material we have had access to. The comparison with Gemini 3.8 Flash is the company's own assessment, based on a pricing model defined by the company itself.

Second: the numerical disagreement between the sources (29 versus 30 evaluations, more than 25 versus more than 26 percent) is unresolved and may indicate different versions of the company's own figures.

Third: several of the figures connected to TechNode and IT Home — such as the RMB 0.8 price — rest on secondary sources, not Alibaba's original communications.

What would change the picture is independent evaluations, or the model appearing on public leaderboards. Until that happens, the safest things one can say are that Qwen has launched a documentedly inexpensive multimodal model with one million tokens of context and a tool suite aimed at audio and video agents — and that the performance claims, for now, rest solely on the company's own word.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. Qwenqwen.ai
  2. Qwen3.8-Omni-Flash Makes a Powerful Leap with 1M-token Context for Audio and Video AI - Techgenyztechgenyz.com
  3. Alibaba’s Qwen releases Qwen3.8-Omni-Flash with 1M-token context · TechNodetechnode.com