Qwen 3.7 Flash Spotted: What We Know About the Next Open-Weights Qwen (July 2026)
A model called qwen3.7-flash appeared on OpenRouter on July 27, 2026, before any Qwen blog post or Hugging Face repo. The listing is real and we verified it: 1M native context, $0.03 per million input tokens, $0.13 per million output. That is roughly 6x cheaper on input than Qwen3.6 Flash. What the listing does not tell us is parameter count, architecture, or whether weights ship at all. This page separates the two: what is confirmed on the listing, and what is inference. We'll update it when weights land.
Waiting on weights that may not fit anyway?
See our AI training options. We'll pick a model that runs on the machine you already own and wire it into OpenClaw today.
Bottom Line (July 2026)
- Confirmed:
qwen3.7-flashis listed on OpenRouter with a July 27, 2026 release date. We fetched the page. - Confirmed: 1,000,000-token context window, up to 65,536 output tokens.
- Confirmed: $0.03 per million input tokens, $0.13 per million output tokens.
- Confirmed: multimodal — text, image, and video in, text out. OpenRouter describes it as a vision-language reasoning model aimed at multimodal agents, visual coding, search, and computer interaction.
- Confirmed by comparison: Qwen3.6 Flash on OpenRouter is $0.1875 input / $1.125 output. The new listing is about 6x cheaper on input and 8.7x cheaper on output.
- Not confirmed: parameter count, active-parameter count, architecture, license, whether open weights ship at all, or when.
- Community inference, not fact: the Flash name has tracked Qwen’s small MoE tier, so people expect something near the 35B-total / 3B-active class. Pricing this low is consistent with that. It is still a guess.
- Why local users care: a small MoE with 1M native context would land directly on the gap the local crowd has been complaining about all year.
What Is Actually Confirmed
There is no Qwen blog post, no Hugging Face repo, and no model card. What exists is an API listing, and we checked it directly rather than repeating secondhand numbers.
| Field | Value on the OpenRouter listing |
|---|---|
| Model ID | qwen/qwen3.7-flash |
| Listed release | July 27, 2026 |
| Context window | 1,000,000 tokens |
| Max output | 65,536 tokens |
| Input price | $0.03 / 1M tokens |
| Output price | $0.13 / 1M tokens |
| Modalities | text + image + video in, text out |
That is the whole factual base. Everything below this line is either comparison against published Qwen models or explicitly labeled inference.
The Price Drop Is the Loudest Signal
Set the two Flash listings side by side:
| Qwen3.6 Flash | Qwen3.7 Flash | |
|---|---|---|
| Listed | April 27, 2026 | July 27, 2026 |
| Input | $0.1875 / 1M | $0.03 / 1M |
| Output | $1.125 / 1M | $0.13 / 1M |
| Context | 1M (tiered pricing above 256K) | 1M |
Roughly 6x cheaper input and 8.7x cheaper output in three months, at the same advertised context. Providers do not cut serving prices by that much without something changing underneath — usually fewer active parameters per token, a better-quantized serving stack, or both.
One detail worth flagging: the 3.6 Flash listing prices context in tiers above 256K, which lines up with the open-weights Qwen3.6-35B-A3B shipping 262,144 native tokens and stretching to about 1M with YaRN. The 3.7 Flash listing shows no such tier break. That is thin evidence, but it is the kind of thin evidence that points at 1M being native this time rather than extended.
What the Community Thinks It Is
Attribute this correctly: this is the read circulating among local-model users, not a Qwen statement.
The argument is simple. Qwen has used “Flash” for its small, sparse, cheap tier. The open-weights model that sat behind that tier in the 3.6 generation was Qwen3.6-35B-A3B — 35 billion total parameters, 3 billion active, Apache 2.0, released April 16, 2026, and it scores 73.4% on SWE-bench Verified. If 3.7 Flash inherits the naming convention, people expect another small MoE, plausibly in the same 30-40B total / low-single-digit-billions active neighborhood.
Nobody has published a parameter count. Treat “small MoE” as the most probable shape, not a spec.
Why This One Matters to Local Users
The loudest recurring request in local-model threads this year has not been for a better frontier model. It has been for releases that fit on hardware people own. One widely-upvoted comment put it plainly — “We could really use releases in 27B, 35B, 122B sizes” — and drew 541 upvotes. The frustration is a direct response to 2026’s pattern: open weights that are not actually local, like Kimi K3 at ~1.4TB.
The second complaint is context. Agent loops eat context fast, and the practical ceiling on a good local MoE has been 262K. Long sessions hit that wall, then you are managing compaction instead of working. If 3.7 Flash really carries 1M natively and an open-weights version follows, that is the first time the small tier and the long-context tier are the same model.
Those two asks together — small enough to load, long enough to agent with — describe exactly what this listing looks like. That is why the thread moved fast on an endpoint with no model card.
What We Do Not Know
Be honest about the size of this list. It is longer than the confirmed one.
- Parameter count and architecture. Not published anywhere. Total, active, expert count, all unknown.
- Whether open weights ship at all. An API listing is not a weights release. Qwen’s track record is good — Apache 2.0 on 3.6-35B-A3B — but track record is not an announcement.
- Timing. No date, no repo, no teaser. Qwen has shipped weights before and after API availability in different generations.
- License. Assume nothing. 2026 has already seen a major lab tighten terms between generations.
- Whether 1M is native on the weights. The served endpoint says 1M. Served endpoints run on datacenter hardware with tricks that a GGUF on your desk does not get.
- Real agentic quality. No benchmarks. The OpenRouter description mentions visual coding and computer interaction; that is marketing copy, not a score.
- Quantized footprint. Without a parameter count there is no honest VRAM math to do yet.
If you see a post confidently listing 3.7 Flash’s parameters, VRAM requirement, or SWE-bench score right now, it is invented.
What to Expect on Fit — If the Small-MoE Read Holds
Conditional section. If 3.7 Flash lands as open weights in the same class as Qwen3.6-35B-A3B, the fit story is already known, because that model is published and measured.
Qwen3.6-35B-A3B is 35B total / 3B active, Apache 2.0, 262,144 native context extensible toward ~1M with YaRN, and 73.4% on SWE-bench Verified. A dynamic 4-bit quantization of it fits on a 24GB unified-memory Mac with room for the OS. Those are published numbers for 3.6, and they are the only defensible baseline until 3.7 weights exist.
So the reasonable expectation, stated as an expectation: a 24GB card or a 32GB Mac is the tier to watch. If you want current guidance for the machine you own, use the 3.6-generation pages — they reflect what actually runs today:
Do not buy hardware for 3.7 Flash. There is no spec to buy against. Buy for what runs now; if 3.7 arrives in the class people expect, the same machine will hold it.
What Would Change This Page
Three events, in order of how much they would settle:
- A Hugging Face repo appears. That gives parameter count, license, native context, and file sizes in one shot. Every unknown above collapses.
- A Qwen blog post or model card. Confirms architecture and benchmarks even before GGUFs exist.
- Unsloth or Ollama builds land. That is when the practical question — does it fit 24GB — gets a real answer instead of an analogy to 3.6.
We’ll update this page when weights land. Until then the honest summary is: a cheap, long-context, multimodal Qwen endpoint exists, it is priced like a small model, and nothing has been said about running it yourself.
Common Mistakes
- Treating the OpenRouter listing as a weights release. It is a served API endpoint. The two have shipped weeks apart before, and sometimes only one ships.
- Quoting a parameter count. None is published. “35B-A3B class” is a naming-based guess and should be labeled as one every time it is repeated.
- Assuming 1M context survives quantization. Even where native long context is real, KV cache at long context costs memory that a 24GB card does not have spare. See context window traps for local agents.
- Assuming Apache 2.0. Qwen has been generous historically. Historical generosity is not a license file.
- Delaying a local setup to wait for it. There is no announced date. A working OpenClaw + local model stack today swaps to a new model in one config line later.
Related Guides
- Can You Run Kimi K3 Locally? — the other side of 2026: frontier open weights that no desk can hold
- Open Weights Aren’t Local Anymore — why “open” and “runnable” split into separate axes this year
- Qwen 3.6 vs Gemma 4 for Agentic Work — the current small-MoE comparison, with real numbers
- Best Local LLMs for 24GB RAM — the tier a 35B-A3B-class model lands in
- Best Local LLM by RAM (hub) — pick your tier first, then the model
- OpenClaw Best Local Models — which local models hold up in long agent loops
- Context Window Traps for Local Agents — why a 1M context number rarely means 1M usable tokens on your machine
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session