Step 3.7 Flash
StepFun released 2026-05-28
Qwen3.6 35B A3B is both better and cheaper.
It scores +1.2 higher and costs 19% less ($0.355/M against $0.438/M) on this workload — and it does everything this model does.
DeepSeek V4 Flash 0731 is cheaper still (76% less) but drops no image, video input.
Every price dimension
| Input | $0.200/M |
|---|---|
| Output | $1.15/M |
| Cached input | $0.040/M |
Independent scores
| Intelligence | 30.9 |
|---|---|
| Coding | 39.6 |
| Agentic | 21.7 |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Best rankings by task
- asciiart #27of 83 1181
Capability
| Context window | 262K |
|---|---|
| Max output | 256K |
| Input modes | text, image, video |
| Tool use | yes |
| Reasoning | always on |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...