DoublewordDoubleword
Get started

Model Name

Qwen/Qwen3.5-35B-A3B-FP8-dottxt

Qwen3.5 35B A3B dottxt

  • Type: Generation
  • Capabilities: reasoning, enhanced_structured_generation, vision
  • Cache read: $0.20 per 1M input tokens (0.5714× standard input price). See prompt caching.

Overview

Qwen3.5-35B-A3B is a high-intelligence, mid-sized model that hits a very compelling price/performance point for async workloads. In Qwen's published benchmarks, this model outperformed GPT-5-mini, GPT-OSS-120B, and Claude Sonnet 4.5.


Thinking Mode:

This model reasons step-by-step before responding by default.

This model does not support graduated thinking levels. Parameters such as reasoning_effort are not supported and will have no effect.

Reasoning efforts

  • Supported: none, minimal, low, medium, high, xhigh, max

See the reasoning effort guide for request examples.

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$0.35$0.20$1.20
Async$0.14$0.08$0.60
Batch (24h)$0.10$0.06$0.40

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API.