Model Name
meta-models/Muse-Glimmer-30Bmeta-models/Muse-Glimmer-30B
- Type: Generation
- Capabilities:
vision,reasoning - Cache read: $0.01 per 1M input tokens (0.1× standard input price). See prompt caching.
Overview
Muse Glimmer is a 30-billion-parameter language model with a dedicated image encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a small model.
Reasoning: This model controls reasoning via prompting, not request parameters. Reasoning strength can be defined as part of the system prompt as Reasoning strength:
Reasoning efforts
- Supported:
none,minimal,low,medium,high,xhigh,max
See the reasoning effort guide for request examples.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $0.11 | $0.01 | $0.32 |
| Async | $0.07 | $0.01 | $0.24 |
| Batch (24h) | $0.05 | $0.01 | $0.15 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩