September focused on making failover safer to trust and keeping the model catalog current. We made provider failover credential-aware across every request path, surfaced the full attempt history in API responses and request logs, refreshed the catalog with the latest frontier models, and brought the pricing page to Chinese.
Credential-aware failover
Failover now treats a broken credential differently from a provider outage. When a primary attempt fails with an authentication error, ModelRiver only falls back to a different provider — it will not retry another model on the same provider with the same invalid key. If a workflow's first backup is another model from the same provider, that backup is skipped and the next different-provider backup is used instead.
Other changes in this release:
- A missing or unconfigured provider credential now fails the request directly instead of quietly routing to a backup, so a configuration mistake surfaces immediately.
- Authentication is detected more precisely: a 403 is only treated as a credential problem when the provider's error explicitly says the key is invalid or revoked, so policy and entitlement denials no longer trigger a cross-provider retry.
- The same rules apply to sync, streaming, async, and Playground requests, so Test Mode reflects how production routing actually behaves.
Failover visibility in responses and logs
You can now see exactly which provider served a request and every attempt it took to get there.
- API responses include a
meta.attemptsarray with each attempt's provider, model, status, and timing, plusmeta.used_providerandmeta.used_modelfor the model that ultimately answered. Error responses include the attempt list too. - Each failed attempt is recorded as its own request log entry with an attempt sequence, and is linked to the request that eventually succeeded, so one timeline shows the whole chain.
- The request detail view lists the failed models alongside their error responses.
Read provider failover attempts for the timeline and provider reliability for tracking failover rates.
Model catalog refresh
We refreshed the catalog to current releases across providers, with updated pricing:
- OpenAI / Azure OpenAI: GPT-6 Astra (current flagship for the hardest reasoning and coding, $10/$50 per million tokens), GPT-6 Sol (coding and agentic, 1.05M-token context, $2/$10), and GPT-6 Luna (focused, high-volume workloads, $0.10/$0.50).
- Anthropic: Claude Opus 5.5 (adaptive thinking, 1M context, 128K output, $4/$20).
- xAI:
grok-4.7, the current flagship for coding and agents ($2/$6). Grok 4.6 is now the previous flagship. - Google: Gemini 3.8 Flash is the new GA Flash for agentic and multimodal work.
- DeepSeek: V4.1 Flash replaces V4 Flash as the multimodal default for chat and high-volume workloads, now $0.15/$0.60 per million tokens.
- Groq: added Safety GPT OSS 20B, a safety-focused open model, alongside Qwen 3.8 27B.
- Qwen: added the Qwen 3.8 family — 27B (1M-context multimodal for coding and vision), 2.4T A95B (open flagship reasoning), Omni Flash (text, image, audio, and video input), and new Coder Next/Flash and VL Plus/Flash models — plus pinned Qwen 3.8 Max and 3.7 snapshots.
- Amazon Bedrock: added GPT-6 Astra, Sol, and Luna plus GPT-5.6 models. Sol and Luna run on Bedrock's OpenAI-compatible Chat Completions API, while Astra and the Claude models use the Converse API. Claude Opus 5.5 is also available.
Browse the full lineup and current pricing on the provider catalog.
Pricing in Chinese and clearer enterprise contact
The pricing page body copy is now available in Chinese, so teams comparing plans can read the full plan descriptions in their own language. The enterprise plan's "Talk to Us" button now opens the contact page instead of an email link.
Token calculator improvements
The token calculator now explains how its estimates are derived (the characters-per-4 rule of thumb for English GPT and Claude models), notes that cost comparisons assume equal input and output at standard on-demand rates, and clarifies that your text is counted entirely in the browser. We also added a "Which models does this estimate approximate?" FAQ.
Also this month
We streamlined the console by removing the legacy guided walkthrough.
