支持的平台

对比 AI 提供商与模型

接入 13 家主流 AI 提供商,从 109 款模型中选择,涵盖 ChatGPT、Claude、Gemini、Grok、Azure OpenAI 和 Amazon Bedrock。查看可用选项、对比差异,无需重建产品即可切换提供商。.

AI 提供商
13

团队熟悉的主流平台:OpenAI、Anthropic、Google、xAI、Cohere、Mistral、DeepSeek、Qwen、Groq、Kimi、MiniMax、Azure OpenAI 和 Amazon Bedrock。.

可用模型
103

此处列出的模型均可在 ModelRiver 上直接使用,无需您额外配置。.

一次接入
1 次配置

团队只需接入一次。后续更换模型或提供商,无需推倒重来。.

OpenAI

GPT-5.6 Sol/Terra/Luna, GPT-5.5/5.4, GPT-4.1/4o, and GPT-5.3 Codex

20 models
Frontier, mini, codex, and reasoning tiers
gpt-5.6-sol
GPT-5.6 Sol - flagship frontier model for complex reasoning and coding
Chat Completion Text Vision Streaming Structured output Tools 1.1M ctx 128K out
1.1M context · 128K max out · Released 2026
$4 in
$20 out
gpt-5.6-terra
GPT-5.6 Terra - balances intelligence and cost for everyday professional work
Chat Completion Text Vision Streaming Structured output Tools 1.1M ctx 128K out
1.1M context · 128K max out · Released 2026
$2 in
$12 out
gpt-5.6-luna
GPT-5.6 Luna - fast low-cost model for high-volume workloads
Chat Completion Text Vision Streaming Structured output Tools 1.1M ctx 128K out
1.1M context · 128K max out · Released 2026
$0.2 in
$1.20 out
gpt-5.5
Previous frontier GPT model for coding and professional work
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2026
$5 in
$30 out
gpt-5.5-pro
Highest-compute GPT-5.5 variant
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2026
$30 in
$180 out
gpt-5.4
Production workhorse GPT model for professional work
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$2.50 in
$15 out
gpt-5.4-mini
Strong mini GPT model for coding and subagents
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2026
$0.75 in
$4.50 out
gpt-5.4-nano
Smallest current GPT-5.4 variant
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2026
$0.2 in
$1.25 out
gpt-5.4-pro
Highest-precision GPT-5.4 variant
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$30 in
$180 out
gpt-5.3-codex
Latest Codex model for agentic coding
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2026
$1.75 in
$14 out
gpt-5.2
Previous frontier GPT model with configurable reasoning
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2025
$1.75 in
$14 out
gpt-5.2-pro
Previous pro GPT model for professional work
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2025
$21 in
$168 out
gpt-5.1
Flagship GPT model for coding and agentic tasks
Chat Completion Text Vision Streaming Structured output Tools 400K ctx 128K out
400K context · 128K max out · Released 2025
$1.25 in
$10 out
gpt-4.1
Strong non-reasoning model for instructions and tools
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 32.8K out
1.0M context · 32.8K max out · Released 2025
$2 in
$8 out
gpt-4.1-mini
Efficient GPT-4.1 variant
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 32.8K out
1.0M context · 32.8K max out · Released 2025
$0.4 in
$1.60 out
gpt-4.1-nano
Lightweight GPT-4.1 model
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 32.8K out
1.0M context · 32.8K max out · Released 2025
$0.1 in
$0.4 out
gpt-4o-mini
Low-cost multimodal small model
Chat Completion Text Vision Streaming Structured output Tools 128K ctx 16.4K out
128K context · 16.4K max out · Released 2024
$0.15 in
$0.6 out
gpt-4o
Widely used multimodal flagship model
Chat Completion Text Vision Streaming Structured output Tools 128K ctx 16.4K out
128K context · 16.4K max out · Released 2024
$2.50 in
$10 out
text-embedding-3-small
Latest OpenAI small embedding model — efficient default for RAG
Embedding Text 8.2K ctx
8.2K context · Released 2024
$0.02 in
$0 out
text-embedding-3-large
Latest OpenAI large embedding model with highest accuracy
Embedding Text 8.2K ctx
8.2K context · Released 2024
$0.13 in
$0 out

Anthropic

Claude Fable 5.1/5, Opus 5, Sonnet 5, and Haiku 4.5

claude-fable-5-1
Claude Fable 5.1 - current Fable for demanding reasoning and long-horizon agentic work
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$10 in
$50 out
claude-opus-5
Claude Opus 5 - current Opus flagship for coding, agents, and enterprise work
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$5 in
$25 out
claude-sonnet-5
Latest Claude Sonnet - best speed/intelligence balance
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$2 in
$10 out
claude-fable-5
Previous Claude Fable model - still available; prefer Fable 5.1 for new work
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$10 in
$50 out
claude-opus-4-8
Previous Claude Opus model - still available; prefer Opus 5 for new work
Chat Completion Text Vision Streaming Structured output Tools 1M ctx 128K out
1M context · 128K max out · Released 2026
$5 in
$25 out
claude-haiku-4-5-20251001
Claude Haiku 4.5 - low latency and cost-efficient current Haiku
Chat Completion Text Vision Streaming Structured output Tools 200K ctx 64K out
200K context · 64K max out · Released 2025
$1 in
$5 out

Google

Gemini 3.8/3.7/3.6 Flash, 3.5 Flash/Flash-Lite, 3.1 Pro/Flash-Lite, and Gemini 2.5 family

gemini-3.8-flash
Gemini 3.8 Flash - latest GA Flash for agentic and multimodal work
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$0.75 in
$3.75 out
gemini-3.7-flash
Gemini 3.7 Flash - previous workhorse Flash for coding and agents
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$0.75 in
$3.75 out
gemini-3.6-flash
Gemini 3.6 Flash - previous GA Flash for agentic and multimodal work
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$1.50 in
$7.50 out
gemini-3.5-flash-lite
Gemini 3.5 Flash-Lite - fastest low-cost 3.5 model for high throughput
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$0.3 in
$2.50 out
gemini-3.5-flash
Gemini 3.5 Flash - frontier speed with strong search and grounding
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$1.50 in
$9 out
gemini-3.1-flash-lite
Gemini 3.1 Flash-Lite - cost-efficient high-volume Gemini 3.1 model
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$0.25 in
$1.50 out
gemini-embedding-2
Latest Gemini embedding model (GA Apr 2026) for text and multimodal RAG
Embedding Text 8.2K ctx
8.2K context · Released 2026
$0.2 in
$0 out
gemini-3.1-pro-preview
Gemini 3.1 Pro Preview - latest multimodal reasoning model
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2026
$2 in
$12 out
gemini-2.5-flash-lite
Gemini 2.5 Flash-Lite - ultra cost-efficient model
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2025
$0.1 in
$0.4 out
gemini-2.5-flash
Gemini 2.5 Flash - fast high-throughput model
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2025
$0.3 in
$2.50 out
gemini-2.5-pro
Gemini 2.5 Pro - high-capability multipurpose model
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 65.5K out
1.0M context · 65.5K max out · Released 2025
$1.25 in
$10 out

xAI

Grok 4.6/4.5 flagships and Grok 4.3 for coding, chat, and agents

grok-4.6
Grok 4.6 - latest frontier model for coding, agents, and knowledge work
Chat Completion Text Streaming Structured output Tools 500K ctx
500K context · Released 2026
$2 in
$6 out
grok-4.5
Grok 4.5 - previous flagship for coding, chat, and agents
Chat Completion Text Streaming Structured output Tools 500K ctx 64K out
500K context · 64K max out · Released 2026
$2 in
$6 out
grok-4.3
Grok 4.3 - lower-cost 1M-context model for everyday chat and coding
Chat Completion Text Streaming Structured output Tools 1M ctx 64K out
1M context · 64K max out · Released 2026
$1.25 in
$2.50 out

Cohere

Command A+, Command A, and current Command R family models

command-a-plus-05-2026
Command A+ - latest enterprise agentic model; hosted API is free with capped usage
Chat Completion Text Streaming Structured output Tools 128K ctx 64K out
128K context · 64K max out · Released 2026
$0 in
$0 out
command-a-reasoning-08-2025
Command A Reasoning - free rate-limited hosted reasoning model
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2025
$0 in
$0 out
command-a-translate-08-2025
Command A Translate - free rate-limited translation model
Chat Completion Text Streaming Structured output Tools 8K ctx 8K out
8K context · 8K max out · Released 2025
$0 in
$0 out
command-a-vision-07-2025
Command A Vision - free rate-limited multimodal model for image inputs
Chat Completion Text Vision Streaming Structured output Tools 128K ctx 8.2K out
128K context · 8.2K max out · Released 2025
$0 in
$0 out
command-a-03-2025
Command A - advanced instruction-following model (March 2025)
Chat Completion Text Streaming Structured output Tools 256K ctx 8K out
256K context · 8K max out · Released 2025
$2.50 in
$10 out
embed-v4.0
Latest Cohere embedding model for search, RAG, and multimodal retrieval
Embedding Text 128K ctx
128K context · Released 2025
$0.12 in
$0 out
command-r7b-12-2024
Command R7B - fastest low-cost Command model
Chat Completion Text Streaming Structured output Tools 128K ctx 4.1K out
128K context · 4.1K max out · Released 2024
$0.0375 in
$0.15 out
command-r-plus-08-2024
Command R+ - strong long-context agent model
Chat Completion Text Streaming Structured output Tools 128K ctx 4.1K out
128K context · 4.1K max out · Released 2024
$2.50 in
$10 out
command-r-08-2024
Command R - lower-cost long-context model
Chat Completion Text Streaming Structured output Tools 128K ctx 4.1K out
128K context · 4.1K max out · Released 2024
$0.15 in
$0.6 out

Mistral AI

Mistral Large/Medium/Small, Ministral, and Codestral

mistral-medium-latest
Mistral Medium 3.5 - frontier multimodal model for agents and coding
Chat Completion Text Streaming Structured output Tools 256K ctx 64K out
256K context · 64K max out · Released 2026
$1.50 in
$7.50 out
mistral-small-latest
Mistral Small 4 - fast and efficient multimodal model
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2026
$0.15 in
$0.6 out
mistral-large-latest
Mistral Large 3 - flagship general-purpose model
Chat Completion Text Streaming Structured output Tools 256K ctx 64K out
256K context · 64K max out · Released 2025
$0.5 in
$1.50 out
ministral-14b-latest
Ministral 3 14B - compact text and vision model
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2025
$0.2 in
$0.2 out
ministral-8b-latest
Ministral 3 8B - efficient edge-capable model
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2025
$0.15 in
$0.15 out
ministral-3b-latest
Ministral 3 3B - smallest Ministral 3 model
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2025
$0.1 in
$0.1 out
codestral-latest
Codestral - specialized coding model
Chat Completion Text Streaming Structured output Tools 128K ctx 32K out
128K context · 32K max out · Released 2025
$0.3 in
$0.9 out
codestral-embed
Mistral code-focused embedding model
Embedding Text 8.2K ctx
8.2K context · Released 2025
$0.15 in
$0 out
mistral-embed
Mistral embedding model for semantic search and RAG
Embedding Text 8.2K ctx
8.2K context · Released 2023
$0.1 in
$0 out

DeepSeek

DeepSeek V4 Pro and V4 Flash

deepseek-v4-pro
DeepSeek V4 Pro - frontier coding and long-horizon agents
Chat Completion Text Streaming Structured output Tools 1M ctx 384K out
1M context · 384K max out · Released 2026
$0.435 in
$0.87 out
deepseek-v4-flash
DeepSeek V4 Flash - default chat and high-volume workloads
Chat Completion Text Streaming Structured output Tools 1M ctx 384K out
1M context · 384K max out · Released 2026
$0.14 in
$0.28 out

Qwen

Qwen 3.8 Max/Flash, 3.7 Max/Plus, 3.6 Plus/Flash, and Coder (Alibaba Cloud)

qwen3.8-max
Qwen 3.8 Max - latest flagship Qwen model
Chat Completion Text Streaming Structured output Tools 1M ctx 131.1K out
1M context · 131.1K max out · Released 2026
$2 in
$6 out
qwen3.8-flash
Qwen 3.8 Flash - latest fast 1M-context cost-efficient model
Chat Completion Text Streaming Structured output Tools 1M ctx 131.1K out
1M context · 131.1K max out · Released 2026
$0.15 in
$0.47 out
qwen3.7-plus
Qwen 3.7 Plus - current balanced production model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2026
$0.4 in
$1.60 out
qwen3.7-max
Qwen 3.7 Max - previous flagship Qwen model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2026
$2.50 in
$7.50 out
qwen3.6-flash
Qwen 3.6 Flash - fast long-context cost-efficient model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2026
$0.25 in
$1.50 out
qwen3.6-plus
Qwen 3.6 Plus - previous balanced long-context production model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2026
$0.5 in
$3 out
qwen3.5-plus
Qwen 3.5 Plus - balanced production model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2026
$0.4 in
$2.40 out
qwen3.5-flash
Qwen 3.5 Flash - fast 1M-context cost-efficient model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2026
$0.1 in
$0.4 out
qwen3-coder-plus
Qwen3 Coder Plus - current coding and agent model
Chat Completion Text Streaming Structured output Tools 1M ctx 65.5K out
1M context · 65.5K max out · Released 2025
$1 in
$5 out
qwen3-max
Qwen3 Max - previous flagship Qwen model
Chat Completion Text Streaming Structured output Tools 262.1K ctx 65.5K out
262.1K context · 65.5K max out · Released 2025
$1.20 in
$6 out
qwen-flash
Qwen-Flash - fast and cost-efficient flagship model
Chat Completion Text Streaming Structured output Tools 1M ctx 32.8K out
1M context · 32.8K max out · Released 2025
$0.05 in
$0.4 out
qwen-max
Qwen-Max - flagship model, highest capability
Chat Completion Text Streaming Structured output Tools 131.1K ctx 16.4K out
131.1K context · 16.4K max out · Released 2025
$1.60 in
$6.40 out
qwen-plus
Qwen-Plus - balanced performance and cost
Chat Completion Text Streaming Structured output Tools 131.1K ctx 16.4K out
131.1K context · 16.4K max out · Released 2025
$0.4 in
$1.20 out

Groq

Ultra-fast LPU inference - GPT OSS and Qwen 3.8/3.6 27B

qwen/qwen3.8-27b
Qwen 3.8 27B - latest fast open model for coding and multilingual workloads
Chat Completion Text Streaming Structured output Tools 131.1K ctx 16.4K out
131.1K context · 16.4K max out · Released 2026
$0.8 in
$4 out
qwen/qwen3.6-27b
Qwen 3.6 27B - previous fast open model for coding and multilingual workloads
Chat Completion Text Streaming Structured output Tools 131.1K ctx 16.4K out
131.1K context · 16.4K max out · Released 2026
$0.6 in
$3 out
openai/gpt-oss-120b
GPT OSS 120B - open model with strong quality/cost balance
Chat Completion Text Streaming Structured output Tools 131.1K ctx 65.5K out
131.1K context · 65.5K max out · Released 2025
$0.15 in
$0.6 out
openai/gpt-oss-20b
GPT OSS 20B - very fast open model for lightweight workloads
Chat Completion Text Streaming Structured output Tools 131.1K ctx 65.5K out
131.1K context · 65.5K max out · Released 2025
$0.075 in
$0.3 out

Kimi

Kimi K3 flagship, K2.7 Code (standard and high-speed), K2.6, and K2.5 models

kimi-k3
Kimi K3 - flagship 2.8T multimodal model with 1M context
Chat Completion Text Streaming Structured output Tools 1.0M ctx 131.1K out
1.0M context · 131.1K max out · Released 2026
$3 in
$15 out
kimi-k2.7-code
Kimi K2.7 Code - latest coding model with 256K context
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2026
$0.95 in
$4 out
kimi-k2.7-code-highspeed
Kimi K2.7 Code Highspeed - the same coding model with higher output throughput
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2026
$1.90 in
$8 out
kimi-k2.6
Kimi K2.6 - flagship multimodal agent and coding model (256K context)
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2026
$0.95 in
$4 out
kimi-k2.5
Kimi K2.5 - multimodal model with thinking and non-thinking modes (256K context)
Chat Completion Text Streaming Structured output Tools 256K ctx 32K out
256K context · 32K max out · Released 2026
$0.6 in
$3 out

MiniMax

MiniMax M3 and M2.7/M2.5 coding/agent models

MiniMax-M3
MiniMax M3 - latest agentic reasoning and coding model (1M context)
Chat Completion Text Streaming Structured output Tools 1M ctx 64K out
1M context · 64K max out · Released 2026
$0.3 in
$1.20 out
MiniMax-M2.7
MiniMax M2.7 - flagship MoE coding and agent model (205K context)
Chat Completion Text Streaming Structured output Tools 204.8K ctx 128K out
204.8K context · 128K max out · Released 2026
$0.3 in
$1.20 out
MiniMax-M2.7-highspeed
MiniMax M2.7 Highspeed - same M2.7 quality with higher throughput
Chat Completion Text Streaming Structured output Tools 204.8K ctx 128K out
204.8K context · 128K max out · Released 2026
$0.6 in
$2.40 out
MiniMax-M2.5
MiniMax M2.5 - previous flagship coding/agent model, still supported
Chat Completion Text Streaming Structured output Tools 204.8K ctx 128K out
204.8K context · 128K max out · Released 2026
$0.3 in
$1.20 out
MiniMax-M2.5-highspeed
MiniMax M2.5 Highspeed - previous M2.5 with higher throughput
Chat Completion Text Streaming Structured output Tools 204.8K ctx 128K out
204.8K context · 128K max out · Released 2026
$0.6 in
$2.40 out

Azure OpenAI

Azure-hosted OpenAI deployments (GPT-5.6 Sol/Terra/Luna, GPT-4.1/4o) — set your resource /openai/v1 base URL; model must match deployment name

gpt-5.6-sol
GPT-5.6 Sol deployment — frontier Azure OpenAI model (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 1.1M ctx 128K out
1.1M context · 128K max out · Released 2026
$4 in
$20 out
gpt-5.6-terra
GPT-5.6 Terra deployment — balanced intelligence/cost (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 1.1M ctx 128K out
1.1M context · 128K max out · Released 2026
$2 in
$12 out
gpt-5.6-luna
GPT-5.6 Luna deployment — fast low-cost high-volume (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 1.1M ctx 128K out
1.1M context · 128K max out · Released 2026
$0.2 in
$1.20 out
gpt-4.1
GPT-4.1 deployment — strong instructions and tools (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 32.8K out
1.0M context · 32.8K max out · Released 2025
$2 in
$8 out
gpt-4.1-mini
GPT-4.1 mini deployment — efficient GPT-4.1 variant (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 1.0M ctx 32.8K out
1.0M context · 32.8K max out · Released 2025
$0.4 in
$1.60 out
gpt-4o-mini
GPT-4o mini deployment — low-cost multimodal (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 128K ctx 16.4K out
128K context · 16.4K max out · Released 2024
$0.15 in
$0.6 out
gpt-4o
GPT-4o deployment — workflow model must match your Azure deployment name (Global Standard pricing)
Chat Completion Text Vision Streaming Structured output Tools 128K ctx 16.4K out
128K context · 16.4K max out · Released 2024
$2.50 in
$10 out

Amazon Bedrock

Amazon Nova 2/Nova and Claude 5.1/5-family models via Bedrock Converse API (Bedrock API key)

us.anthropic.claude-fable-5-1
Claude Fable 5.1 on Bedrock (US inference profile) - current Fable for long-horizon work
Chat Completion Text Vision 1M ctx 128K out
1M context · 128K max out · Released 2026
$10 in
$50 out
us.anthropic.claude-opus-5
Claude Opus 5 on Bedrock (US inference profile) - current Opus flagship
Chat Completion Text Vision 1M ctx 128K out
1M context · 128K max out · Released 2026
$5 in
$25 out
us.anthropic.claude-sonnet-5
Claude Sonnet 5 on Bedrock (US inference profile) - speed/intelligence balance
Chat Completion Text Vision 1M ctx 128K out
1M context · 128K max out · Released 2026
$2 in
$10 out
us.anthropic.claude-fable-5
Claude Fable 5 on Bedrock (US inference profile) - previous Fable, still available
Chat Completion Text Vision 1M ctx 128K out
1M context · 128K max out · Released 2026
$10 in
$50 out
us.amazon.nova-2-lite-v1:0
Amazon Nova 2 Lite (US inference profile) - current cost-efficient multimodal reasoning model
Chat Completion Text Vision 1M ctx 64K out
1M context · 64K max out · Released 2025
$0.3 in
$2.50 out
us.anthropic.claude-haiku-4-5-20251001-v1:0
Claude Haiku 4.5 on Bedrock (US inference profile) - fast low-cost Claude
Chat Completion Text Vision 200K ctx 64K out
200K context · 64K max out · Released 2025
$1 in
$5 out
amazon.nova-lite-v1:0
Amazon Nova Lite - fast low-cost multimodal model (on-demand)
Chat Completion Text Vision 300K ctx 8.2K out
300K context · 8.2K max out · Released 2024
$0.06 in
$0.24 out
amazon.nova-pro-v1:0
Amazon Nova Pro - capable multimodal production model (on-demand)
Chat Completion Text Vision 300K ctx 8.2K out
300K context · 8.2K max out · Released 2024
$0.8 in
$3.20 out
amazon.nova-micro-v1:0
Amazon Nova Micro - lowest-cost text model (on-demand)
Chat Completion Text 128K ctx 8.2K out
128K context · 8.2K max out · Released 2024
$0.035 in
$0.14 out

一次搭建,随 AI 演变保持灵活

一次配置 AI 工作流,当价格、性能或战略调整时,可在不同提供商之间切换(含 Azure OpenAI 与 Amazon Bedrock),无需从零开始。.