Consider an AI feature for damaged-package support tickets. The product goal is to analyze the ticket, use order information, route the case, and draft a reply. The prototype starts smaller, with one round trip:
The backend authenticates the customer, sends the ticket text to a model, and returns a suggested reply. The first implementation works. It has no application-data lookup or side effect yet. For a short request, this is a sound design: it is easy to understand, debug, and operate. A direct provider API may be all the feature needs.
The model call is simple until the model needs something the application owns. The following walkthrough evolves this one support workflow; it is illustrative, not a claim about a particular customer deployment.
Stage 1: The model needs application data
The model can read the ticket, but it does not know whether the order exists, when it shipped, or which support policy applies. That information belongs to the application. The backend needs to retrieve it and provide only the relevant context:
The model should not become the application's data layer. Giving it unrestricted database credentials would couple generated output to privileged access and remove a reliable boundary for deciding whose records it may see.
The backend authenticates the request, checks which account and ticket the customer may access, fetches relevant fields, and passes bounded context to the model. It remains the authority for authorization, application state, business rules, data access, and result validation. The model can interpret evidence; it does not decide whether that evidence may be read.
The engineering question is how AI participates in the application's data flow without becoming the data layer. The prompt can contain the ticket, shipment status, and the relevant policy excerpt the backend selected. It does not need the entire customer record.
Stage 2: The model needs to trigger an action
Suppose the model classifies the case as urgent and recommends moving it to the damaged-goods queue. The product now needs to change a ticket, not just display generated prose.
The model can return a structured proposal:
The model's decision: this ticket appears urgent and may belong in the damaged-goods queue.
The application's authority: the backend checks that the queue exists, that the customer and workflow may update this ticket, and that the transition is valid for its current state. It then makes the change through the application's domain code and records the outcome.
The proposal can be wrong, out of date, or outside the allowed set. The backend validates it and owns the side effect. If processing is retried after a timeout, the update must not happen twice. The model can suggest what should happen; application code decides whether it is allowed and executes it.
Now the request crosses a boundary between model judgment and application authority.
Stage 3: One request becomes several steps
The support team wants the final reply to mention the verified ticket status and next step. A useful flow is now: classify the ticket, let the backend check the order and update the ticket, then ask a second model workflow to draft a reply using the verified result.
Each stage needs the right inputs from the earlier ones. The second model should receive the verified shipment status and the ticket action that actually succeeded, not the first model's assumptions. That handoff is intermediate state: facts and results that must survive while work moves between systems.
Once one user action creates several model and backend interactions, the system needs to know which event belongs to the original request. ModelRiver's channel_id identifies the async request, and pipeline webhooks include a step index to identify the configured stage. The identifier ties the events together; it does not make them execute exactly once.
Ordering and partial failure now matter. Suppose a webhook handler updates the ticket, but its HTTP acknowledgement never reaches ModelRiver. Delivery may retry, so the application needs an idempotency check before applying the same action again. The callback boundary has a separate guard: ModelRiver checks callback step order and treats a repeated completed callback idempotently. A provider retry is different again: failover may select another configured model and add latency, which is visible in the provider timeline. Webhook delivery has its own retry policy. Each boundary needs an error path that says which step failed and whether the workflow can continue.
Model choice becomes part of the workflow too. Classification may need a fast, constrained output; reply drafting may use a different model or configuration. Logs should let an engineer locate the failing stage and inspect its provider attempts and intermediate result. ModelRiver's multi-step pipelines run ordered backend events with optional linked AI workflows and carry data between steps. The important design choice is that each handoff is explicit, not that every feature needs an agent framework.
Stage 4: The browser is still waiting
The multi-step flow may take several seconds or longer. With POST → wait → response, the browser holds a request open until the entire pipeline finishes. If the connection drops, the customer may not know whether the ticket was updated. The system needs to separate starting the request from delivering its result:

The client needs a return path while the backend and AI work continue.
An asynchronous request separates starting work from receiving its result. The application backend calls ModelRiver's POST /v1/ai/async endpoint and gets a pending response containing a channel_id, a single-use ws_token, and WebSocket connection details. The backend can return the request handle to the browser without exposing its ModelRiver API key. The browser then connects to the /socket endpoint and joins the request channel:
The token authorizes the socket connection; the project and channel_id identify its topic. ModelRiver broadcasts request updates and the final response on that channel. A streaming async workflow can also emit stream_start, stream_delta, stream_done, and stream_error. Token deltas are different from pipeline progress: a backend callback may mark a step ready or complete even when no model tokens are being streamed.
If the browser disconnects while a request is pending, the work can continue. The client can ask its own backend for a fresh one-time token and reconnect to the pending request. The application still owns the user-facing state: it decides whether the result belongs in the open ticket, whether the customer may view it, and what to do if the ticket changed meanwhile.
WebSockets are not mandatory for every AI interaction. For short synchronous requests, a normal HTTP response is simpler. For token streaming in a synchronous request, ModelRiver also supports Server-Sent Events (SSE). The async WebSocket path is useful when the work continues beyond the initial request or when the client needs updates while it runs. See the async API, React client SDK, and streaming guide for the specific connection patterns.
Stage 5: The application backend still owns the application
ModelRiver does not own the support product's end-user identity, domain rules, records, or sensitive actions. The application backend remains responsible for:
- Authenticating the user and authorizing access to the ticket and account.
- Reading and changing records through its own data layer.
- Applying support policy and validating model output against domain rules.
- Executing sensitive operations and recording the resulting application state.
- Deciding what result the user is allowed to see.
For a configured ModelRiver backend pipeline, the initial workflow sends its result to the application's webhook endpoint. The signed webhook includes the channel_id, generated data, and a callback_url. The backend verifies it using the documented signature verification process, applies its own rules and actions, then calls the callback URL. ModelRiver can advance to a linked AI step or another backend event; when the configured pipeline completes, it broadcasts the final response on the request's WebSocket channel.
This is a division of responsibility, not a transfer of application ownership. ModelRiver coordinates configured AI work and the communication around it; the application backend remains the authority over application data and actions. The backend pipeline guide covers the webhook, callback, delivery, and error paths.
Stage 6: The connections become the system
Step back from the support ticket. The model call is one component; the lifecycle around it is the system:
Nothing here is exotic on its own. The maintenance burden is connecting it reliably. If building the lifecycle in application code, the team may need provider integrations, routing and failover, workflow coordination, request correlation, webhook verification and retries, callbacks, WebSocket delivery, client reconnection, and request-level error visibility. The question is not how many features a platform has; it is how many boundaries the application team must connect and maintain.
The team needs to trace the initial model attempt, any provider failover, webhook delivery, backend callback, linked model step, and final client result. The request timeline connects those operational events. Otherwise, a successful model response can hide a failed ticket update, or a working backend can appear broken because the browser never received the final response.
This is the boundary question behind the gateway, framework, and ModelRiver distinction:
Different infrastructure layers simplify different boundaries in the same application.
An AI gateway primarily addresses reliable application-to-provider connectivity, routing, and gateway operations. An LLM framework helps construct AI logic, tools, workflow state, and orchestration. ModelRiver includes meaningful gateway capabilities, including provider connectivity, routing, and failover; its async pipeline also connects configured AI work with backend webhooks, callbacks, and WebSocket delivery. These layers can be combined. ModelRiver does not replace the support backend or require replacing its AI framework.
Once the problem is viewed as a lifecycle instead of a model call, the infrastructure boundary changes. The useful question is which parts of that lifecycle the application team wants to operate itself.
When the simple architecture is enough
If a feature only needs Backend → Model → Response, keep that architecture. A direct provider API may be easier to reason about, with fewer moving parts and fewer operational dependencies. A gateway can still be useful when multiple providers, routing, or failover matter; it does not require the application to adopt an asynchronous pipeline.
The broader lifecycle earns its cost when a request needs application data, backend actions, multiple AI interactions, asynchronous processing, external services, or real-time client updates. Even then, add only what the feature needs. A short synchronous classifier does not need a WebSocket channel just because another feature uses one.
What we learned
The model should not own your application. It can interpret evidence and propose an action. The backend remains responsible for permissions, business rules, data, validation, and side effects.
The model call is not always the request lifecycle. One user action may span several model attempts, backend steps, and callbacks. State and correlation make that flow understandable.
Delivery belongs in the design. If work continues asynchronously or streams incrementally, returning its result to the right client is part of the architecture. It is not an afterthought.
Infrastructure boundaries solve different problems. Gateways, frameworks, application backends, and client transports overlap in places, but each has a different center of responsibility.
The connections can become the expensive part to maintain. Teams often have good reasons to keep their backend and AI logic. The hard part is the surrounding infrastructure that lets both participate in one reliable request.
The first architecture still has its place:
When the request needs data, actions, or asynchronous delivery, the shape changes:
The difficult part of production AI is often not getting a model to answer. It is making that answer participate safely and reliably in the rest of the application. ModelRiver is one infrastructure option at that boundary: it combines gateway capabilities with configured backend callbacks and client delivery. For the concrete async flow, see the backend pipeline guide.

