Concepts

How it works

Three layers turn a sentence into rendered UI and executed actions.

Every turn moves through three layers:

The anatomy of one turn: a POST opens an event stream, the stream ends when the server needs the device, and the device answers with another POST.On the device“Where is my order?”POST /v1/chatOn our serversRedact, retrieve, routeEmails and phone numbers becomeplaceholders before anything leaves hereHaiku for a simple turn, Sonnet for a hard onethe turnBlocks, streamedblock.start · delta · readyturn.end awaiting_tool_resultstext/event-streamNative componentsdrawn by your app, not a web viewIf the turn asked for somethingYour handler runsin your app, as your userThe user confirmsfor a write, not a readThe next turntool_resultsSame conversation, same id, and the usage row was written once.A retry carrying the same turn_id is not a second turn.
One turn is one POST, one stream, and one row of usage. The stream ends whenever the server needs something from the device; the device answers by opening the next one. Nothing is held open waiting.
  1. Conversation surface. The branded launcher and thread UI your users see. Text in, streamed components out.
  2. Generative UI engine. The model composes an answer from a typed component catalog. The server validates every block against the catalog schema before it ever reaches the device; the SDK never renders raw model output.
  3. Action engine. Registered actions execute inside your app, through your own backend and the user's own permissions. The copilot proposes; your code performs.

The request lifecycle

user message
  → gate        key, rate limits, quota
  → redact      PII masked before any model call
  → ground      knowledge retrieval + app state
  → route       fast model for simple turns, frontier for complex ones
  → generate    render_ui tool + your registered actions
  → validate    every block schema-checked, then streamed
  → render      native widgets on device
  → act         handlers run in-app, results return to the model

The stream closes whenever the copilot needs something from the device (an action result, a confirmation). The SDK answers with a continuation request. Mobile networks drop connections; this design makes that boring instead of fatal.

What the copilot cannot do

  • Call anything that is not a registered action.
  • Execute a write or destructive action without the user tapping a confirmation card.
  • See raw emails, phone numbers or card numbers: they are masked before the model call and restored after validation.
  • Render UI that failed schema validation: invalid blocks degrade to their plain-text fallback.