Concepts
How it works
Three layers turn a sentence into rendered UI and executed actions.
Every turn moves through three layers:
POST, one stream, and one row of usage. The stream ends whenever the server needs something from the device; the device answers by opening the next one. Nothing is held open waiting.- Conversation surface. The branded launcher and thread UI your users see. Text in, streamed components out.
- Generative UI engine. The model composes an answer from a typed component catalog. The server validates every block against the catalog schema before it ever reaches the device; the SDK never renders raw model output.
- Action engine. Registered actions execute inside your app, through your own backend and the user's own permissions. The copilot proposes; your code performs.
The request lifecycle
user message
→ gate key, rate limits, quota
→ redact PII masked before any model call
→ ground knowledge retrieval + app state
→ route fast model for simple turns, frontier for complex ones
→ generate render_ui tool + your registered actions
→ validate every block schema-checked, then streamed
→ render native widgets on device
→ act handlers run in-app, results return to the modelThe stream closes whenever the copilot needs something from the device (an action result, a confirmation). The SDK answers with a continuation request. Mobile networks drop connections; this design makes that boring instead of fatal.
What the copilot cannot do
- Call anything that is not a registered action.
- Execute a write or destructive action without the user tapping a confirmation card.
- See raw emails, phone numbers or card numbers: they are masked before the model call and restored after validation.
- Render UI that failed schema validation: invalid blocks degrade to their plain-text fallback.