Rate limits and caps
Every ceiling in the system, what it counts, and what happens when it is reached.
There are five ceilings. They exist for different reasons and fail in different ways, and knowing which one you hit is usually the whole diagnosis.
| Ceiling | Counts | Scope | Default | What happens |
|---|---|---|---|---|
| Messages per minute | Chat, identify and feedback requests | Per device, per app | 20 | HTTP 429 before the stream opens. |
| Messages per day | Turns that start work | Per app | 400 | rate_limited on the stream. |
| Conversations per month | New billable conversations | Per app | Plan | quota_exceeded on the stream. |
| Sandbox conversations per month | New conversations on a test key | Per app | 200 | quota_exceeded, with a message written for a developer. |
| Daily model budget | Model spend in dollars | Per app | None | quota_exceeded on the stream. |
Messages per minute
An in-process token bucket, per (app, device). Serverless instances each keep their own, so this is a soft first line rather than an exact number — it exists to absorb a stuck retry loop, not to meter anything.
GET /v1/apps/config has its own two buckets: 12/minute per device and 600/minute per app. The app-wide one is what still holds when a caller rotates X-Rendel-Device on every request, and what it actually protects is the database connection pool /v1/chat shares.
Messages per day
The durable version of the minute bucket. It counts rows rather than trusting per-instance state, and it applies to sandbox traffic too: a test key is extractable from any debug build you ship, and unmetered means unbounded model spend.
Two things are exempt on purpose:
- Retries of a turn already counted. Resend the same
turn_idand the one-tap recovery is never answered with "come back tomorrow". - Continuations carrying
tool_results. Blocking a confirmation mid-flight would strand an invocation the user has already answered.
Conversations per month
The plan's meter — or your organization's, if you are on a pooled allowance.
A portfolio can share one monthly number across every app instead of carrying a cap each. That is the right shape for a studio with a long tail of quiet apps: under per-app caps you buy a hundred allowances to use two, and the ninety-eight quiet apps become a reason not to install the copilot in them at all. On a pool there is no per-app cap; the whole portfolio stops together, and Portfolio in the console shows where the allowance is going. Ask us to put your org on one.
The rest of this section is the same either way. A conversation becomes billable at its first user message and is free for the rest of its life, so this counts new conversations, not messages.
The cap stops new work; it never strands work in flight. A thread the user is halfway through finishes normally even if the month rolled over its limit while they were typing.
If hard_cap is off for your app, the API keeps serving past the limit and the overage appears on your invoice instead.
You are emailed at 80%, 95% and 100% — every member of the organization, once per threshold per month.
Sandbox conversations per month
Sandbox traffic is never billed, which means the monthly conversation cap never looks at it and nothing else bounded the month. A test key ships in every debug build; 400 messages a day, every day, is a real bill arriving on a plan whose price assumed it could not.
The refusal happens before a conversation row is written, so a script hammering a leaked test key does not accumulate rows either. Live traffic is unaffected, and the message says so — it is the one gate whose wording is written for a developer rather than for an end user.
Daily model budget
Optional, off by default, and set by us on request. It stops one app's turns once that app's own token spend for the day passes an amount you name.
The platform also has a global daily ceiling behind all of this. It is the guard for the case no per-app limit describes — a retry storm across many apps at once — and if it ever fires you will hear about it from us before you notice it.
Payload ceilings
| Thing | Limit |
|---|---|
context.app_state | 8 KB serialized |
| Action result payload | 16 KB |
| Manifest actions | 100 |
Action timeoutMs | 120 s |
tool_results per continuation | 10 |
Feedback reason | 500 characters |
Raising a limit
Email support@neonapps.co. Per-app limits are a row in a table; changing one takes a minute and applies to the next request. There is no self-serve upgrade during the pilot, on purpose — we would rather have the conversation.