Reference

Rate limits and caps

Every ceiling in the system, what it counts, and what happens when it is reached.

There are five ceilings. They exist for different reasons and fail in different ways, and knowing which one you hit is usually the whole diagnosis.

CeilingCountsScopeDefaultWhat happens
Messages per minuteChat, identify and feedback requestsPer device, per app20HTTP 429 before the stream opens.
Messages per dayTurns that start workPer app400rate_limited on the stream.
Conversations per monthNew billable conversationsPer appPlanquota_exceeded on the stream.
Sandbox conversations per monthNew conversations on a test keyPer app200quota_exceeded, with a message written for a developer.
Daily model budgetModel spend in dollarsPer appNonequota_exceeded on the stream.

Messages per minute

An in-process token bucket, per (app, device). Serverless instances each keep their own, so this is a soft first line rather than an exact number — it exists to absorb a stuck retry loop, not to meter anything.

GET /v1/apps/config has its own two buckets: 12/minute per device and 600/minute per app. The app-wide one is what still holds when a caller rotates X-Rendel-Device on every request, and what it actually protects is the database connection pool /v1/chat shares.

Messages per day

The durable version of the minute bucket. It counts rows rather than trusting per-instance state, and it applies to sandbox traffic too: a test key is extractable from any debug build you ship, and unmetered means unbounded model spend.

Two things are exempt on purpose:

  • Retries of a turn already counted. Resend the same turn_id and the one-tap recovery is never answered with "come back tomorrow".
  • Continuations carrying tool_results. Blocking a confirmation mid-flight would strand an invocation the user has already answered.

Conversations per month

The plan's meter — or your organization's, if you are on a pooled allowance.

A portfolio can share one monthly number across every app instead of carrying a cap each. That is the right shape for a studio with a long tail of quiet apps: under per-app caps you buy a hundred allowances to use two, and the ninety-eight quiet apps become a reason not to install the copilot in them at all. On a pool there is no per-app cap; the whole portfolio stops together, and Portfolio in the console shows where the allowance is going. Ask us to put your org on one.

The rest of this section is the same either way. A conversation becomes billable at its first user message and is free for the rest of its life, so this counts new conversations, not messages.

The cap stops new work; it never strands work in flight. A thread the user is halfway through finishes normally even if the month rolled over its limit while they were typing.

If hard_cap is off for your app, the API keeps serving past the limit and the overage appears on your invoice instead.

You are emailed at 80%, 95% and 100% — every member of the organization, once per threshold per month.

Sandbox conversations per month

Sandbox traffic is never billed, which means the monthly conversation cap never looks at it and nothing else bounded the month. A test key ships in every debug build; 400 messages a day, every day, is a real bill arriving on a plan whose price assumed it could not.

The refusal happens before a conversation row is written, so a script hammering a leaked test key does not accumulate rows either. Live traffic is unaffected, and the message says so — it is the one gate whose wording is written for a developer rather than for an end user.

Daily model budget

Optional, off by default, and set by us on request. It stops one app's turns once that app's own token spend for the day passes an amount you name.

The platform also has a global daily ceiling behind all of this. It is the guard for the case no per-app limit describes — a retry storm across many apps at once — and if it ever fires you will hear about it from us before you notice it.

Payload ceilings

ThingLimit
context.app_state8 KB serialized
Action result payload16 KB
Manifest actions100
Action timeoutMs120 s
tool_results per continuation10
Feedback reason500 characters

Raising a limit

Email support@neonapps.co. Per-app limits are a row in a table; changing one takes a minute and applies to the next request. There is no self-serve upgrade during the pilot, on purpose — we would rather have the conversation.