Billable routes
Trace each chat, speech, image, or agent route from HTTP request to provider invocation. Identify whether the key belongs to the server or the end user.
A server-held AI key can be charged every time a public chat or text-to-speech route calls the provider. Bound request size and model output, set per-user or session quotas, limit concurrency, enforce a budget before each provider call, and log usage without recording private prompt text. Authentication is appropriate when the feature is for account holders; a deliberately public demo still needs abuse controls.
This guide covers check 06: Performance and scalability in the Zenveus Production Readiness Standard.
For builders
Your server pays for AI calls even when a visitor makes them repeatedly. A limit displayed in the browser cannot stop a direct API request. Enforce shared quotas and request limits before spending provider credits.
A scoped repair request
Use this prompt in Lovable, Cursor, Replit, or Claude Code with the relevant server files available.
Audit my v0 AI route before the provider call. Add input and output bounds, atomic shared per-user quotas, concurrency limits, a global budget stop, and a timeout without relying on browser counters. Work on a branch with synthetic data and mocked external services. Show the smallest diff, identify required adapters and deployment settings, and add allowed and denied tests that prove side effects cannot happen before checks pass. Do not disable security checks to make a test pass.The same failure may appear as
Find the failure layer
Each check removes a class of causes. Keep the first failing result, its timestamp, and the production log beside it.
Trace each chat, speech, image, or agent route from HTTP request to provider invocation. Identify whether the key belongs to the server or the end user.
Count messages and bytes, bound text length and model output, and check who may choose model, voice, or other expensive settings.
Verify identity or anonymous-session handling, request rate, concurrent jobs, daily spend, and a provider-side budget. A response limit applied after generation does not cap the call already made.
Over-limit requests should return a clear 400, 413, or 429 and never reach the provider. Use mocks so verification incurs no billable requests.
Ranked diagnosis
The endpoint invokes a provider without an effective caller or session policy.
The route forwards arbitrary message counts or text lengths and allows costly output.
Each request is valid alone, but aggregate requests exceed the product's budget.
No per-user or global usage signal reveals a cost spike early.
Step-by-step repair
Edit app/api/chat/route.ts or app/api/elevenlabs/speak/route.ts before the provider SDK call. Put shared quota state in a database or Redis rather than an in-memory counter.
Production safety ruleNever disable access controls, expose service keys, or add wildcard CORS as a routine shortcut.
Choose a fixed model and reject unsupported caller settings. Enforce request byte size at ingress, message count, text length, and provider output cap. Example starting limits are 8 messages, 4,000 characters per message, and 500 output tokens; tune them to your product and provider.
In one shared transaction, check caller rate, active jobs, and remaining budget, then reserve the maximum expected cost. Reserve global budget in the same transaction. Reject excess requests before the provider call; never check then decrement in separate operations.
Record actual usage, release unused reserved cost, and decrement the active-job count in finally. Use expiring reservations and recovery for worker crashes. When billing is uncertain, retain the reservation until reconciled.
Use your own budget gate before every billable route. Provider alerts can arrive after spending and are not always hard caps. Fail closed when the quota store is unavailable, and add provider-side controls where supported.
Implementation example
function validMessages(value: unknown): boolean {
return Array.isArray(value) && value.length > 0 && value.length <= 8 &&
value.every(m => m && ['user', 'assistant'].includes(m.role) &&
typeof m.content === 'string' && m.content.length <= 4000);
}
// At the start of your authenticated POST handler:
// 1. Reject invalid input with 400.
// 2. Atomically reserve user + global budget and a job slot.
// 3. If reservation fails, return 429 (quota) or 503 (store outage).
// 4. Call only the server-selected model with an output cap.
// 5. Settle actual usage and release the slot in finally.This input guard is executable; the numbered integration steps require your shared quota implementation and provider SDK. Character limits do not replace token limits or the model-specific output parameter.
Prove the repair
Run these checks with synthetic data in your test environment, then repeat the relevant acceptance checks after deployment.
When the built-in AI fix makes it worse
Pause generated changes, restore a known working branch, and capture one failing request with its logs. Change one layer and rerun the allowed and denied checks before proceeding.
Engineering handoff
We trace one production request through the complete path, isolate the failing boundary, and leave behind evidence your team can repeat.
Caller policy and input/output limits before each billable invocation.
Atomic reservation, concurrent calls, retries, and budget exhaustion.
Usage alerts, account-level limits, and the global stop mechanism.
Mock provider call counts for oversized and over-quota requests.
Typical repair pattern
Mock-provider tests cover normal use, oversized input, repeated requests, and concurrent calls before release.
Before the next release
Clear answers
It keeps the key out of browser code, but the public route can still spend it unless the route limits use.
Yes, if it has suitable session/IP abuse controls, small input and output bounds, concurrency limits, and a spend ceiling.
It is a useful backstop. App-level quotas provide faster, user-specific protection before a provider call.
Run the verification checks on this page against your test environment, then repeat the relevant checks after deployment. Example code needs your app's authentication, data model, and configuration; reading the guide alone does not verify your deployment.
Official documentation and library references
Free next step
The free tool helps you inspect this symptom. Its result does not establish whether the whole app is production ready.
The Verdict
We can see the symptom from here. What we cannot tell you from outside is whether it is contained or structural. A scanner collects evidence. A named senior engineer makes the decision. For $299, a named senior engineer reads your code and signs a written Verdict against the nine checks in the Zenveus Production Readiness Standard. The 48-hour clock begins when the required access and context are available. If the report does not give your developer a list they can act on, you do not pay.
The 48-hour clock starts when the required access and context are available.
Free decision aid
Get a senior view of the constraint, the evidence you have, and the next decision that removes the most risk.
No email required for this decision aid. Dismiss once and this popup stays closed for the session.