Caps act before the model call
Spend is enforced where security is enforced. Every metered interaction crosses the same four functions, and a cap is a decision made at enforcement time.
| LIVE SPEND | Every interaction metered as it happens on the enforcement path, attributed at request level, across providers |
| BUDGETS THAT HOLD | Caps enforce at the same point policy does, inline, before the model call; at the cap, your policy decides what happens next |
| REAL ALLOCATION | Spend attributed to user, team, department, cost center, model, and agent from runtime context, carried on the request itself |
| ONE CONSOLE | Spend and adoption sit beside threat activity, fed by the same runtime decisions |
Common questions
How does the runtime's metering relate to provider invoices?
The runtime meters usage as it happens on the enforcement path, per request, with the user, team, model, and provider attached at enforcement time. Provider invoices remain the bill; they arrive later and aggregate. Runtime metering exists for control and attribution at the moment of use. It enforces caps, allocates spend to the teams that created it, and catches a runaway workload while it is still running.
What happens when a budget hits its cap?
Whatever your policy says. At budget enforcement the runtime supports allow (record the overage), alert (notify owners and keep going), hold for human review (further calls wait for an approval), or block (calls stop until the budget is raised or resets). Caps enforce inline, before the model call, so a cap is a decision made at enforcement time rather than a number discovered at month end.
How granular is attribution? Can we run chargeback?
Request-level. Every metered interaction carries user, team, department, cost center, model, and provider, and agent workloads are attributed the same way. Dashboards and exportable reports support chargeback, showback, finance reviews, and board updates without spreadsheet reconstruction.