AI spend anomaly detection: baselines and investigation

A spend anomaly is a departure from an expected workload, not a fraud classification. To explain it, separate changes in request count, tokens per request, model mix, and pricing. Then connect the change to jobs and authorization records.

Use comparable baselines

Compare the same principal and workload across equivalent periods. An overnight batch should not be judged against a quiet afternoon. Mark deployments, promotions, scheduled evaluations, and model migrations. For a new account with little history, say that its baseline is immature instead of presenting a precise confidence score.

Keep recorded usage separate from invoice amounts. Cached-token pricing, retries, provider adjustments, and missing cost data can change reconciliation. Fix the accounting unit before interpreting the anomaly.

Decompose the increase

ChangeCheck
More requestsNew customers, retries, repeated jobs, or unfamiliar traffic
More input tokens per requestContext accumulation or larger legitimate documents
More output tokens per requestChanged output limits or task type
A more expensive model mixRouting changes, fallback, or a new permitted model
Different reported cost with similar usagePricing configuration, cached usage, currency, and billing adjustments

A worked calculation

Synthetic example: a job makes 1,000 comparable requests at an average recorded cost of $0.01, totaling $10. The next run makes 2,000 at $0.03, totaling $60. At the earlier average cost, the extra 1,000 requests account for $10. Applying the $0.02 average-cost increase to the new 2,000-request volume accounts for the remaining $40 increase. This decomposition explains the arithmetic, not the cause.

Investigate the model and token mix behind the average before saying prices tripled. Also check whether the apparent request increase is duplicate log delivery. Deduplicating events must not accidentally discard real paid retry attempts.

Turn a baseline into a review policy

Choose thresholds using known workloads, review capacity, and the cost of missed events. Evaluate absolute impact as well as relative change: doubling a tiny test account may matter less than a modest increase on a large production workload. Record why a case crossed the threshold.

Keep the baseline window, minimum data requirement, comparison method, and exclusions with the result. Avoid automatically absorbing a suspected incident into the normal baseline while it is still under investigation.

Report the result honestly

Show total recorded cost, the portion explained by known changes, and the portion still under review. Do not multiply an anomaly score by spend to invent fraud loss. Track confirmed explanations and unresolved cases so later threshold changes can be evaluated against the same evidence.

For an active spike, follow the incident triage sequence. For consumption controls, use the rate-limit and budget guide. For the fields needed, see security logging.