Cost Management

Know where your AI spend
comes from

Track tokens, cost, and latency for model calls made through Omni, then break usage down by team, person, model, workload, or individual request.

Analyze usage by date, team, and model

Compare token consumption and cost over time, then open the request records behind any change.

Omni Omni

Usage analytics

Total tokens
3.4M
2.6M in · 796K out
Model spend
$22
Across selected usage
Model calls
2,887
Requests in selected range
Average cost
$0.01
Per model call

Token consumption by date

Updates with the filters above

Input Output
075K149K224K298KMar 1: 141K input, 42K outputMar 1Mar 2: 147K input, 44K outputMar 3: 154K input, 46K outputMar 4: 160K input, 49K outputMar 4Mar 5: 167K input, 51K outputMar 6: 174K input, 53K outputMar 7: 180K input, 56K outputMar 7Mar 8: 187K input, 58K outputMar 9: 193K input, 60K outputMar 10: 200K input, 63K outputMar 10Mar 11: 206K input, 65K outputMar 12: 213K input, 67K outputMar 13: 220K input, 70K outputMar 13Mar 14: 226K input, 72K outputMar 14

Usage by workload

Share of tokens in the selected view

Support summaries 847K
Company research 833K
Data analysis 846K
Task automation 837K

Recent model calls

Request-level audit trail
Alex Morgan
Operations · Data analysis
$0.46
GPT-4.1 42K tokens 1467ms
Alex Morgan
Operations · Company research
$0.50
Claude Sonnet 36K tokens 1244ms
Maya Patel
Sales · Company research
$0.04
Gemini Flash 34K tokens 1286ms
Maya Patel
Sales · Support summaries
$0.01
Llama 3.1 29K tokens 1063ms
Devon Lee
Engineering · Support summaries
$0.66
Claude Sonnet 47K tokens 1105ms

Understand usage at every level

Start with company-wide totals, then break usage down by team, model, workload, person, or individual request.

One ledger for Omni model calls

Commercial APIs and self-hosted models appear in the same view with input tokens, output tokens, model, latency, and cost.

Usage tied to your organization

Break usage down by function, department, team, person, model, or workload instead of working backward from a provider invoice.

Request-level audit records

Review the model, token counts, latency, cost, user, and workload behind an individual request, then export the records you need.

Comparable usage data

Compare models and workloads using the same usage dimensions so teams can make model decisions with their own operating data.

Turn visibility into better decisions

Use detailed usage and cost data to understand where AI spend is going and where a different model may make sense.

1

Compare model usage

See how token consumption, latency, and cost differ across the models handling each kind of workload.

2

Find costly workloads

Identify the teams and recurring tasks responsible for the largest share of tokens and model spend.

3

Review individual requests

Move from an aggregate trend to the request records behind it when finance, security, or engineering needs an explanation.

Start tracking model usage with Omni

Join early access for Omni Cloud, or run the open-source deployment in your own environment.