Vantage Launches Token Allocation Support for LiteLLM

Allocate AI provider costs by team, user, customer, or application using LiteLLM request metadata, with totals that stay reconciled to the provider bill.

Vantage Launches Token Allocation Support for LiteLLM
Author:Vantage Team
Vantage Team

Today, Vantage is launching a native LiteLLM integration for Token Cost Allocation, joining per-request logs with native provider costs in Vantage to help organizations measure, allocate, and control AI spend. Organizations using LiteLLM can enrich costs from supported AI provider integrations with request-level metadata identifying the teams, users, customers, and applications driving that spend, giving engineering and finance teams the most granular view of where costs originate.

Cost Report with AI vendors such as OpenAI and Anthropic with ingested metadata tags for App and Environment
Cost Report with AI vendors such as OpenAI and Anthropic with ingested metadata tags for App and Environment

Vantage recently published the Token Cost Allocation Specification, allowing customers to enrich their native provider costs with request-level metadata through log telemetry in a common schema. The spec defines how to attribute token usage proportionally, keep totals aligned with what vendors bill, and carry those dimensions into reporting. Any gateway or logging pipeline that can write that common format to S3 can use it. For customers on gateways like LiteLLM, that used to mean they would have needed to create their own ETL in order to transform their log data and store in S3 in order to send to Vantage to use enriched LLM cost data.

Now, with the LiteLLM integration for Token Cost Allocation, customers deploy a collector into their LLM proxy, which sends data back to Vantage automatically and allocates native provider costs to the teams, users, customers, applications, and workloads that generated them. The Vantage collector runs alongside your LiteLLM proxy, receiving token usage and metadata from the Vantage callback, batching it locally, and securely uploading it to Vantage for cost allocation. It does not collect prompts or completions, and collector interruptions do not block AI requests. Once received, Vantage joins LiteLLM usage telemetry with costs from supported native provider integrations, such as OpenAI, Anthropic, or SpaceXAI, exposing LiteLLM’s metadata.tags and metadata.spend_logs_metadata as tags as if they were directly from the provider, as well as creating Managed AI Tags. Customers can group and filter Cost Reports by any of these dimensions, create Virtual Tags, establish budgets and alerts, and perform showback or chargeback using the request attributes already present in LiteLLM.

LiteLLM enrichment is generally available to Enterprise customers with contracted Build Hours. Enrichment through Token Allocation consumes those Build Hours. To get started, select LiteLLM under the LLM Enrichment section of the Integrations page, and follow the steps to set up the collector on your LiteLLM proxy. See the LiteLLM enrichment documentation for detailed setup instructions.

Frequently Asked Questions

1. What is being launched today?

Vantage is launching a native LiteLLM integration for Token Cost Allocation. The integration uses request telemetry from the Vantage LiteLLM Collector to send per request event data back to Vantage to enrich and allocate costs imported through supported native AI provider integrations.

2. Who is the customer?

The integration is generally available to Vantage Enterprise customers with contracted Build Hours who also route AI requests through LiteLLM and have connected the corresponding AI providers to Vantage. It is particularly useful for organizations that share provider accounts or API keys across teams, applications, or customers and therefore cannot determine cost ownership from provider billing data alone.

If eligible, you can add a LiteLLM source under LLM Enrichment on the Integrations page and follow the setup documentation to configure the callback, collector, and provider cost integrations. If you are not currently eligible, you can reach out to support@vantage.sh.

3. How much does this cost?

Enrichment through Token Allocation consumes your contracted Build Hours in order to add tags to your provider cost data. The amount of build hours consumed depends both on log records ingested and cost rows per month.

As a baseline, an average for a customer with 30M cost rows per month and 5M log records per month would estimate to consume up to ~40 Build Hours per month with daily refreshes.

4. Which providers are supported?

At launch, supported providers are OpenAI, Anthropic, AWS Bedrock, Google Cloud (Vertex AI Gemini and Marketplace Claude), Azure, SpaceXAI, and Baseten. Azure support covers direct Azure integrations; Azure CSP billing accounts are not supported.

Customers must have an active Vantage cost integration for each provider whose costs they want to enrich and select exactly one cost integration per provider. LiteLLM usage records do not include a provider account identifier, so Vantage cannot distinguish multiple accounts for the same provider.

5. How do I set up the LiteLLM integration?

Connect the supported AI providers whose costs you want to allocate. On the Vantage Integrations page, add LiteLLM under LLM Enrichment and select Connect Source. Copy the installation token from the Collector Setup screen.

Install the Vantage callback in the LiteLLM proxy’s Python environment and register it in the proxy configuration. Deploy the Vantage collector as a sidecar or adjacent container, with a shared Unix socket and persistent local storage for its event buffer. On a VM, the collector can also run as a system service. Run the callback package and collector on the same version and upgrade them together.

Then, provide both the installation token and a separate Vantage API token with integration permissions through your platform’s secret management. The collector needs outbound HTTPS access to Vantage and the presigned Amazon S3 upload URLs it receives. Customers do not need to create an S3 bucket or grant an AWS role.

Restart or redeploy LiteLLM to load the callback. Then, on the Collector Setup screen, select Configure Providers and choose a single connected cost integration per provider. Only a single integration can be used, as LiteLLM logs do not contain information about billing account or organization in order to differentiate between different accounts. Integrations connected later are not enabled automatically; add them through Manage Providers. Follow the deployment instructions for platform-specific configuration.

6. What role do I need in Vantage to set up the integration?

A user must have the Organization Owner or Integration Owner role in Vantage to create and configure the LiteLLM integration. These roles provide the permissions necessary to manage integrations and obtain the installation token used by the collector.

Other users can view the resulting allocated costs according to their existing workspace and Cost Report permissions.

7. What permissions do I need in LiteLLM to set up the integration?

The integration does not require a specific LiteLLM user role or access to the LiteLLM administrative interface. The person performing the setup must have access to the LiteLLM deployment environment so they can modify the proxy configuration, add the Vantage callback, supply the installation token and Vantage API token through secret management, deploy the collector, configure its shared socket and persistent storage, and restart or redeploy the proxy.

In many organizations, these permissions are held by the platform, infrastructure, or AI gateway team rather than an individual LiteLLM user.

8. How does Token Cost Allocation work after setup?

The Vantage callback sends LiteLLM usage events to the collector running alongside the LiteLLM deployment. The collector groups compatible events, compresses them, and uploads them to Vantage.

Vantage normalizes the provider, model, token, and metadata fields, matches the telemetry with eligible cost rows from the corresponding native provider integration, and allocates those costs according to the token consumption associated with each tag combination.

For more detailed information on Token Cost Allocation, see our LLM Enrichment documentation.

9. Is LiteLLM the source of cost data?

No. LiteLLM is the source of usage telemetry and allocation metadata. Costs and provider-reported usage come from the corresponding native provider integration in Vantage and remain the financial source of truth.

This differs from treating LiteLLM’s calculated spend as an independent cost import. The integration uses LiteLLM to explain and allocate provider-billed costs rather than replacing those costs with an estimate.

10. Will enrichment change my provider cost totals?

No. Vantage divides matching provider cost rows according to the token share associated with each LiteLLM tag combination. Allocated portions and unallocated remainders sum to the original provider cost imported into Vantage. Cost rows with no matching usage remain unsplit.

The resulting amounts are calculated allocations of provider costs, not charges the provider directly itemized for each team or user. Attribution depends on the coverage and matching of your telemetry; preserving cost totals does not guarantee that every request has been attributed.

Detailed examples can be found in our LLM Enrichment Cost Row Splitting documentation.

11. What happens if LiteLLM and the provider report different token totals?

Provider cost and usage data remain the source of truth.

If LiteLLM reports fewer tokens than the provider, Vantage allocates costs covered by the LiteLLM telemetry and creates a leftover cost row for the tokens the telemetry did not cover. The leftover row retains the model tag but does not receive LiteLLM allocation tags. This can happen when some requests do not pass through LiteLLM or telemetry is interrupted.

If LiteLLM reports more tokens than the provider, Vantage allocates the entire provider cost proportionally across the LiteLLM telemetry. There is no leftover unallocated cost.

In both cases, the resulting rows always sum to the original provider cost. Token Cost Allocation reattributes existing costs; it does not create, remove, or reprice them.

12. Which LiteLLM tags can be used for allocation?

The callback combines supported fields from LiteLLM’s metadata.tags and metadata.spend_logs_metadata. When both contain the same key, metadata.tags takes precedence. Custom dimensions can include customer, application, environment, feature, or purpose. LiteLLM’s organization, team, project, user, end-user, and key attributes are captured automatically when present.

Each record supports up to 32 tags, with keys up to 128 characters and string values up to 256 characters. Values must be strings, booleans, integers, or floats; null, nested, and array values are dropped. Request identifiers are not turned into tags, and Vantage may exclude noisy or sensitive keys.

Custom metadata keys appear without a provider prefix, allowing consistent keys such as team to be used across providers. Vantage also adds managed tags such as vntg:ai:model, vntg:ai:model_provider when recognized, and vntg:ai:token_type. Existing provider tags are preserved; when a usage-slice tag has the same key, its value takes precedence.

13. Can costs be allocated to an individual user?

Yes, when a stable user identifier is available in the LiteLLM telemetry. The callback automatically captures LiteLLM user and end-user attributes when present; customers can also attach a custom identifier through supported metadata. These tags can be used to group and filter Cost Reports, and in Virtual Tags and budgets.

The collector groups compatible events rather than preserving a separate cost row for every request. Enrichment tags follow Vantage’s cost allocation rules: an enrichment tag can belong to only one allocation chain.

14. How are cached tokens handled?

Vantage distinguishes input, output, cache-read, and cache-write tokens when matching usage to provider costs. Enriched rows identify the token kind through the vntg:ai:token_type tag. See the LiteLLM enrichment tag reference for details.

15. Can I connect multiple LiteLLM deployments or replicas?

Yes. Multiple deployments can reuse the same installation token when they process distinct request streams. Multiple LiteLLM replicas behind a load balancer can also share an installation token because each collector maintains its own identity and upload sequence.

Each collector needs its own stable collector identifier and persistent spool; a spool must not be shared between collectors. Customers should not send overlapping telemetry for the same requests through multiple enrichment sources.

16. What happens if the collector becomes unavailable?

LiteLLM continues processing model requests if the collector cannot receive or upload telemetry. The callback delivers events in the background so a collector interruption does not prevent applications from reaching their model providers.

The collector stores pending telemetry locally and retries delivery when service is restored. If the local storage limit is reached or events cannot be recorded, some telemetry may be lost. The associated provider costs will remain in Vantage but may appear partially or fully unallocated.

17. How can I see the health of the LiteLLM collector?

The Collector Setup screen shows a Data Received indicator: Waiting, Received, or Unavailable. The Connected Collectors table shows import status separately, so customers can distinguish receiving telemetry from successfully enriching costs.

Open View import history from the collector’s menu to inspect enrichment runs. Tokens Kept shows how much logged token usage matched provider costs, Log Lines Kept shows how many log lines were read and indexed, and Bill Match shows how many provider cost rows received matching usage. Token Metrics provides disposition details and sample records to help investigate unmatched usage.

For infrastructure monitoring, the collector also exposes a local readiness endpoint and Prometheus metrics. See the documentation for endpoint configuration and troubleshooting.

18. How quickly does telemetry appear in Vantage?

The collector batches and uploads usage, by default at least every 10 minutes. Display of the enriched cost depends on when the provider sends cost data for the time period measured and follows that provider’s refresh cadence.

Recent days are reprocessed within a rolling three-day window to pick up late-arriving batches. The integration supports ongoing cost management rather than real-time request observability.

19. Does Vantage collect prompts and completions?

No. The integration processes the usage fields and request metadata needed for Token Cost Allocation, such as provider, model, token quantities, service tier, batch status, and customer-defined tags. Prompt and completion content and credentials are not collected, stored, or written to a Vantage-owned artifact. This data is not used to train models.

20. Can I backfill historical LiteLLM telemetry?

No. The LiteLLM integration begins collecting telemetry after the callback and collector are configured. It cannot backfill request telemetry from before the integration was installed.

Historical provider costs remain available in Vantage, but costs from before LiteLLM telemetry collection began cannot be retroactively allocated using this integration.

If you are updating to the native LiteLLM integration from the Custom LLM Enrichment source, pause the Custom LLM Enrichment integration and then enable the LiteLLM integration in order to retain previous allocation history.

21. What happens if I stop enrichment?

Stopping a collector in Vantage pauses enrichment of new cost data and preserves existing enriched history. It does not roll back or re-import previously enriched costs. The collector remains visible and can be resumed.

Stopping enrichment in Vantage does not stop the collector process from sending usage. To stop uploads, also stop the collector process. Disabling an individual cost integration in Manage Providers likewise preserves its enriched history.

22. Can I use LiteLLM alongside another Token Cost Allocation enrichment source?

Yes. A cost integration can be enriched by more than one source at once. For example, it can be enriched by a LiteLLM collector and a Custom LLM Enrichment source, or by two LiteLLM collectors. Enrichment only splits existing cost rows, so totals always match the original cost no matter how many sources are connected.

Within a single LiteLLM collector, you can enable only one cost integration per provider. This is because LiteLLM usage doesn't include the provider account a request came from.

To move a provider from one source to another, enable it in the new source's Manage Providers screen, then disable it in the old one. Disabling a provider stops new enrichment from that source but keeps the enriched history already in place.

Avoid sending the same requests through more than one source, because the duplicate usage counts twice when Vantage splits costs.

23. How long does Vantage retain LiteLLM telemetry?

LiteLLM telemetry follows the same retention policy as your other data in Vantage. No separate retention period applies to the LiteLLM integration.

Sign up for a free trial.

Get started with tracking your cloud costs.

Sign up

TakeCtrlof YourCloud Costs

You've probably burned $0 just thinking about it. Time to act.