Private inference cost: an enterprise budgeting method
Budget implementation, capacity and operation, then calculate cost per accepted task with an explicit worked example.
Choose a business unit before requesting prices
Private inference cost depends on workload and service scope. A monthly server price cannot compare two proposals when one includes support and the other requires the customer to operate everything. A price per million tokens also does not establish the cost of a usable business result. Start with a unit such as an accepted document classification, reviewed case summary or completed request.
Define when a task counts as accepted. Record correction time if human review is needed. Count all executions when a model retries. Do not include discarded outputs in the denominator of useful tasks. Otherwise a system that looks inexpensive per token can shift substantial work onto the team that fixes its answers. Acceptance criteria belong in the financial model as well as the quality test.
State the reference month, service hours and expected volume. Explain any difference between working-day demand and infrastructure reserved around the clock. The capacity planning guide translates that workload into concurrency and peak requirements. A budget should identify the capacity it assumes rather than treating infrastructure as an unlimited flat fee.
Minimum total-cost breakdown
Include a cost even when another department pays it. Existing hardware has opportunity cost and requires maintenance; internal engineers still spend time on the service. You can distinguish cash expenditure from economic cost, but neither should disappear from the decision. Finance should be able to replace assumptions without rebuilding the entire model.
| Category | Examples | Budget basis |
|---|---|---|
| Implementation | Design, integration, testing and handover | Initial cost and allocation horizon |
| Infrastructure | Hardware or hosting, storage and networking | Monthly price or explicit depreciation |
| Operation | Patches, incidents, certificates and monitoring | Staff time or support agreement |
| Continuity | Alternative capacity, backups and rehearsals | Reserved resources and test effort |
| Usage | Energy, traffic and metered consumption | Demand and volume sensitivity |
| Quality | Review, corrections and retries | Minutes per task and acceptance rate |
| Exit | Export, migration and retirement | Separate end-of-service estimate |
A formula stakeholders can inspect
For a reference month, calculate total cost as allocated implementation plus infrastructure, operation, continuity, consumption and review. State the implementation allocation horizon, such as twelve or twenty-four months; this is a planning assumption rather than a universal accounting rule. Divide the total by accepted tasks in the same period. Keep each component visible so procurement and finance can substitute measured values.
When comparing an API with a private environment, use the same tasks and acceptance criteria. Include API input and output charges, retries and auxiliary services. Include idle reserved capacity and administration in the private option. Review effort may differ between models, so estimate it per alternative rather than copying an unmeasured constant across the comparison.
Worked example: 20,000 accepted tasks per month
Illustrative assumptions. The following figures are invented to demonstrate the calculation. They are not Primedefence prices or customer results. Replace them with project quotations and measurements. All figures use euros for a consistent example.
Assume €6,000 of implementation allocated over twelve months, €900 per month for infrastructure, €600 for operation and €200 for continuity and other expenditure. Equivalent fixed cost is €2,200 per month. At 20,000 accepted tasks, fixed cost is €0.11 per task. At 5,000 accepted tasks, it rises to €0.44 before human review is added.
Now assume 10% of the 20,000 tasks require two minutes of review at €30 per hour. That is 4,000 minutes, approximately 66.67 hours and €2,000 of review cost. Total monthly cost becomes €4,200, or €0.21 per accepted task. Improving task usefulness may therefore matter more than a small reduction in infrastructure expenditure.
| Scenario | Fixed cost | Assumed review cost | Cost per task |
|---|---|---|---|
| 20,000 accepted tasks | €2,200 | €2,000 | €0.21 |
| 5,000 accepted tasks | €2,200 | €500 | €0.54 |
The lower-volume scenario keeps the same review proportion and duration while fixed cost remains unchanged. Do not extrapolate if model, document length or output quality changes. The example also does not establish that the assumed infrastructure can serve the workload; capacity needs a separate test. Taxes and financing treatment should be handled consistently in an actual comparison.
Calculate break-even without hiding assumptions
In a simplified model, let a private option have fixed monthly cost F and variable cost v per task, while the API costs a per task. The economic threshold is F divided by a minus v, provided a is greater than v. That comparison applies only within the workload the private environment can support at the required quality and response time.
If a is equal to or below v, this model has no positive break-even point in favour of the private configuration. If growth requires another capacity unit, F is no longer constant. Use stepped scenarios instead of extending a straight line indefinitely. Data restrictions or version control can justify an option that is not cheapest, but those requirements should remain explicit rather than being disguised as savings.
What to measure before approving recurring spend
Record input and output size distributions, waiting time, retries, failures and accepted tasks. The NVIDIA NIM observability documentation illustrates operational measurements covering latency, throughput and queues. Those measurements help explain infrastructure demand; the cost model additionally needs business acceptance and review observations.
Build low-demand, expected-demand and peak scenarios. For each, state reserved capacity, acceptable response time and behaviour beyond the limit. Add sensitivity to review effort, which depends on both task and model. The private AI pilot should produce these observations before a long-term commitment is made.
Request comparable proposals
Give every supplier the same workload brief, data conditions and acceptance criteria. Request separate implementation, operations, capacity and expansion prices, with exclusions and customer obligations. Ask what creates additional charges and what happens if the pilot misses its objectives. The vendor assessment guide provides a structure for reviewing those responses.
The budget should be easy to recalculate when users or document sizes change. Primedefence scopes private inference around requirements and evidence. This article provides an economic evaluation method, not a universal tariff or a guaranteed saving. Keep the decision record so that a later expansion can be compared against the assumptions originally approved.
Frequently asked questions
Can we compare token prices alone?
Token prices describe part of consumption, not total cost. Include setup, reserved capacity, operation, retries and human review, and establish comparable quality first.
Do open-weight models remove licensing cost?
Review the actual model and component licences. Even without per-token charges, infrastructure, integration and operating costs remain.
What does unlimited tokens mean in a proposal?
It does not mean unlimited capacity. Check concurrency, context limits, speed, priorities, usage policy and reserved resources to understand the service available.

Written by
Daute DelgadoCEO & Co-founder, Primedefence
Daute Delgado is CEO and co-founder of Primedefence. He spent more than a decade defending airlines, managed SOCs and international organizations, first as an operator and later leading security teams.
View full profileIs private inference right for your business?
Define your use case, data requirements and pilot acceptance criteria.
Related articles

Private inference · Deep-dive
LLM inference capacity planning: memory, load and latency

Private inference · Comparison
Self-hosted LLM vs API: an enterprise decision framework

Private inference · Playbook
Private AI vendor assessment: an enterprise checklist

Private inference · Playbook

