To keep AI inference cost below subscription revenue, model cost per active user as requests per day multiplied by cost per request, and compare it with net revenue per user. Cap free usage on the server, as SlimAI does with 3 photo scans per day, route simple requests to cheaper models, cache repeated inputs, and track cost per user daily. Aim for inference at a predictable share of net revenue.
Most software has a marginal cost close to zero. AI features do not. Every request has a price, and that price is paid whether the user is a subscriber or a free-tier visitor who installed the app for a day. This post walks through the arithmetic that keeps that price below what users pay, using SlimAI, an AI calorie tracker Techparser built and has operated since October 2025, as the running example.
The equation you have to win
The unit economics of an AI subscription come down to one comparison per user per month.
Three of those variables are under your control before launch: requests per active day (by capping), cost per request (by model choice, prompt size, and caching), and the split between free and paid users (by conversion design). The fourth, active days, is what you want to be high, so the design has to make heavy use affordable rather than discouraged.
A worked table with illustrative numbers
The figures below are illustrative. They are not SlimAI’s financials and the per-request cost is a placeholder in the range typical for a vision model call with a short structured response in 2026. The structure is the point: fill in your own numbers and the conclusion falls out.
| Line | Free user (capped) | Paid user | Blended per 100 users at 5% conversion |
|---|---|---|---|
| Requests per active day | 3 (hard cap) | 8 (observed average, illustrative) | |
| Active days per month | 12 | 20 | |
| Requests per month | 36 | 160 | 95 x 36 + 5 x 160 = 4,220 |
| Cost per request (illustrative) | $0.004 | $0.004 | |
| Inference cost per month | $0.14 | $0.64 | $16.88 |
| Gross subscription price | $0 | $9.99 | 5 x $9.99 = $49.95 |
| Net after 15% store commission | $0 | $8.49 | $42.46 |
| Inference as share of net revenue | n/a | 7.5% | 39.8% |
Read the last column carefully. At a 5 percent conversion rate, the 95 free users cost $13.30 in inference and the 5 paid users cost $3.20, while only the paid users bring in revenue. Inference is 40 percent of net revenue, which is survivable but tight. Remove the cap and let free users average 8 scans per day and the free tier alone costs $36.48, or 86 percent of net revenue. Raise conversion to 10 percent with the cap in place and inference drops to about 21 percent of net revenue. The cap and the conversion rate are the two levers that matter most.
The 3-scans-per-day free tier, enforced server-side
SlimAI allows free users three photo scans per day. Subscriptions, handled through RevenueCat, unlock unlimited scanning plus features such as voice logging, a body-fat estimate from a photo, and an AI coach. Three was chosen because it covers a typical day of meals, so the free experience is complete rather than crippled, while making the worst-case free cost a known number.
The enforcement lives on the server, not in the app. A client-side counter is a suggestion. A modified APK, a cleared app storage, or a reinstall defeats it in seconds. In SlimAI the check happens where the request is processed: the backend looks up the user’s scan count for the current day and their subscription entitlement before any call to Gemini through firebase_ai is made. If the quota is exhausted and there is no entitlement, the model is never called and the app shows the paywall.
- The quota check and the model call happen in the same trusted place, so there is no window between them.
- The day boundary is computed on the server in a fixed timezone, so a device clock change does not reset the count.
- Entitlement comes from the subscription provider’s server-side webhook state, not from a flag the app sends.
- The counter is written before the model call, not after, so a client that drops the connection mid-request still consumes a scan.
- Failed model calls refund the scan, because charging a user for an error is how you earn a one-star review.
The result is a free tier whose cost is bounded by a number you chose. That is the property that makes the rest of the model work. SlimAI has passed 10,000 installs on Google Play with a 4.6-star rating and ships in seven languages, and the free-tier cost has stayed a predictable line because the cap has never been optional.
Reducing cost per request: routing and caching
Once usage is bounded, the next lever is the cost of each request. Four techniques cover most of the savings available to a subscription app in 2026.
| Technique | How it works | Typical saving | Watch out for |
|---|---|---|---|
| Model routing | Send simple or low-stakes requests to a smaller, cheaper model; reserve the large model for hard cases or paid users | Large when most traffic is simple | Route decisions need their own evaluation set; quality drift on the cheap path |
| Prompt and output trimming | Shorten system prompts, request structured output, cap response tokens, downscale images before upload | Proportional to tokens removed | Over-trimming hurts accuracy; re-run evals after every change |
| Response caching | Hash the input and reuse the result for identical or near-identical requests, such as a barcode or a repeated packaged product | High for repeat inputs, near zero for unique photos | Cache invalidation when the model or prompt changes; privacy of cached content |
| Provider prompt caching | Use the provider’s cached-prefix pricing for the fixed part of long prompts | Moderate for long, stable system prompts | Only pays off above the provider’s minimum cacheable length |
| Batch and off-peak processing | Move non-interactive work such as weekly summaries to batch endpoints or scheduled jobs | Provider batch discounts where offered | Not usable for anything the user is waiting on |
For a photo-first product like SlimAI, caching helps mainly for packaged goods and barcodes, while image downscaling and structured output give the reliable savings. For a text-heavy assistant, routing and prompt caching dominate. Measure before assuming which one applies to you.
Pricing the paid tier against heavy users
Unlimited plans are a promise that a small number of users will test. Model the 95th percentile paid user, not the average. If a heavy user scans 30 times a day for 30 days at the illustrative $0.004 per request, that is $3.60 a month, still well under an $8.49 net subscription. If your per-request cost is ten times higher, the same user costs $36 and the unlimited promise is a loss. Either raise the price, add a generous soft cap, or route that user’s later requests to a cheaper model. Decide before launch, not after the first invoice.
Monitoring: the four numbers to watch daily
- Inference cost per active user per day, split by free and paid. This is the primary health metric.
- Requests per active user per day, with the distribution, not only the mean. A fat tail is your future bill.
- Cost per request by model, so a provider price change or a routing regression is visible the day it happens.
- Conversion rate from free to paid, because it sets how many free users each subscriber is funding.
Put a hard budget alert on the provider account and a daily anomaly check on cost per active user. An abuse pattern, a retry loop after a client bug, or a prompt change that doubled output tokens should page someone within hours, not appear on a monthly statement.
A pre-launch checklist
- Write the per-user cost equation with your own numbers for free and paid tiers and for the 95th percentile user.
- Choose a free-tier cap that gives a complete daily experience and a bounded worst case.
- Enforce the cap and the entitlement check on the server, in the same code path as the model call.
- Downscale images, request structured output, and cap response tokens.
- Put the model behind an abstraction so routing and provider changes are configuration.
- Build the cost-per-active-user dashboard before launch and set a budget alert.
- Re-run your evaluation set after every prompt, model, or routing change.
Frequently asked questions
- How do you calculate AI inference cost per user?
- Multiply requests per active day by cost per request and by active days per month. Do it separately for free and paid users, then blend by conversion rate. Compare the result with net revenue per user, which is the subscription price minus store commission. Model the 95th percentile heavy user as well as the average.
- Why should a free tier be enforced server-side?
- A client-side limit can be bypassed with a modified app, a cleared storage, or a reinstall. Server-side enforcement checks the user’s daily count and subscription entitlement in the same code path that calls the model, so the model is never invoked without authorisation. SlimAI enforces its three-scans-per-day free tier this way.
- What is a good free tier limit for an AI app?
- Pick a limit that covers a complete typical day of use while bounding the worst case. SlimAI chose three photo scans per day because it covers three meals. The right number depends on cost per request and conversion rate; run the per-user equation at your cap and at a 5 to 10 percent conversion rate to check it.
- How can you reduce AI cost per request?
- Route simple requests to smaller models, trim prompts and cap output tokens, downscale images before upload, cache results for repeated inputs such as barcodes, use provider prompt caching for long fixed prompts, and move non-interactive work to batch jobs. Re-run your evaluation set after each change so savings do not come at the cost of accuracy.
- What share of subscription revenue should AI inference cost?
- There is no universal figure, but a healthy target keeps inference at a small, predictable fraction of net revenue, with room for store commission, backend hosting, and support. In the illustrative model in this post, a capped free tier with 5 percent conversion put inference near 40 percent of net revenue; 10 percent conversion brought it near 21 percent.
About the author
Zoraiz Ejaz
Co-founder, Techparser
Zoraiz Ejaz is a co-founder of Techparser and leads its engineering and product practice. He has spent close to a decade designing, building and scaling mobile, web and AI products for startups and enterprise teams across health, fintech, payments, social and education, from first architecture and release pipelines through to launch and years of production support. He writes about how to scope, cost and ship software that lasts.


