The task stays the same. Why does the bill grow?
Ask an AI to browse a catalog, collect prices and return a table. It appears to move forward with each action. Behind the scenes, each model request may send earlier page observations and screenshots again. A longer history creates more repeated input. Removing old material to save space can also break reuse of cached input.
Asana’s recent browser-agent study illustrates a useful tension: deleting context does not necessarily make a task cheaper, and keeping more does not necessarily make it more expensive. The question is how much needs to be processed again—and whether the agent finishes the job.
This independent analysis separates the comparisons in the study, then proposes a way for teams to evaluate their own costs and outputs. It is not a Onevium performance test or a promise of equivalent savings.
Separate the parts of the 76x headline
Asana’s engineering article appeared on October 8, 2026; OpenAI’s case study followed on October 9. The experiment varied models, history budgets and screenshot-retention policies on a demo-bookstore task.
The headline compares an older model’s original configuration with a newer model’s optimized configuration. Both the workflow and the model changed. It cannot establish a 76x benefit from switching models alone. The report also includes within-model comparisons: about 29x for Model B’s workflow changes, and about 4x for Sol with the same larger history budget.
Before buying a cheaper model, ask what drives your bill: unit prices, repeated input, revisiting pages or failed attempts? A lower unit price can still produce a higher task cost if the agent needs more steps.
Think of caching as a bookmark in an ordered stack
Imagine each request as a stack: working rules, the original task, the first observation, the second observation and the newest material. Caching can reuse an earlier portion that has stayed unchanged.
This original illustration shows the idea, not an actual API payload:
Before: rules → task → page A → page B
After: rules → task → page A → page B → page C
earlier material stays unchanged
Change a screenshot inside page A:
After: rules → task → page A′ → page B → page C
the unchanged prefix ends here
Keeping only the newest screenshot has two effects: a request can shrink, while reusable history can shrink too. You need both quantities to judge the tradeoff.
A cache setting cannot fix a workflow that continually rewrites earlier material. Such a workflow may pay to write cache entries that it rarely reads. Actual rules, lifetimes and prices depend on the service. Do not copy one provider’s settings into every model integration.
Sending less and doing less are different savings
Consider a fictional procurement assistant comparing twelve products. By the ninth product, the first few prices have been trimmed from history, so it opens those pages again. Even if an individual call costs less, revisits add model calls, browser time and more opportunities to fail.
Keeping everything also has a cost. Irrelevant pages, repeated screenshots and wrong turns can accumulate. An agent stuck in a loop is not rescued by an unlimited history budget.
Track three levels separately:
| Level | Question |
|---|---|
| One model call | Which inputs are charged as ordinary input, cache writes or cache reads? |
| One complete task | How many calls, revisits and retries does it take? |
| One acceptable deliverable | How much review, repair and rerunning do people need? |
Total tokens are one measurement. For the business, the more useful question is the cost of obtaining an acceptable result.
Seeing the facts is not delivering the requested answer
The full report measures seeing the required facts, retaining them, producing an answer and calculating a correct total separately. It notes that Sol returned a total and selected summary details rather than the requested full table. Final-answer tuning was outside the study’s scope.
That does not remove the engineering value of better caching. It does mean retrieval success should not be mistaken for delivery success. A procurement team usually needs itemized prices, sources and calculations; a correct total alone may be difficult to audit.
For our fictional twelve-product task, write the acceptance conditions first: no missing products; name, specification, currency, price and source in each row; explicit missing values; no adding different currencies together; a total that can be recalculated from the rows. Only outputs meeting those conditions enter the cost comparison.
These are proposed business acceptance checks, not an independent reproduction of Asana’s experiment.
Change one thing in the next experiment
If you build agent systems, fix the task, inputs, model and answer criteria before comparing history policies. Changing the model, prompt and budget together makes it hard to attribute a result.
A practical experiment plan:
- Choose a repeatable, non-sensitive task with a reference answer and retain the original configuration’s results.
- Record usage per call: ordinary input, cache reads, cache writes where the service reports them, and output. Mark missing values as unknown.
- Change one retention policy. Preserve request order and record where pruning occurs.
- Repeat runs and retain completion status, actual deliverables, elapsed time, model cost and human repairs.
- Inspect failures and the worst run before expanding use.
Keep budgets and stop conditions. Stop for human review when spending exceeds the agreed limit, pages are repeatedly revisited or sources cannot be verified. Do not let a task run indefinitely just to preserve cache reuse.
If you use an AI application without these low-level controls, require itemized evidence and a clear output format, and record rework. Do not search for a cache switch that your product does not expose.
Four questions to ask about a savings claim
The experiment covers one architecture and one task, with few repetitions per condition. The best configuration was selected after comparison; model phases and protocols also differed. The findings motivate testing, rather than establishing a universal business outcome.
When reading a similar case, ask:
- Was the model held constant?
- Did cheaper runs meet the same delivery requirements?
- Is the amount estimated model usage or total cost including infrastructure and people?
- Does the advantage appear across repeated runs, or only in the best attempt?
Average performance helps identify a pattern. Failure frequency and worst-case spending determine whether you can leave a task unattended.
Evaluate the work, as well as the model
The useful shift is from comparing model prices to examining how the entire job runs. Unit prices, history management, repeated work and output quality belong in the same evaluation.
Start with one recurring task. Save its inputs and outputs, define an acceptable result, then compare bills. Teams using Onevium can keep task materials and outputs in a project for review; this article does not claim that Onevium exposes the study’s cache controls. See files and result review.
Continue with measuring AI workflow value, improving evaluations through trace review, and natural-language browser operation.