Dispatches
Trust · Athena Governance System · Hard Budgets Per Run

Control agent spend with hard budgets per run

Cap every workflow in dollars or model calls. The runtime stops at the limit, tracks cost by workflow and user, and fails closed if the ledger is unreachable.

Two smoke runs cost $200 and $101. Then we built budgets.

Owner dashboard · cost and governanceAll AOPsLast daysAll userscontract-review AOP run$18.40Budget cap$25.00Rows processed412/500contra…audit…invoic…legal…compli…2 stale procedures flagged
Fig. 01Run pauses at $25 cap; dashboard shows cost by AOP, user, project, and flags stale procedures
  • Hard budgets per run
  • Atomic across sub-agents
  • Fail-closed ledger
  • Cost by workflow, agent, user, project
  • Stale and duplicate detection
  • Every middleware observable
The problem

Autonomy without limits is a liability

Agents that fan out do real work and spend real money. A spreadsheet formula dragged down 500 rows is 500 runs. A sub-agent per audit sample is 125 concurrent runs. Without a hard cap, "the agent got stuck in a loop" is a finance conversation.

A monthly invoice rarely shows per-run detail.

Autonomy with a limit is a tool. Without one, it is a liability.

Invoice onlyOwner dashboardAgent usage$2,140Details unavailableNo per-workflow costMonthly invoice and a shrugCost per AOP shownRun paused at capApprove to continueAutonomy with a limit is a tool
Fig. 02Before: agent usage invoice with no detail. After: per-AOP cost and runs that pause at cap
How it works

Set. Enforce. Watch. Prune.

  1. Set. Per run: max_cost_usd, max_model_calls, or both. Per user, per project, per Skill (reusable component) on the dashboard. Set it in the AOP editor or in one field on the API call.

  2. Enforce. The runtime reserves cost before each model call in a shared transaction ledger—a real-time spend tracker that concurrent sub-agents read and update atomically so fan-out cannot overspend. If the ledger is unreachable, the run fails closed.

  3. Watch. Cost by workflow, agent, user, and project. Run steps, exports, and a per-run summary written by the run itself.

  4. Prune. Stale and duplicate procedures flagged. Role-based visibility into who can see and run which procedures. Approval and cost estimation required for high-volume sub-agent fan-out.

01Budget onrunmax_cost_usd,max_model_calls02AtomicledgerReserves costbefore each …03Next callwithin cap?Concurrentsub-agents …a person approves04Model callor pauseIf cap reached05Fail closedIf ledgerunreachableOwner dashboard shows cost by workflow, agent, user, project
Fig. 03Runtime reserves cost atomically; if cap is reached, run pauses for approval; fail closed if ledger is down
In practice

Budgets and observability in production

A cap on the spreadsheet fan-out.

=AOP dragged down 500 rows with a $25 budget. At $25 the run pauses and asks. Usually it finishes under.

=AOP() dragged down 500rows with $25 budgetRow 1 contract-reviewcompleteRow 2 contract-reviewcompleteRow 3 contract-reviewcompleteRow 412 contract-reviewcompleteBudget cap reachedPause at $25.00×500Completed · $18.40of $25.00 budget
Fig. 04Spreadsheet formula fan-out with $25 cap; run paused at limit, finished under budget at $18.40

Cost per user, per practice group.

A team asks whether an internal knowledge bot is worth it. The dashboard answers by user and by group.

Knowledge bot cost by user and groupPractice groupMonthTotal monthly spendUsage by groupLit.CorpIPTaxUsage trend shows most value
Fig. 05Cost by practice group answers whether the knowledge bot is worth its monthly spend

Every governance layer signals when it runs, enforced by the build.

Each governance layer emits one production signal every time it runs. A test fails the build if any middleware goes silent, so a "silent no-op" regression cannot ship.

CI check: middleware signalsTrigger · Build · all agent middleware requiredBudget ledger signal emittedCost tracking signal emittedApproval gate signal emittedStale detection signal emittedAll 75 signals presentBuild passes · no silentmiddleware can shipmiddleware-signals · 75/75
Fig. 06The build fails if any middleware stops emitting its production signal

Stale and duplicate detection.

Procedures nobody has run in 60 days, or two that do the same thing, surface for review. Retire or merge.

Stale and duplicate detectionTrigger · Procedures flagged for reviewinvoice-check · last run over 60 days agolegal-research · not run in 60 daysaudit-sample / audit-sample-v2 · duplicate paircontract-review · active2 stale procedures · 1duplicate pair · retire ormergeOwner dashboard · governance
Fig. 07Dashboard flags procedures not run in 60 days and duplicate pairs for review, retire, or merge

Approval before the fan-out.

High-volume sub-agent runs show a cost estimate and require approval before they start.

Athena wants toRun 100+ sub-agents for audit samplefan-outEstimated cost: $14–$22100+ concurrent sub-agentsApproval required before high-volume fan-outWAITING FOR A PERSONApprove to start · run will pause at capCost estimation and approval on high-volume sub-agent runs
Fig. 08High-volume sub-agent fan-out shows cost estimate and requires approval before starting
What customers say
“It's good to understand the capability of the platform. That how far, and how quickly you guys can get these things going [...] that has been impressive overall.”
VP-level AI leader · A Fortune 500 retailer
“Lets us get a lot more out of our data than [...] We can sit on top of all of the data in [our data warehouse] as opposed to just, like, our little ingested, you know, three gigabytes semantic models. I think we have a lot of power here.”
Analytics lead · A Fortune 500 retailer
“Gave Athena the spreadsheet, explained which column I was trying to figure out, and then ten secs it told me... exactly how the calculation was based on the other data.”
Analyst · A global manufacturer
“He wrote the entire stuff. By itself. And I was there, and I was thinking, alright. If I had to write all this stuff, I mean, I would've spent days, days, days.”
Manufacturing analyst · A global manufacturer
“I send that to the programmer and he said, alright. I read your file. I copied the section of the code because he has the right interface, and it works. Wow. Just like magic.”
Manufacturing analyst · A global manufacturer
“One workspace. All different, let's say, avatars of their output data. Right? Suddenly, it's dashboard. Excel is here. PowerPoint is when it's all together. It's very impressive.”
Global Sustainability lead · A global manufacturer
“Use the Athena platform probably closer to the concept and how it was designed, like, with the spaces architecture where you've got a set of primitives and capabilities, and you bring together the right set of capabilities to solve a task.”
Audit lead · A global professional services firm
“Digital worker, it gets spun up, it gets given access to the toolkits that it needs. It gets given access to the data that it needs. Just exists for the period it needs to exist to perform the task. And then it gets torn back down again afterwards.”
Audit lead · A global professional services firm
“I like that direction, that concept of the agent is the agent. It becomes an entity in its own right with its own ownership and its own missioning, and you get rid of the gray area and the blurry lines.”
Audit lead · A global professional services firm
“It was PowerPoint... there was no kind of distinction between the two. So that was killer. And I literally had a meeting on it today, and people said how great it was.”
Sales analyst · A global CPG provider

Verbatim from customer calls. Customers anonymized.

Stated plainly

What budgets and governance do

  • Hard budgets per run; shared ledger across sub-agents; fail-closed.

  • Dashboards by workflow, agent, user, project; exports.

  • Approval and estimation on high-volume fan-out.

  • Observability enforced in CI.

What to plan for

A budget stops spend. It does not finish the job. Set it with headroom, and read the run summary when it pauses.

FAQ

How do we control AI spending?

Set hard budgets per run, workflow, and workspace; runs stop when a budget is reached.

Can we see where usage goes?

Yes, by user, team, and workflow.

Who sets the rules?

Your admins control tools, models, connectors, and approvals.

The rest of the platform

Related products and stories

Pick the workflow you were afraid to let loose.

We will set a cap and let it run.