Dispatches
Data · Athena Lakehouse: Analytics on your data

Store and query billions of rows at millisecond speed in your storage

Analytics on open Iceberg format in your S3, GCS, or Blob—with catalog, semantic models, dashboards, and an engine built for the questions agents actually ask.

Billion-row tables. Three- to eight-way joins.

Sales transactions · 1.2B rowsRegionDate rangeCategoryMatching rows70,000Engine time1.6 msWestEastSouthNorthCent.Cross-filtering as cursor
Fig. 01Query over a 1.2B-row table returning 70,000 matches with engine time 1.6 ms
  • Open Iceberg format
  • Your object storage
  • 79 ms on 2M rows
  • 1.2B rows at millisecond engine time
  • Per-user warehouse permissions
  • Coexists with Snowflake and Databricks
The Problem

You pay for compute you cannot use

The warehouse and the BI layer monetize compute, so they throttle it. Dashboards refresh in hours. Capacity SKUs get bought and then rationed. Analysts wait, and agents, which ask a hundred questions where a human asks one, wait more.

And the data is in a proprietary format, so leaving is a migration.

Open format. Your storage. An engine you are allowed to floor.

BI admin consoleLakehouse dashboardCapacity throttled31% of days this monthRefresh ~4 hHours to refresh, throttled computeLast refresh 4 sCross-filter 0.4 sSub-second cross-filter, no throttling
Fig. 02From throttled hours-long refreshes to sub-second cross-filter on Lakehouse
How It Works

Land. Model. Query. Own.

  1. Land the data. Migrate or replicate tables to Iceberg in your bucket using connectors or agents. Or point the Lakehouse at existing Snowflake, Databricks, or BigQuery and query through it with each user's own permissions.

  2. Model it. Athena drafts a semantic model from the catalog and context you provide. You review it. Everything downstream queries through it.

  3. Query it. From dashboards, from =ATHENAQUERY in a sheet, from an agent, from a notebook. Millisecond engine time on billion-row tables.

  4. Own it. Iceberg is open. One retailer migrated between cloud warehouses; dashboards queried Iceberg throughout with no changes.

01Snowflake ·DatabricksMigrate orreplicate to …02IcebergtablesYour S3, GCS,or Azure Blob03SemanticmodelAthena drafts,you reviewa person approves04Dashboards ·=ATHENAQUER…Millisecondengine time05PortableOpen IcebergformatOne retailer moved between cloud warehouses and the dashboards did not notice
Fig. 03Migrate data to Iceberg in your storage, model it, query it, own it in open format
People and agents

What people and agents each do

Lakehouse. Governed storage for everything connectors bring in, from Snowflake, Databricks, files, and systems of record, stored where Palladium says it lives.

CapabilityPeople canAgents canTogether
LakehouseConnect sources, set retention and residency.Load, clean, and join data; keep pipelines running.Both read from the same governed copy, with the same permissions.

Anything a person can do here, an agent can do with the same permissions, and both land in the same audit trail.

Use cases

Put the Lakehouse to work

The store-operations live dashboard.

A table with hundreds of millions of rows behind an operations dashboard at a Fortune 500 retailer. Sub-second cross-filter for thousands of concurrent users, included in Athena services.

Store operations · 100M+ rowsRegionStoreWeekConcurrent users1000sCross-filter timeSub-secondMonTueWedThuFriSatSunFilter interaction · sub-second
Fig. 04Sub-second cross-filter for concurrent users on a dashboard over 100M+ rows

Consolidating the Wild West.

Several ungoverned BI capacities, multi-hour refreshes, and frequent throttling. One governed platform, fresher data.

Ungoverned capacitiesOne LakehouseBI capacity 1BI capacity 2BI capacity 3BI capacity 4BI capacity 545-minute to 4-hour refreshes, throttlingDashboard 1Dashboard 2Dashboard 3One governed platform, fresher data
Fig. 05Multiple ungoverned BI capacities consolidated to one governed Lakehouse

Data-science tooling without a second vendor.

Notebooks with isolated compute and per-user volumes on billion-row datasets. Ephemeral Postgres up to 1 TB.

SELECT store_id, SUM(revenue) FROM transactions WHEREstore_idrevenuetxn_countavg_basketregionstore_alphahigh volumemany txnstypicalNortheaststore_betahigh volumemany txnstypicalSoutheaststore_gammamedium volumemoderate txnstypicalMidweststore_deltahigh volumemany txnstypicalWeststore_epsilonmedium volumemoderate txnstypicalSouthwestWHO CHANGED WHATAnalystNotebook on billion-rowdatasetAthenaIsolated compute · 40 GBvolumeAnalystEphemeral Postgres up to 1TB
Fig. 06Notebook querying billion-row tables with isolated compute per user

"Why this number?"

Click a chart. See the exact upstream table, the transformation, and freshness. Finance stops calling the data team.

transactionsstoresproductsONE DEFINITIONWeekly salesSUM(amount) WHERE week= current - 1Fresh as of 06:58current
Fig. 07Click a chart to see the exact upstream table, transformation, and freshness

Digests without dashboards.

Email or Slack digests from the Lakehouse on a schedule: what moved, where, why.

Weekly Lakehouse DigestRevenue moved up week-over-week, driven by Westregion stores.Average basket size declined but transaction countrose significantly.Inventory turnover improved in Electronics and Homecategories.Several stores flagged for margin compression belowthreshold.transactions100M+ rowsinventoryfresh as of 06:58Lakehouse 8:15 · scheduled
Fig. 08Scheduled Lakehouse digest showing what moved, where, and why
What customers say
“This combination is a winning combination for us, and this is where we see maximum traction, which is how do you apply all the intelligence that you have within [our data platform] and not worry about how you render that output.”
Analytics lead · A global CPG provider
“I envision Athena and [our data platform] as the combination... that's the stack for me.”
Analytics lead · A global CPG provider
“It's good to understand the capability of the platform. That how far, and how quickly you guys can get these things going [...] that has been impressive overall.”
VP-level AI leader · A Fortune 500 retailer
“Lets us get a lot more out of our data than [...] We can sit on top of all of the data in [our data warehouse] as opposed to just, like, our little ingested, you know, three gigabytes semantic models. I think we have a lot of power here.”
Analytics lead · A Fortune 500 retailer
“Gave Athena the spreadsheet, explained which column I was trying to figure out, and then ten secs it told me... exactly how the calculation was based on the other data.”
Analyst · A global manufacturer
“He wrote the entire stuff. By itself. And I was there, and I was thinking, alright. If I had to write all this stuff, I mean, I would've spent days, days, days.”
Manufacturing analyst · A global manufacturer
“I send that to the programmer and he said, alright. I read your file. I copied the section of the code because he has the right interface, and it works. Wow. Just like magic.”
Manufacturing analyst · A global manufacturer
“One workspace. All different, let's say, avatars of their output data. Right? Suddenly, it's dashboard. Excel is here. PowerPoint is when it's all together. It's very impressive.”
Global Sustainability lead · A global manufacturer
“Use the Athena platform probably closer to the concept and how it was designed, like, with the spaces architecture where you've got a set of primitives and capabilities, and you bring together the right set of capabilities to solve a task.”
Audit lead · A global professional services firm
“Digital worker, it gets spun up, it gets given access to the toolkits that it needs. It gets given access to the data that it needs. Just exists for the period it needs to exist to perform the task. And then it gets torn back down again afterwards.”
Audit lead · A global professional services firm

Verbatim from customer calls. Customers anonymized.

The bet

Any customer with real spend on BI, warehouse, or lakehouse tools: we want that business. Not to rip it out on day one. To sit on top with your users' own permissions and let the numbers make the argument. Coexistence first. Consolidation as the outcome.

Governance, stated plainly

  • Tables live in your object storage, in open Iceberg format.

  • Every query runs with the user's own warehouse permissions; agents included.

  • Lineage and freshness on every chart.

  • Deploy in your VPC, on-prem, or air-gapped with the rest of the platform.

What to plan for

New Lakehouse figures are engine time, not end to end. We publish both when we have both, and we do not round up.

FAQ

Why does this matter to the business?

It makes large amounts of data fast to use for people and agents, without locking it into a vendor.

Where does our data live?

In your own storage, in open formats.

Do we have to migrate?

No. Start on top of your existing warehouse.

One Platform Underneath

AGS · Athena Governance System

Records every change by a person or an agent, rolls back one contributor's edits without losing anyone else's, and keeps the model, instructions, and sources behind each agent action.

Palladium · deployment

How the platform is deployed: Athena's managed cloud, your cloud on AWS, GCP, or Azure, on-prem, air-gapped, or GovCloud. Same platform in every option.

How it fits together

Build, Work, and Data on top; 150+ connectors and every surface in and out; one map of the whole platform.

The rest of the platform

Related products and stories

Bring one ugly table and one slow dashboard.

We will land the table, rebuild the dashboard, and show you the cross-filter.