Data platform optimisation
Cloud data platforms are sold on elasticity and paid for on consumption. The gap between those two facts is where a great deal of money quietly goes, and almost nobody reviews it after the migration business case is signed.
This is the least glamorous work I do and often the fastest to pay for itself. I have saved $1.1M rationalising platforms and $267K a year modernising a legacy estate, and neither required anyone to buy anything new.
What usually needs fixing
Consumption nobody owns. The bill arrives centrally, no team sees the cost of their own workloads, and consumption grows at the rate of enthusiasm rather than value. Attribution is almost always the single highest return change.
Compute sized for the worst day. Clusters and warehouses provisioned for a quarter end peak and left running the rest of the year. Auto suspend never configured, auto scaling never tuned.
Pipelines that rebuild the world nightly. Full refreshes where incremental would do, because full was easier to write and nobody has revisited it since.
Two platforms doing one job. A Databricks and a Snowflake estate that arrived through different sponsors, with overlapping workloads, duplicated storage and two sets of licences.
Licences bought against a forecast that did not happen, and never renegotiated.
How I approach it
Measure before touching anything. Which workloads cost what, who owns them, and which produce something a business person would miss. That last question usually removes more spend than any tuning exercise, because a material share of cost sits behind pipelines feeding reports nobody opens.
Then the sequence is boring and effective: attribute cost to owners, retire what is dead, right size what remains, tune what is genuinely hot, and only then argue about the licence.
Most platform overspend is not a tuning problem. It is an ownership problem that presents as a tuning problem.
Where I am careful
I am cautious about replatforming as a cost exercise. Moving from one warehouse to another to save on licence rarely survives contact with migration cost, parallel running and retraining. It is occasionally right, and it is proposed far more often than it is right.
I am also wary of aggressive right sizing before workload patterns are understood. Saving money by making the platform slow is a trade that gets reversed within a quarter, usually loudly.
Platforms I work across
- Databricks. Certified Data Engineer Associate, Generative AI Engineer and Machine Learning Associate. Lakehouse architecture, cluster policy, Unity Catalog, Model Registry and MLflow.
- Snowflake. Warehouse sizing and auto suspend, credit attribution, clustering and storage strategy.
- AWS. Certified Cloud Practitioner. Storage tiering, compute right sizing and FinOps practice.
- Azure, and the BI layer above all of it: Power BI, Tableau, Qlik and Looker.
- Legacy estates. Alteryx, SQL Server and Access, which is where a surprising amount of business critical logic still lives.
What you get
- A cost attribution model so every workload has an owner who can see what it costs.
- A ranked savings list, each item with the saving, the effort and the risk of doing it.
- A rationalisation plan where two platforms overlap, with the honest migration cost included.
- A decommissioning list covering pipelines and reports nobody uses.
- A licence and consumption position to take into your next renewal.
- Guardrails so the savings do not quietly reappear in six months.
How this connects
Platform optimisation sits alongside data modernisation and usually pays for it. It is also the fastest way to find out how healthy your estate really is, because consumption data does not flatter anybody.
The free gate check will tell you whether this is your blocker or a symptom of something further up the chain.
Find out what your platform actually costs you.
Bring last quarter's bill and, if you know it, who owns which workload. That is normally the whole conversation.