Datacube: Your Invoices Know More Than You Think

Every invoice you send contains far more than a revenue line - it holds churn, expansion, cross-sell and cohort data that most businesses never fully mine. This note covers what a single, governed model of that data looks like, why it matters to operational teams, finance and the board, and where we've seen these projects go wrong.

7 min read
01 - Why it matters

Where the hours actually go

Most finance teams spend the bulk of their month-end cycle reconciling billing exports instead of interpreting them - time not spent on the thing that actually moves the business: understanding why revenue moved, not just confirming that it did.

A datacube automates that reconciliation across millions of invoice lines a month, cutting the work from days to minutes and freeing the team for analysis and action instead. It's also the piece of infrastructure that due diligence teams look for directly: a business that can show clean, restated, cohort-level KPIs at a moment's notice is a business that's easier - and faster - to sell.

Finance

Month-end stops being a reconciliation exercise. The numbers are already governed - the job becomes interpreting them, not producing them.

Tech

One governed model instead of five spreadsheets and a dozen ad hoc queries - built once, trusted everywhere it's used.

Board

Answers on billing performance in the meeting itself, not two weeks after it - the same question never needs asking twice.

Investors

The same clean, cohort-level KPIs due diligence teams ask for, ready before they've asked for them - keeping the business exit-ready at any point, not just once a deal is already in motion.

Where the hours go

Illustrative - the typical split of a finance team's monthly billing-analysis time
Same headcount, same billing data - the difference is what the time is spent doing.
Where projects go wrong

Underestimating the time and effort to build the right platform

  • It's tempting to think you can flip the balance of hours above overnight. In reality, building the right platform properly means investing more time and effort for a while first - defining assumptions, mapping products, testing the model - before any of that work starts paying off. We've already made that investment, so you start seeing the benefit from day one.
02 - Why it isn't the accounts

Statutory accounts are a photograph. A datacube is a film you can re-cut

Once a month-end close is signed off, it's fixed - restating it means re-opening the books. A datacube doesn't have that constraint. If you reclassify a product, redefine a customer segment, or change what counts as "contracted" versus "one-off," the entire history restates instantly and consistently.

Statutory accounts

Stuck with month-end close

Product mapping and revenue classification are locked in the period they were booked. Change your mind about how a category should be treated, and last year's numbers stay exactly as they were - no longer comparable to today's.

Datacube

Fully restatable, always like-for-like

Product mapping, revenue-type mapping, even the customer hierarchy can all be rebuilt and reapplied across the full history in one pass - so a five-year trend line means the same thing in every single month.

Statutory accounts also aggregate up - by product family, by legal entity - collapsing detail that a datacube keeps intact. Every underlying invoice line, contract and product code stays queryable at its original granularity, however many ways the accounts choose to summarise it.

Where projects go wrong

Expecting it to tie to the accounts

  • Treating any gap between the datacube and the statutory accounts as an error to fix, rather than a consequence of answering a different question - one aggregated and legally fixed, one granular and restated. We help format that explanation, for audit and peace of mind.
03 - What it does

From Churn to Cohorts: The Complete KPI Set

Every invoice line exists to answer one question: is this business actually growing, and where's the opportunity in the base. Answering the first half means being able to trust the numbers - watching how customers move month to month through the Snowball, and comparing a newer cohort fairly against an older one through cohort analysis. Answering the second half means giving a sales team something to act on: product penetration and whitespace show exactly which products are under-sold today, and which specific accounts are the next cross-sell opportunity.

Every invoice line starts with a revenue type

Illustrative revenue base over time, £m

Contracted

A fixed value agreed in the contract, unaffected by how much the customer actually uses.

Recurring

Ongoing revenue that fluctuates with usage or volume - the opposite of fixed.

One-Off

A single transaction with no expectation of repeating.

The Snowball KPIs below are calculated on contracted revenue specifically - the portion of the base with real predictive value.

The Snowball bridge

Our name for the month-on-month revenue bridge - how a base of contracted revenue actually moves, illustrative figures, £m
Green = growth, red = erosion, amber = noise, blue = the organic growth checkpoint, teal = GRR and NRR.

Cohort analysis

Illustrative - how a cohort of customers behaves once it's on the books
Average revenue per new client, by the financial year they were won - illustrative, £.

Product penetration

Illustrative - share of the customer base holding each product, by quarter
Share of the customer base already holding each product, by quarter - illustrative.
Where projects go wrong

Getting the KPIs right

  • Spending most of the build effort on the model itself leaves little time to check the input assumptions behind it - like how revenue types are defined. We remove the build time entirely and put that effort into getting the assumptions right instead.
  • A simplified bridge with a single "net upsell" bucket - no split between volume, price, migration and noise - makes it hard to tell whether growth is real. We've already built the features to unlock that detail, down to volume, price, migration and noise.
  • Teams often lose weeks debating how each KPI should be calculated. We bring tried-and-tested, market-standard definitions, so that debate doesn't have to happen.
04 - The infrastructure behind it

Built to Scale: Why the Warehouse Matters

A spreadsheet, or a database bolted together with a legacy ETL tool, can just about handle a few years of billing history. It can't do what comes next: transform 100m+ rows in seconds so the analytics team stays agile, connect straight into modern tools like AI agents and data middleware without a rebuild, and stand ready for what's next - a customer 360 view pulling in data from other sources, and unstructured data like calls, emails and tickets. The part shown below just needs your data: plug it in, and it's ready from day one - already built, tested and automated on our side.

Source

Raw billing exports
Plug in your data - ready from day one

Data Lake (Blob)

Untouched, permanent record

Data Warehouse

System-agnostic - Snowflake, Fabric, etc.

Power BI / AI

Dashboards, or ask directly
<£100/mo
what we aim to keep Snowflake compute cost under, for our clients

Storage and compute scale with actual usage, and there's no dedicated infrastructure sitting idle between refreshes - a useful side effect of building this way, alongside the scale and speed it unlocks.

It also refreshes in under 15 minutes end to end, so the model can be brought current as often as billing changes rather than batched into a monthly exercise. And because every step happens inside your own warehouse environment - never a third-party server, never a separate copy of the data - it carries the same security posture as the rest of your infrastructure already does.

Where projects go wrong

Building on the right foundation

  • Using the wrong tool for warehousing - Excel, Power BI as a database, MySQL, SSIS - runs out of road exactly when the business needs it most. Our recommended out-of-the-box enterprise infrastructure is built to scale and cheap to run.
  • Complicated queries that take hours to run because the underlying model was never built to scale. We have invested the R&D time to ensure the model is fast and agile, so you can focus on the analysis from day one.
  • No testing on the model itself, so a broken assumption reaches a dashboard before anyone catches it. Our experience has helped us identify where things go wrong - the datacube model runs over 300 tests after each run.
05 - How the datacube is consumed

The same governed numbers, reached two different ways

Built on the warehouse covered above, the model's architecture is designed for BI and AI consumption at once - not two separate systems that happen to agree, but one governed source of truth that both are built directly on top of.

One warehouse, two ways to ask it a question

Both routes query the same governed model - never a separate copy of the data
Option 1

Dashboards

A fixed set of Power BI views covering the KPIs the business checks every week. Built to be investor-ready and operational at once - the same reports our clients have used with internal analytics teams and, in due diligence, directly with PE backers.

Option 2

Agentic AI

A conversational layer over the same model, for the questions a fixed dashboard can't anticipate - asked in plain English, answered from the governed numbers directly.

Where projects go wrong

Keeping AI as trustworthy as the model it's built on

  • An AI layer is only as reliable as what it's grounded in. We build it directly against the same governed warehouse, with the same access rules as the dashboards, so every answer carries the same trust from day one.

Curious what this looks like on your own billing data?

We'll walk through the model, the KPIs, and what it would take to get this running on your data.

Talk to Kolayo →