One definition of conversion
What building a metric semantic layer taught me about teams. The schema was the easy part. Ownership, and making the shared definition easier than a copy, decided whether it got used.
Two dashboards show a metric called “conversion” for the same week. The numbers differ. Someone asks which one is right, and the next hour goes on reading two pieces of SQL side by side. By the end nobody has learned anything about the product. The conversation has become a debate about queries.
I designed a metric semantic layer for an experimentation platform that served dozens of product teams: a centralised metadata schema, plus a query engine on BigQuery that computed metrics from it. The catalogue held more than 100 reusable metrics, and every product team’s experiments used them. This post is about what that work taught me. Most of it was about people, and the schema was the smaller part.
What a semantic layer is
A semantic layer is a place where the meaning of a metric is written down once, in a form that both people and programs can read. Instead of each dashboard or experiment analysis carrying its own SQL for “conversion”, they all refer to one definition, and a query engine turns that definition into the query that runs against the warehouse.
Vendors have plenty to say about this. I’m going to skip the theory and describe the problem it solves.
How drift happens
Nobody sets out to define a metric wrongly. Drift happens because every team has a reasonable local reason:
- One team counts a conversion when an order is created. Another counts it when the order is paid.
- One divides by all sessions. Another divides by sessions from logged-in users.
- One excludes internal test traffic. Another never got round to it.
- One counts per user per day. Another counts per session.
Each choice is defensible. The trouble is that all of them carry the same name, and a name is what people compare.
For an experimentation platform this costs more than for a dashboard. An experiment reports the effect of a change on a metric. If the metric can mean different things in different teams, results are not comparable, and a headline like “this change lifted conversion” cannot be checked against the last one.
What a definition has to say
A definition that both a person and a program can rely on has to answer a few questions explicitly. This is the anatomy I’d expect from any semantic layer:
- Name and description in plain language, so a product manager can tell whether it is the metric they want.
- Owner, a person or a team that answers questions and approves changes.
- Grain: what one unit of the metric is. Per user, per session and per order give different answers to the same question.
- Numerator, denominator and filters: the exact arithmetic, including what is excluded.
- Version, so a change is visible and old results stay interpretable.
The schema had to be precise enough for a program to compute from it. That was the job of the query engine: it took a definition and produced the BigQuery query for it, so the same definition gave the same number wherever it was used.
Ownership matters more than the schema
You can build the schema and engine and still end up with two versions of conversion, one governed and one not. The technical work leaves open the harder question: who decides what “conversion” means?
In my experience the answer has to be a named owner for each metric. Not a committee, and not “the data team” in general. The person who owns a definition can decide when a request to change it is right, and can say no when a team wants a local variant that would fragment the meaning. Without that, the catalogue turns into a list of opinions.
A few practices follow, and they are the ones I’d insist on:
- Every metric has an owner in the definition itself. If you can’t say who owns it, don’t publish it.
- Changes are reviewed and versioned. A change to a definition changes past comparisons, so it should be visible and deliberate. Old versions stay readable.
- Disagreements get resolved in the definition. If two teams need different things, that usually means two metrics with two names, not one metric with a silent fork.
- Deprecation is explicit. A metric nobody owns any more should be marked so, rather than lingering.
Adoption is the real test
A catalogue that nobody uses is documentation. What made the difference to me was that the governed definition had to be the easiest option, not just the correct one. If a team can write its own SQL in ten minutes but must file a request and wait a week to use a shared metric, the local copy wins every time, and it wins for good reasons.
That pointed the design at a few things. Some of these I did and some are what I’d emphasise now, and I’ll keep to the general principle rather than the detail of any one system:
- Meet people where they already work. The natural home for a governed metric was the experimentation workflow, because that is where teams pick the metrics an experiment is judged on. If choosing a shared metric is a step in a task people already do, it costs less than writing a query.
- Cover the metrics people argue about first. A catalogue with the thirty metrics that cause the most disagreement is worth more than one with three hundred that nobody trusts. The catalogue grew past 100 reusable metrics, and the growth mattered less than whether it contained the ones teams actually reached for.
- Make search and description good. People pick the metric they can find and understand. A clear name and a plain description save a lot of local reinvention.
- Treat requests as a signal. When a team wants something the catalogue doesn’t have, either the catalogue has a gap or an existing metric is described badly. Both are fixable, and both are better than letting the team quietly go its own way.
- Make the first use painless. The moment someone tries a governed metric and gets a number they trust is the moment they stop writing their own. Anything that makes that first attempt fail, such as a confusing name or a missing filter, sends them back to local SQL.
Every product team ended up using the catalogue in its experiments. I’d credit that to the definitions being useful and reachable, and I don’t think a mandate would have produced the same result. A rule can make people register a metric. Only a metric that saves them work makes them keep using it.
Ownership needs a process, not just a field
An owner field is a start. What it needs behind it is a path for changing a definition that people can follow without heroics. A common shape looks like this:
- Someone proposes a change to a definition, with the reason and an example of the difference in the number.
- The owner reviews it, and any team that depends on the metric is told.
- The new version is published alongside the old one for a while, so results can be compared across the change.
- The old version is retired on a date everyone knows.
None of this is complicated. The point is that changing what “conversion” means is a decision that affects other people’s results, and the process makes it visible. Otherwise the definition changes silently, and the next dashboard comparison is a debate again.
What I’d watch for
If I were starting again, I’d check for these early:
- Over-modelling. It is easy to design a schema that can express everything, and then nobody can fill it in. Start with what the common metrics need.
- Definitions without tests. A definition should be checkable against a known result, so a change to the engine can’t silently change a number.
- Local variants hiding in plain sight. Watch what people compute outside the layer. It shows you the gaps.
- Ownership drifting. People move teams. An owner field that is out of date is worse than none, because it looks like governance.
What I took from it
The schema and the engine mattered, but a shared definition only becomes a working practice through reuse, and reuse depends on people. Someone has to own each definition, changes have to be visible, and the shared version has to be easier to reach than a copy. Get those right and a dashboard comparison goes back to being a conversation about the product.
References
- Chang, R. How Airbnb Achieved Metric Consistency at Scale. The Airbnb Tech Blog. https://medium.com/airbnb-engineering/how-airbnb-achieved-metric-consistency-at-scale-f23cc53dea70 — Minerva, a metric platform with one definition per metric
- dbt Labs. dbt Semantic Layer. dbt documentation. https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl — centralised metric definitions, reused across tools
- dbt Labs. About MetricFlow. dbt documentation. https://docs.getdbt.com/docs/build/about-metricflow — generating SQL from metric definitions
- Kohavi, R., Tang, D., and Xu, Y. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020. https://experimentguide.com/
- Google Cloud. BigQuery overview. Google Cloud documentation. https://docs.cloud.google.com/bigquery/docs/introduction