Definition

Taxonomy and Classification: Getting It Right

A taxonomy is the set of categories you sort records into. Classification is the act of sorting them. Both look like administration and behave like architecture: once a few thousand records have been classified one way, the cost of changing your mind is high, and every report built on the taxonomy inherits its assumptions.

What it covers

Taxonomy work is deciding what the categories are and what belongs in each. Classification work is applying those decisions at volume, consistently, including in the awkward cases.

  • Defining or documenting the categories, and the rule for each one
  • Writing down the edge cases and how they are treated, which is where consistency is actually won or lost
  • Classifying and tagging records against the taxonomy
  • Reclassifying when the taxonomy changes, and keeping a record of what changed
  • Quality checking classification decisions, particularly at the boundaries
  • Flagging records that do not fit any category cleanly, rather than forcing them

What it does not cover

We do not decide what your categories should be. A taxonomy encodes how your firm sees its market, and that is a commercial judgement. We will document what you decide, apply it consistently, and tell you where it is producing odd results.

We do not silently reinterpret a rule. Where a record does not fit, it is escalated rather than classified by assumption. This is slower, and it is the difference between a taxonomy that holds and one that quietly drifts.

We do not fix a bad taxonomy by classifying harder. If the categories overlap or a large share of records will not fit, the answer is to change the categories. We will say so.

Who normally owns it

Usually nobody, formally. The taxonomy tends to have been set by whoever built the system, then extended by whoever needed a new category, then inherited by people who were not there for either decision.

That is why classification drifts. Two people applying the same category to different things is not carelessness; it is the predictable result of a rule that was never written down. Most classification problems are documentation problems.

The most common symptom is a report nobody trusts. Somebody runs a breakdown by sector, the numbers look wrong, and the reason is that "sector" has meant three different things over four years.

How Clandon supports it

We document the taxonomy and its edge cases, classify to it consistently, and quality check at the boundaries where errors actually occur. Where the categories are producing results that do not make sense, we will show you where and why, and the decision about changing them stays with you.

Related resources

Filling the gaps in a dataset, and checking that what is already there is still true.

Building and maintaining a target universe, and the factual research that comes off it.

What clean CRM data actually means, and what it takes to keep it that way.

Does your sector breakdown mean the same thing it did three years ago?