How to build a consumer decision tree: why the best CDTs start with shopper behavior, not supply logic

Sep 29, 2026 • 10 min

Building a consumer decision tree sounds straightforward enough until someone stops to ask, “Who built it?” and “How did they build it?”

Most category managers didn’t build their CDT. They inherited it from a brand supplier, a consultancy, or an internal category team that has long since moved on. Yet they have been optimizing within it ever since. They update products each year, debate node names, and run assortment reviews against a structure built once by someone else for reasons that may no longer apply.

That matters more than it sounds. A consumer decision tree sits at the foundation of every assortment decision a retailer makes. It determines which products are selected or not selected across different need states to ensure the right breadth and depth for an optimal assortment.

Get the structure wrong, and a retailer can score well on every financial metric while quietly leaving an entire consumer need state unserved. No red flags will appear in the data because the tree doesn’t reflect how shoppers actually make decisions.

Updating the existing decision tree isn’t enough. The fix is building the next one from actual shopper behavior — captured in how products co-occur across baskets and visits over time.

How consumer decision trees (CDTs) are built today

There are a few common ways category managers arrive at a consumer decision tree. Each one is created in its own way, but they all share the same fundamental limitation: these trees capture a single snapshot that stays fixed while the category moves on.

A diagram showing three traditional ways to build a CDT: vendor-built, workshop-built, and attribute-based.
Fig. 1: All three traditional build methods lock in a single snapshot — from a supplier’s research, a workshop, or an internal attribute hierarchy.

Vendor- and supplier-built trees

This is the most common source for CDTs. Brand suppliers and category management consultancies bring genuine category knowledge and research capabilities that most retailers don’t have in-house, making them a natural fit for this kind of work. Research typically combines shopper surveys, market data, and category expertise to produce a hierarchy of attributes ranked in the order in which the supplier believes shoppers make decisions.

The resulting tree maps out which products belong together and how the category should be navigated, from the top-level category node down to the individual SKUs at the leaf level (the bottom-most nodes).

Category reset workshops

The alternative for retailers who build their own CDTs is the category reset workshop. Store visits, competitor observation, and category manager judgment are combined into a structured exercise that draws on firsthand category knowledge over a single period. The output is a hierarchy built from direct observation of how categories are organized in-market. A hierarchy that’s interpreted through the lens of the category team’s expertise, and overlaid with sales performance data to validate the structure.

Internal attribute logic

The third approach is the one that often goes unexamined. Decision trees are built from product attribute hierarchies, defined internally by the category team based on how the category is organized. Attributes are ranked in a sequence that reflects how the team understands the category, and products are assigned to nodes based on where their attributes place them in that hierarchy.

The attributes that are ranked may include:

Each lowest-level node then contains the products that share the same attribute path from top to bottom of the tree.

Why do inherited CDTs leave consumer need states unserved?

Each of the typical CDT build methods has its own specific flaw, but they all share the same consequence, regardless of how the tree was built. The CDT determines what gets protected. If a need state isn’t represented in the tree, the system has no way to flag its absence. It simply fills the assortment from the nodes that do exist, and the gap stays invisible.

An icon set showing three flaws of inherited CDTs: brand bias, outdated trees, and attribute-based logic.
Fig. 2: Every inherited CDT fails differently, but all three leave gaps in shopper need states that can go unnoticed.

Supplier bias pushes brand to the top

When brand sits near the top of the hierarchy because a supplier put it there, the tree reflects supplier priorities, not shopper decisions. A shopper whose first criterion is occasion, or format, or country of origin will navigate the category in a sequence that the tree doesn’t model. The assortment can look comprehensive by the tree’s logic, while systematically underserving shoppers who don’t lead with brand. There’s no mechanism to surface that problem because the tree itself defines what “comprehensive” means.

Workshop trees go stale

A category reset workshop produces a snapshot. While this is useful in the short term, shopper behavior shifts, new formats emerge, competitive dynamics change, private label grows, yet the tree doesn’t move with it. The category manager may sense something is off, but the CDT doesn’t confirm it, because it was built to reflect a moment that is now outdated. Updating it means running the workshop again, which in turn means waiting until the next cycle to make the necessary adjustments.

Attribute logic mistakes organization for insight

A tree built from product attributes (such as brand, size, flavor, or format) reflects how the category is organized in a database, not necessarily how a shopper moves through a decision. The two can look similar on paper, yet diverge significantly in practice. If a shopper reaches for a different brand before they switch from one format to another, but the tree has brand above format, the hierarchy is inverted relative to actual behavior. The assortment decisions that follow will be optimized against the wrong logic.

The better way to build a consumer decision tree: start with substitution data

The methodology that addresses all the problems with inherited CDTs starts with a different question. Instead of asking “how should we categorize these products?”, ask: “Which products do shoppers actually treat as substitutes for each other?”

The input is POS transaction data: the actual purchase records of what shoppers bought across baskets and visits over time. Rather than imposing a hierarchy based on product attributes or supplier input, category managers let co-occurrence patterns in the data define the structure.

The tree that emerges reflects how shoppers behave when their first choice isn’t available, or when they’re weighing up options within a category. This CDT isn’t based on how a supplier thinks they should behave, or how the category looked during a reset workshop two years ago.

Rather than imposing a hierarchy based on product attributes or supplier input, category managers let co-occurrence patterns in the data define the structure.

In practice, the hierarchy this produces is often surprising. A supplier-built tree for a personal care category might place brand high in the hierarchy, although transaction data can tell a different story. The application method may matter more to shoppers than the brand when their first choice isn’t in stock.

For instance, a shopper might reach for a different brand of deodorant before they’d switch from roll-on to spray. If a pattern like that holds consistently across thousands of transactions, it should be reflected in the tree rather than built from product attributes or vendor input.

The underlying principle is straightforward. The CDT should reflect pure consumer behavior, with category knowledge used to explain and name what the data surfaces. The data defines the structure; the category manager interprets it, names the nodes, and makes it usable.

Why substitution-based CDTs produce better assortment outcomes

The immediate advantage of substitution-based CDTs is coverage. Without an accurate tree, top sellers in a category tend to cluster into one or two nodes. The assortment optimizer, working correctly from the structure it’s given, fills those nodes well, which means depth in areas that were already winning and nothing in the areas that needed covering. A substitution-based tree distributes nodes more accurately across the real shape of shopper demand, so the optimizer protects the right things.

The second advantage is that transaction-scale co-occurrence data is difficult to dispute in the way a vendor’s CDT is not. A supplier-built hierarchy can be challenged on intuition, but consistent co-occurrence patterns across hundreds of thousands of baskets are harder to dismiss. When the data reveals something unexpected, it’s usually pointing to something the last category workshop missed — and that’s the point.

The third is scalability. A manually built CDT is one person’s effort, done once per category per cycle, limited by how many categories that person can work through in a year. A model that identifies co-occurrence patterns from transaction data can run across many categories simultaneously, with the practitioner reviewing and validating output rather than building from scratch. The work shifts from construction to interrogation, which is a better use of category expertise.

What it takes to build a consumer decision tree from shopper behavior data

Building a behavior-based CDT is more than a choice of methodology. It has certain practical data and technology requirements.

Fig. 3: A behavior-based CDT replaces assumption with evidence at every stage, but a person still signs off on a new tree before it goes live.

Up-to-date, basket-level POS data and product attributes

The foundation involves two things working together:

  1. Transaction data that captures what shoppers bought across baskets and visits over time
  2. A clean product attribute export that describes each SKU’s properties (format, brand, flavor, price point, etc.)

The substitution signal comes from individual purchase patterns, not category totals, so aggregate sales data isn’t sufficient. The model needs both inputs to identify which attributes explain the differences between products that shoppers treat as substitutes.

A trained model that ranks attributes by explanatory power

The key difference between a behavior-based tree and a conventional one lies in how the hierarchy is determined.

In a traditional CDT, someone decides in advance that brand sits above format, or format above size, based on judgment or convention.

In an approach based on co-occurrence patterns, a trained model reads the customer’s transaction data and decides which attributes to split on and in what order. The decision is based on how well each attribute explains differences between products at each branch.

The result is a hierarchy derived from actual shopper behavior. Products that land in the same leaf node are, by that logic, treated as substitutes: what a shopper reaches for when their first choice isn’t available.

A direct connection to the optimization engine

A CDT that sits in isolation from the assortment planning process only solves part of the problem. The downstream decisions, such as “which products are core?,” “which are optional?,” and “what gets rationalized?,” need to inherit the same shopper logic the tree was built from.

When the CDT feeds directly into a planning system’s assortment optimization engine, the tree’s structure flows through to every assortment decision. The optimizer can then use each terminal node as a coverage constraint. This ensures the assortment contains at least one product for each consumer need state, rather than filling it from aggregate category totals.

An activation model that keeps the practitioner in control

A model that generates decision trees automatically and deploys them without review creates a different problem. Category managers lose visibility into what changed and why, and trust erodes quickly.

The right approach is to generate new trees as inactive by default. This means that nothing goes live until a category manager has reviewed the output, labeled the nodes, validated the substitution groups against commercial reality, and explicitly approved the update. It’s the part of the methodology that’s easiest to underestimate, but also the part that makes the rest of it work.

Category managers working this way are clear about what it needs to feel like. The goal is to sit down, be in the right headspace, and feel like the category manager is doing their CDT, not discover it has already been done for them. The model runs on a schedule; the category manager decides when the new version goes live.

The role of judgment shifts from building a tree from a blank page to interrogating what the data found and deciding which parts to trust. That’s a more productive use of category expertise than manual attribute classification.

How RELEX CDT Builder puts this methodology into practice

RELEX CDT Builder brings together the data inputs, modeling approach, and activation model that behavior-based CDT construction requires. It’s part of RELEX Assortment Planning, which sits within RELEX’s unified platform, allowing assortment decisions to connect with space, demand, and replenishment planning.

The process uses a trained topic model to learn behavioral patterns from basket data. A typical build run takes a few minutes per category and can be configured at the category level. Category managers can control maximum tree depth, minimum node size, and which attributes must or must not appear on every path.

The tree it produces connects directly to the Assortment Optimizer, where each terminal node functions as a coverage constraint. The optimizer ensures the assortment contains at least one product per consumer need state, so depth in winning nodes and coverage of underserved ones are managed together rather than in tension. Category managers review and label the output, validate the substitution groups against their category knowledge, and activate the tree when they’re satisfied. This allows them to retain full control over what goes live and when.

A flowchart of the RELEX CDT Builder process, from POS data through to the assortment optimizer.
Fig. 4: The RELEX CDT Builder generates new trees based on basket data, but nothing reaches the assortment optimizer until a category manager reviews and activates it.

For teams managing large or complex category portfolios, the shift in scalability is significant. The CDT Builder runs across many categories simultaneously, with category managers reviewing and validating output rather than building from scratch. The work moves from construction to interrogation, and the categories that would previously have waited months for a manual rebuild get updated on a model-driven schedule instead.

The long-term direction builds on this foundation. Demand transfer-aware forecasting, including new item forecasting via CDT node assignment, will extend the same behavior-based logic from assortment decisions to the broader planning process.

Make the next CDT rebuild count

Most category managers are working from a CDT that someone else built, using a method that made sense at the time, which now desperately needs updating. The question is whether the next version will be built on how shoppers actually behave, or on the same assumptions as the last one.

For those ready to make that shift, RELEX Assortment Planning is the practical next step. Built on the substitution-based methodology, this solution is designed to put category managers back in control of the structure their assortments depend on.

Written by

Jussi Niemi

Product Manager for Assortment

Jussi Niemi is a Senior Product Manager at RELEX Solutions, based in Helsinki, Finland. He has over six years of experience at RELEX, specializing in assortment planning software and product discovery.