Skip to content
All posts

Product categorisation: mapping supplier categories to your own category tree

A supplier’s category column describes how the supplier runs its business, not how your shoppers look for things. Here is how to map it onto your own tree once, path by path, and keep it mapped as new files arrive.

BeforeWater FiltersWater Filter CartridgesAir Purifier Filters
AfterRefrigerator Accessories
Category path

Have a supplier file like this? Send it exactly as it arrived and see the same products before and after.

Why a supplier’s categories are not yours

A supplier’s categories are built around the supplier’s business: its price list, its warehouse and the way its own sales team talks about the range. Product categorisation in your catalogue asks a different question, which is where a shopper would look for this product, and the supplier’s column was never written to answer it.

Suppliers rarely agree with each other, either. A typical category column from an appliance supplier reads Appliances, appliances, APPLIANCES and White goods for the same kind of product, with the odd row left blank. None of it matches the menu on your own site, and a second supplier of the same products would file them differently again. Buy from a dozen suppliers and you are holding a dozen category trees.

The columns move, too. A supplier reorganises its range, renames a section or merges two, and the next file arrives with paths you have never seen. A catalogue that copied the supplier’s tree inherits every one of those changes.

So treat a supplier’s categories as evidence about the product, never as your categories. They are often a useful hint, because a path that says Spares is telling you something, but the output is always a node in your own tree. Categorisation is one of the four kinds of work in product data enrichment, and it comes early, because everything after it depends on where a product sits.

Map the path first, then the exceptions

Map at the level of the supplier’s category path rather than the product: one decision then covers every product filed under that path, in this file and in every file that follows.

A supplier sheet with thousands of rows usually uses far fewer distinct paths, and they are uneven, with a handful of big ones holding most of the products. Deciding path by path turns thousands of small decisions into a list you can finish, and it leaves you with something to keep. The next file from the same supplier maps itself, apart from whatever is new.

Some paths cannot be mapped in one move, and they are where the care goes.

  • Paths that mix different things. “White goods” may hold fridges, washing machines and dishwashers. No single node is true of all of it, so each product under it is filed on its own.
  • Blank and catch-all paths. There is nothing to map, so each product is filed from its own name and specifications.
  • Products the supplier filed wrongly. A path can be right for almost everything under it and wrong for a few products, which is why every mapping gets a spot check.
Supplier’s pathYour categoryDecided
Appliances > Fridge FreezersKitchen > Refrigeration > Fridge freezersOnce, for the whole path
White goodsSplit between refrigeration, laundry and dishwashersProduct by product
SparesThe parts branch of whichever family each part fitsProduct by product
Not suppliedFrom the product’s own name and specificationsProduct by product
An illustrative mapping sheet for one appliance supplier. Most paths are decided once; the ones that mix different things, or say nothing, are decided product by product.

A mapping method that survives the next file

A mapping method has to be a sequence, because each step needs the one before it, and it has to be repeatable, so that the second file from a supplier takes far less effort than the first.

  1. List every distinct path the supplier uses, with the number of products under each. Work from that list rather than from the rows, and start with the biggest paths, because that is where most of the products are.
  2. Map each path to one node in your tree: the most specific node that is true of everything under it. If no single node is, the path is not ready to map.
  3. Flag the paths that need splitting, along with blank and catch-all paths, and file the products under them one at a time from each product’s own name and specifications.
  4. Check a sample under every mapped path before trusting it. Open a few products under each and confirm they belong where the path sent them. A wrong mapping is wrong for every product under it.
  5. Record every correction as a rule, not as a one-off edit: which supplier, which path or kind of product, and where it goes instead.
  6. Apply the mapping and the rules to every import, so the supplier’s next file lands in your tree without anyone repeating the work, and only new paths and broken rules reach a person.

Categorisation is one stage of a longer routine; the guide to supplier data onboarding places it among the others, from the moment a spreadsheet arrives.

Keep parts apart from the products they fit

Parts and accessories belong next to the products they fit, not among them: a shopper browsing fridge freezers wants fridge freezers, and a shopper whose fridge has a split door seal wants the seal.

Supplier files mix the two constantly, because a supplier files by product line. To the supplier, the fridge, its water filter and its door seal are all refrigeration. Keeping them apart in your shop is what keeps a £9 door seal out of the same aisle as a £1,200 fridge freezer.

The damage goes further than a cluttered listing, because the category decides which attributes a product is expected to have. Shopify, for one, ties attributes to categories: assigning a product to a category in its standard taxonomy is what unlocks that category’s own product attributes, which Shopify calls category metafields.1 A door seal filed as a fridge freezer is judged on capacity and energy rating, has neither, and fills your appliance filters with blanks. Filed as a part, it is judged on the thing that matters for a part, which is what it fits.

So give each product family a parts and accessories branch of its own, and link every part to the products it fits, because the link is what sells them together. Supplier pages usually list what goes with a product as plain links with no codes, so the filter and the seal never connect to the fridge sitting in your own catalogue. Recovering the code from each of those links and matching it against what you already stock is part of the same job.

Learn the rules from your own corrections

Your team’s corrections are the best categorisation rules you have, because they record decisions about your shoppers that nobody ever wrote down.

On one of our recent jobs, for an appliance retailer’s refrigeration accessories, we filed a range across nine categories: water filters, filter cartridges, even air purifier filters. The client moved every one of them to the same place, accessories for the fridge. There were nineteen corrections, and all nineteen pointed one way.

We filed it asThe client moved it to
Kitchen Appliances > Water FiltersKitchen Appliance Accessories > Refrigerator Accessories
Water-Filtration Accessories > Water Filter CartridgesKitchen Appliance Accessories > Refrigerator Accessories
Air Purifier Accessories > Air Purifier FiltersKitchen Appliance Accessories > Refrigerator Accessories
Three of the nine categories we first filed that range under, and where the client moved them. All nineteen corrections on the job went to the same destination.

Nobody told us that rule. We read it off their own corrections rather than pushing a standard taxonomy back at them. By what they are, those products are three different kinds of filter. By where the retailer’s shoppers look for them, they are fridge accessories, and it is the shoppers’ logic that sells.

The lesson travels. One correction is an opinion; a run of corrections in the same direction is a policy. When they line up, write the rule down (this supplier, this kind of product, this node) and apply it to every file from then on, so nobody makes the same edit twice. It is the approach behind categorisation in RefynData’s product data enrichment service: products are filed into your own categories rather than pushed into a standard taxonomy, and on that job the rule came from the client’s own corrections.

When one range is scattered across the tree

The opposite failure is one range spread across many categories, and the fix is to bring it back together while keeping the old category as a filter rather than throwing it away.

On another of our jobs, for a bathroom retailer, sixty-two wall panels were spread across fifteen categories. Some were near neighbours, such as wall and ceiling tiles and cladding. Others were nowhere near: sofas, and categories as far off as coffee makers and saw blades. We pulled them into one range and kept the old category as a filter shoppers can still use.

Scattering like that builds up over time. Products are added by different people from different supplier files, each filed wherever seemed closest on the day, and no single decision looks wrong. The range only looks broken when you ask for all of it at once, which is exactly what a shopper does.

Keeping the old category as a filter matters because a range is often scattered along a distinction shoppers care about, such as a finish or a look. Moved into a filter, that distinction survives without splitting the range: it stops being a place and becomes a choice.

Once a range is together, the values inside its attributes have to agree as well, or the filters you have just made possible will offer the same thing under three names. The post on normalising attribute values covers that half of the job.

The taxonomies you do not control

Your tree is not the only one your products have to fit: Google, Shopify and ChatGPT’s product feeds each keep a category of their own, and each reads it partly from your data.

Google assigns every product a category from its own taxonomy automatically.2 Google says accurate titles, descriptions, pricing, brand and GTIN are what help it get that right, and it accepts an override in only three cases: where it has wrongly put a product in a category that brings extra required attributes (clothing, mobile phones and software among them), where your Google Ads campaigns are built on Google’s categories, and for alcohol. For most products, the way to be categorised correctly by Google is accurate data rather than a category field.

Your own tree still has a place there. Google’s product type attribute exists so you can “define your own taxonomy to match your shop’s navigation”, with every level of the path included and > between them, up to 750 characters.3

Shopify is built the same way round. Every product should carry a category from Shopify’s Standard Product Taxonomy, which Shopify can suggest from a product’s name, description and images, and which helps when selling on channels such as Facebook and Google. Your own terms go in its separate product type field.1

ChatGPT’s product feed asks for your tree directly. OpenAI’s specification has a product category field described as “your category path, from broad to specific, separated by >”.4

The practical conclusion is to keep one tree as the master and write it out in full, from the top level down, on every product. The channel categories are mappings you maintain from your tree, node by node, never values typed in product by product. And the fields those channels read to categorise you, titles, brands and barcodes, are the ones you are cleaning for your own shop anyway.

What good categorisation looks like

Good categorisation can be checked, and it is worth checking before a catalogue goes live, because a misfiled product is invisible to the shopper who is looking for it.

  • Every product sits in exactly one leaf of your tree, and none in a catch-all.
  • Every supplier path you have received is either mapped to a node or flagged for filing product by product.
  • Parts and accessories sit in branches of their own, each linked to at least one product it fits.
  • Each category’s products carry the attributes that category calls for, so its filters have values to offer.
  • Your channel categories are derived from your tree, not typed in by hand.

These sit alongside the completeness and consistency checks in the product data quality checklist, and like those, they are worth running on every import rather than once.

Questions

What is product categorisation in ecommerce?

Product categorisation is the job of putting every product in a catalogue on the path a shopper would follow to find it, from the top of your navigation down to the most specific category that is true of it. It decides where a product is listed, which filters it appears in and which attributes it needs, so it has to follow your shoppers’ logic rather than a supplier’s.

Should I use my supplier’s categories?

Use them as evidence, never as your tree. A supplier’s categories follow its own price list and warehouse, and every supplier’s differ, so adopting them leaves you with as many trees as you have suppliers. Map each supplier path to one node in your own tree, file the paths that mix different products one product at a time, and keep the mapping for the next file.

Where should spare parts and accessories go?

In a branch of their own next to the products they fit, not mixed in with them. A shopper browsing fridge freezers wants fridge freezers, and a shopper with a broken one wants the seal. Give each product family a parts and accessories category, and link every part to the products it fits, so each can be found and sold with the other.

Do I need to set Google’s product category myself?

Usually not. Google assigns every product a category from its own taxonomy automatically, helped by accurate titles, descriptions, pricing, brand and GTIN, and accepts an override only for category-specific attribute requirements, Google Ads campaign targeting and alcohol. Your own category tree belongs in the separate product type attribute, written out in full with every level included.

How do I keep categories consistent when new supplier files arrive?

Store the mapping, not just the result. Keep a table of every supplier category path you have seen and the node it maps to, record each correction your team makes as a rule, and apply both to every import. Then only genuinely new paths, and products that break an existing rule, need a person to look at them.

Sources

  1. Shopify Help Center. Shopify’s Standard Product Taxonomy. Read .
  2. Google Merchant Center Help. Google product category [google_product_category]. Read .
  3. Google Merchant Center Help. Product type [product_type]. Read .
  4. OpenAI Developers, Agentic Commerce. Products: product feed specification for ChatGPT. Read .

Try it on a fileyour supplier sent

Nothing goes live without your sign-offEvery value keeps its sourceYour data is never shared

Send one supplier spreadsheet exactly as it arrived. We will show you the same products before and after, with a source on every value we add.

One email with the upload link. Nothing else lands in your inbox.