Product data for AI shopping assistants: what ChatGPT and Google’s AI Mode actually read
When a shopper asks an assistant what to buy, the answer is assembled from product data. Here is what that data has to contain, in the words of the companies that run the assistants, and why a typical supplier row falls short of it.
Have a supplier file like this? Send it exactly as it arrived and see the same products before and after.
What AI shopping assistants actually read
AI shopping assistants read the product data that merchants publish, and they answer a shopper’s question by matching its words against that data. Product data for AI shopping is not a new kind of data. It is the same names, brands, identifiers, categories and specifications your store already holds, read by a machine that has to decide whether your product answers the question in front of it.
There are two main ways in. The first is a feed: a file of your catalogue sent to the platform. OpenAI puts the purpose of its feeds plainly: “Product feeds provide structured catalog data that helps ChatGPT surface the right products with accurate pricing, availability, and seller context.”1 Merchants send a regularly refreshed file “containing key details such as identifiers, descriptions, pricing, inventory, media, and fulfillment options”.2
Google works at a larger scale. It says its AI Mode shopping experience “brings together Gemini capabilities with our Shopping Graph to help you browse for inspiration, think through considerations and narrow down products”, and in May 2025 it put the Shopping Graph at more than 50 billion product listings, with more than 2 billion of them refreshed every hour.3
The second way in is your own product page. Google says a page with Product markup can be eligible for its merchant listing experiences, including the shopping knowledge panel, Google Images, popular product results and product snippets.4 Either way, an assistant has what your data says to work with, and nothing more.
How a question becomes criteria your data has to meet
An assistant answers a shopping question by turning it into criteria, and a product can only be matched against the criteria its data actually states. Google’s own example shows the mechanism. Asked for a bag for rainy weather and long journeys, AI Mode “runs several simultaneous searches to figure out what makes a bag good for rainy weather and long journeys, and then use those criteria to suggest waterproof options with easy access to pockets”.3
Waterproof and easy-access pockets are attributes. A bag whose listing states neither may well be both, and still be passed over, because nothing in its data lets anyone confirm it.
Now take a question closer to a supplier sheet: a 60cm fridge freezer that fits an alcove 1.9 metres high. To answer it, an assistant needs a width, a height and a product type, each stated as a fact it can compare. Our illustrative fridge freezer, as its supplier sent it, has a code in place of a name, its height buried in 1850x595x655 with no unit, and no capacity at all. Every one of those facts is true of the product. None of them is in a form a machine can safely compare.
What OpenAI’s product feed asks for
OpenAI’s product feed specification for ChatGPT requires nine basic fields for every item and lists many optional ones that describe the product in more detail. The nine are an item ID, a title, a description, a product page URL, the brand, the seller’s name, a main image, availability and price.5
Two of the required fields carry rules that supplier data routinely breaks. The title is the “Product name, including the selected variant when relevant”, in at most 150 characters. The description is a “Factual product description for this item”, in plain text and at most 5,000 characters. A dealer code is not a product name, and “see datasheet” is not a description.
The optional fields are where matching happens. They include the GTIN (the barcode number), the manufacturer’s part number, a category path “from broad to specific, separated by >”, and material, colour, size, dimensions and weight. Their rules are strict in the way product data ought to be: a GTIN must include a valid check digit and keep its leading zeros, a part number keeps its punctuation and casing, dimensions need at least two of length, width and height plus a unit, and weight is net of packaging.5 OpenAI’s own guidance says optional fields “can improve answer quality”.6
| Feed field | Status | What the specification asks for | Illustrative supplier row |
|---|---|---|---|
| title | Required | The product name with its variant, at most 150 characters | NX-4821-W FF SS |
| description | Required | Factual, plain text, at most 5,000 characters | See datasheet |
| brand | Required | The brand as shown on the product page | Not supplied |
| gtin | Optional | 8, 12, 13 or 14 digits with a valid check digit, leading zeros kept | 5.01235E+12 |
| mpn | Optional | The manufacturer’s part number, punctuation and casing kept | nx4821w |
| product_category | Optional | A category path, broad to specific, separated by > | Appliances |
| dimensions | Optional | At least two of length, width and height, plus a unit | 1850x595x655 |
| color | Optional | The selected colour, consistent with the product image | Steel/Inox |
Hardly any of those supplier values is false. Each is a true fact in a form the feed cannot use, or a gap where a fact should be. The barcode is the worst of them: a spreadsheet has turned it into scientific notation, and the digits it dropped cannot be worked back out. They have to be found again.
Where a supplier row falls short
A supplier row usually fails an assistant in three places: the name, the specifications, and the consistency of its values.
A code where the name should be
A dealer code in the title matches nothing a shopper would type, and a range name that covers several models matches too much. When a name is too vague to identify one item, the answer that comes back can describe a product you do not sell. The cure is a full name with the brand, the model and the variant, backed by an identifier that ties the listing to exactly one product. The guide to GTINs, EANs and MPNs covers how to keep those identifiers intact on their way through a spreadsheet.
Blank cells for the facts people ask about
Shoppers ask about the things they would filter on: size, capacity, material, compatibility, power. A blank cell cannot be quoted, compared or matched. Enrichment fills those gaps from manufacturer pages and datasheets, with a link to the source on each value, so the facts are both present and checkable. Our guide to product data enrichment explains where those values come from and how to prove they belong to your product.
One value, spelt several ways
An assistant asked for a single ended bath has to decide whether “Single End Bath” means the same thing. Your own filters face the same problem. On one bathroom retailer’s catalogue of 329 baths, a single Type column held twelve spellings covering eight actual bath types, so shoppers got twelve filter options for eight real things. Folding each spelling onto one word fixes the filter, and gives every other reader of the data one fact instead of several. The post on normalising attribute values works through that example.
Fix the facts first, in order
Making products answerable is ordinary enrichment done in order, because each step depends on the one before it.
- Identify the product. Match the supplier’s code or barcode to the real item, and confirm that each source page shows that item.
- Fill the specifications a shopper would ask about, recording where each value came from.
- Split compound cells such as 1850x595x655 into their own fields, and put every measure in one unit.
- Fold every spelling of a value onto one word, so each fact is stated once.
- Keep identifiers as text, so a barcode keeps its leading zeros and a part number keeps its punctuation.
- Only then write the title and description, from the values you now hold.
OpenAI’s guidance puts the principle well: “If an optional field requires brittle transforms, omit it until data quality is stable.”6 An empty field is a gap. A wrong value is worse, because it can be repeated back to a shopper as fact.
Write titles and descriptions from the values
Titles and descriptions for AI shopping should be factual and built only from the specifications you hold, because that is what OpenAI asks for and because an invented detail in your copy becomes an invented detail in someone’s answer. The guidance is short: “Use concise, factual copy that helps users understand products.”6
In practice, a title should name the product the way a shopper would: the brand, the product type, the size or capacity that sets it apart, and the variant. Our fridge freezer goes from NX-4821-W FF SS to Northvale 60cm fridge freezer, stainless steel. Every word of that is a value already in the enriched data, and nothing in it needs checking twice.
That is the rule RefynData’s product data enrichment works to. It writes the shop title, the description, the page title and the search snippet for every product, in your tone of voice, using only values that already sit in your data. Length limits are enforced, not hoped for: if a title runs long, the least important detail is dropped until it fits, never cut mid-word and never padded out. The check then runs twice, once on the raw data and again after the copy is written, so every description is held against the specifications it claims.
Keyword stuffing gains nothing here. An assistant working from the criteria in a question has no use for a phrase repeated five times, and a description padded with search terms is harder for a person to read while giving a machine nothing new to compare.
Make the feed, the page and the image agree
Your feed, your product page and your images should all state the same facts, and OpenAI’s specification builds that expectation in. It asks for the brand “as shown on the product page”, a colour “consistent with the product image”, and a URL for the product detail page “with the variant selected when possible”.5
The reason is practical. A shopper who follows an answer to your page should find the product the answer described, in the colour it described, at the size it described. When a feed is built from one version of the data and the page from another, that promise breaks at the moment it matters most.
It is also the strongest argument for enriching once, at the source. If the supplier file is cleaned before it reaches your store or PIM, every channel built from it inherits the same facts: your pages, your Google feed, and any feed you send to OpenAI. Patch each channel separately and they drift apart. The product data quality checklist covers how to tell whether a catalogue is ready before any of them sees it.
Questions
What product data do AI shopping assistants use?
Structured product facts: the product’s name, brand, identifiers such as the GTIN, its category, specifications, images, price and availability. ChatGPT can read a product feed that a merchant supplies to OpenAI, and Google’s AI Mode draws on its Shopping Graph. The more complete and consistent those facts are, the more questions a product can be matched to.
How do I optimise products for AI assistants?
Fix the data before the copy. Give every product a full name and brand, fill the specifications shoppers ask about with a source on each, put every measure in one unit, fold spelling variants onto one value and keep identifiers as text. Then write titles and descriptions only from those values. There is no separate version of your catalogue to build for AI.
Does ChatGPT need its own product feed?
If you want to send your catalogue to ChatGPT directly, OpenAI’s product feed specification sets out the format, with required fields such as the title, a factual description, the brand, an image, price and availability. The data inside it should be the same data you publish everywhere else. Built from a clean catalogue, the feed is a formatting job; built from a raw supplier file, it inherits every gap.
Why do AI assistants describe the wrong product?
Usually because the listing’s name is too vague to identify one item. A dealer code or a range name can match several products, so an answer assembled from the wrong one describes something you do not sell. A full name with the brand, model and variant, backed by a valid GTIN or part number, ties the listing to exactly one product.
Does keyword stuffing help products appear in AI answers?
There is no sign that it does, and OpenAI’s own guidance asks for concise, factual copy. Assistants work from the criteria in a question, such as size, material or capacity, so what helps is stating those criteria as facts. A description padded with search phrases gives an assistant nothing new to compare and makes the listing harder for a shopper to read.
Sources
- OpenAI Developers, Agentic Commerce. Product feeds. Read .
- OpenAI Developers, Agentic Commerce. Key concepts. Read .
- Google, The Keyword. Shop with AI Mode, use AI to buy and try clothes on yourself virtually. Read .
- Google Search Central. Merchant listing (Product, Offer) structured data. Read .
- OpenAI Developers, Agentic Commerce. Products: product feed specification for ChatGPT. Read .
- OpenAI Developers, Agentic Commerce. Best practices. Read .