AI-ready data is business data an AI can reason about without you in the room. Every number has one written definition, a labeled source and period, and a link to the products, orders, customers and spend it belongs to. Facts only you know, like unit cost, are written down and approved. Answers are checked against the source of record before anyone acts on them.
It is 9:12 on a Monday.
You ask your AI one question. What was revenue last month?
Four seconds later you have a number. Clean. Confident. No hedging.
So you act on it. You hold the ad budget. You tell the team the month was soft.
And you are wrong.
Not because the AI made something up. Because nobody ever told it which revenue you meant.
Three right answers. One bad decision.
Here is the same month, measured three ways.
Three honest answers to one question. Gross sales before discounts and refunds, net sales after them, and the money that actually settled after marketplace fees. An AI with no written definition picks one and does not tell you. Illustrative numbers.
Gross sales: $590,000. Net sales: $540,000. Money that reached the bank: $445,000.
Every one of those is correct. Each answers a different question. Gross sales tells you about demand. Net sales feeds the P&L (your profit and loss statement). The settled amount is what you can spend.
Your AI picked one. It did not say which.
Next month it may pick another. Then it will compare the two and hand you a trend. The trend will be about a month that never happened.
That is the whole problem, and it is not an AI problem. Connecting Claude or ChatGPT to your store takes minutes now. The plumbing is solved. The meaning is not.
The five-minute test
Most brands have plenty of data and almost none of it is ready.
Here is how to tell in five minutes. Data is AI-ready when it is defined, connected, labeled, owned and checked. Miss one and the answers drift, quietly, in the direction that flatters you.
| Property | What it means | What breaks without it |
|---|---|---|
| Defined | Every metric has one written definition that every answer shares | “Revenue” means three things across three answers |
| Connected | Products, listings, orders, customers and ads are linked to each other | The AI cannot tell which ad sold which product, or that five listings are one product |
| Labeled | Every number carries its source, period and known caveats | A half-finished month is compared to a closed one and reads as a collapse |
| Owned | Facts the systems cannot know are written down, dated and approved | Margin is calculated with last year’s unit cost, or with no cost at all |
| Checked | Answers are compared to the source of record before anyone acts | Two tools disagree and the AI averages them into a number nobody recorded |
Take them one at a time. None of this needs a data team.
Defined: write the sentence once
A definition is a sentence plus a formula. Here is one:
“Net revenue is gross sales minus discounts, refunds and cancellations, excluding shipping and tax, by the date the order was placed.”
Write it once. Yours may differ from mine, and that is fine. The point is not whose definition wins. The point is that every answer uses the same one. Otherwise this week and last week were never comparable.
Start with the ten metrics your team argues about most: revenue, new customers, repeat rate, ad spend, return on ad spend, contribution margin, units, average order value, refunds and inventory on hand.
Ten sentences. One afternoon. It is the cheapest thing on this list and it removes half the arguments.
Connected: the bottle with three names
Take one bottle of shampoo.
On Amazon it has an ASIN (Amazon’s code for a product), say B0-something. On Shopify it is variant 4412. On the retail shelf it carries a UPC (the barcode number).
Three codes. One bottle. Nothing in your data says so.
So your best seller arrives at the AI as three modest products, and the AI names something else as your best seller. The sums are right. The answer is still wrong, because the AI is counting one product as three.
Your data describes things. A product family and the variants inside it. The listings for each variant, the orders that bought them, the customers who placed those orders, and the campaigns that brought those customers in. A written-down model of those things and how they relate is called an ontology.
It sounds academic. In practice it is a map. This listing belongs to this family. This campaign sends shoppers to Amazon, not to the store. This customer first bought this product.
Without the map, ask “how is our repair line doing?” and the AI adds up whatever rows happen to have “repair” in the title. Our guide to product hierarchy walks through building it.
Labeled: a number without a label is a trap
Every figure an AI reasons about should arrive carrying three things.
- Its basis. Order data, settlement data, or store gross sales?
- Its period, and whether that period is closed. The current month is still in flight, meaning not finished. Orders will cancel, refunds will land, and ad platforms will restate (revise their earlier numbers).
- Its blind spots. Meta’s pixel (the tracking code Meta puts on your store) cannot see a purchase that happens on Amazon. Amazon buyer emails can arrive days late, so recent customer counts run low. A cohort (a group of customers who first bought in the same period) that is six weeks old cannot tell you its repeat rate yet.
Labels are what let an AI say the most valuable sentence it can say: “this looks like a drop, but the month is not closed.”
Any AI can answer. Yours should know when not to.
Better still, the label travels with the number itself. Then no prompt (the message you send the AI) has to remember to mention it, and no teammate has to know the trap exists. We cover the most common mislabeled numbers in why AI gets your ecommerce numbers wrong.
Owned: the facts that live in nobody’s system
Some of the most important numbers in your business are not in any system you connect.
What a unit actually costs to land. When a product launched. Which ads point to Amazon. What margin you consider healthy. Whether “the 3-pack” means the 3-count bottle or three separate bottles.
Those facts live in a founder’s head, a supplier email, and a spreadsheet somebody renamed. So margin gets calculated with last year’s cost, or with no cost at all, and the answer looks fine either way.
A fact needs an owner, a date and an approval. “Landed cost for the 90-count bottle is $4.10 per unit, effective 1 June, approved by finance” is a fact. A cost figure someone typed into a chat three months ago is not.
One more rule, and it is the one people get backwards. Facts are stored. Metrics are not. Lifetime value (what a customer spends with you over time) and return on ad spend get recalculated every time. Save them as facts and they go stale in silence.
Checked: never average a disagreement
Even good data produces bad answers sometimes. So the last property is a habit built into the system.
Before anyone acts, compare the answer to the source of record: the system that actually recorded the money. On Amazon that is the settlement report, which shows what Amazon actually paid you. On Shopify it is payouts. For ads it is the platforms’ invoices. For costs it is the ledger in your books.
And when two figures for the same metric and period disagree, say so out loud.
Never average them. An average of two numbers nobody recorded is a third number nobody recorded, and it is the easiest lie in analytics to tell by accident.
Where the ontology ends and the harness begins
The five properties split into two jobs.
The ontology holds the meaning. The definitions, the map of things and links, the approved facts. It also holds governed actions, which are changes with rules attached. A proposed new unit cost takes effect only after a person approves it. This is your brand’s memory, written so a machine can use it.
The harness governs the AI. A harness is the layer wrapped around a model. It decides what the model can reach: your brand and nothing else. It decides what the model is handed with each question: the definitions, labels and facts that apply. It decides how answers get checked before you see them.
The model does the reasoning. The harness keeps that reasoning attached to your real numbers.
You need both. An ontology with no harness is a well-organized library the AI may or may not visit. A harness with no ontology is a strict referee with no rulebook.
How to make your data AI-ready
You can do most of this in a few focused weeks, with or without software.
- Pick a source of record for each money number. Settlement for Amazon revenue and fees, payouts for Shopify, invoices for ad spend, the ledger for costs. Orders and clicks stay useful as early signals. They are not the truth.
- Write your product family map. One row per family. Its name, and every SKU (your own product code), ASIN, variant and UPC inside it. Add units per pack, launch date and status.
- Write definitions for your ten most argued-about metrics. One sentence and one formula each, in plain words.
- Write down the facts nobody’s system holds. Landed cost per unit with an effective date, launch dates, targets, which campaigns drive Amazon. Give each one an owner.
- Mark every period as closed or in flight. Decide your close rule, for example a month is closed ten days after it ends, and apply it everywhere.
- Test with questions you already know the answer to. Pick five whose answers are surprising, such as a product that sells well and loses money. If the AI only gets the easy version right, the data is not ready. A test that cannot fail is not a test.
If you want the long version of this, we published it as a standard: twenty-five checks you can score yourself against, free to copy.
The six questions
Answer these out loud, right now:
- Does everyone on the team get the same number for last month’s net revenue?
- Can you list every listing that belongs to your best-selling product family?
- Do you know the landed unit cost of every active product, and when it last changed?
- Can you tell which campaigns send shoppers to Amazon rather than your store?
- Is the current month clearly marked as unfinished everywhere it appears?
- When two tools disagree, is there a rule for which one wins?
Three or more “no” answers, and an AI will hand you confident answers you should not act on yet.
That is not an argument against using AI. It is an argument for spending one week on the five properties first, because the AI is no longer the hard part.
Where Synthesis fits
Synthesis is ontology software and an AI harness for consumer brands. It joins Amazon, Shopify, ad platforms, email, subscriptions and accounting into one model per brand, isolated from every other brand. The definitions and the product family map are written once and shared by every answer. Every number arrives labeled with its basis, period and caveats. Then it serves that model to Claude and other AI tools, and flags figures that disagree instead of quietly picking one.
Questions people ask
What is AI-ready data?
Data an AI can reason about correctly on its own. Each number has one agreed definition and carries its source and period. It connects to the business things it describes. It can be checked against the system that actually recorded the money.
Can I just connect ChatGPT or Claude to Shopify and Amazon?
You can, and it will answer. Access is not understanding. A raw connection hands the AI tables with no shared definitions. So it picks one meaning of “revenue” and mixes finished months with months still in progress. It also cannot see costs or facts that live outside those systems.
Is AI-ready data the same as a data warehouse?
No. A warehouse puts the data in one place. AI-ready data also needs the definitions, the relationships between things, labels on every number and a way to check answers. A warehouse is the floor, not the building.
What is the difference between an ontology and a semantic layer?
A semantic layer is a set of metric definitions, such as how net revenue is calculated. An ontology also models the things in the business and how they connect. It also covers what can be done about them, such as approving a unit cost. Brands need the definitions either way.
What is an AI harness?
It is the layer around an AI model. It decides what data the model can reach and what context it gets with each question. It also checks the answers before a person sees them. The model does the reasoning. The harness keeps it on your real numbers.
How long does it take to make a brand's data AI-ready?
The data connections can be set up in a day. The part that takes thought is writing down definitions, the product family map and the facts only the team knows. For most brands that is a few working sessions. Then come small fixes as questions show gaps.