Synthetic data is a lie businesses want to believe

Synthetic data is the idea that we can create personas with AI, and use these personas to understand what real people want. 

For example, rather than ask one hundred real women over fifty-five what they think of a holiday package, we can have one hundred AI generated women over fifty-five (or even just one that is the archetype of a woman over fifty-five) and ask them questions.

The concept is that because we've had so many women over fifty-five answering questions in the past, we have enough data to create a synthetic woman over fifty-five, and we can use it to answer our questions, give opinions on our products, and guide our business decisions. 

To some readers, this may sound like a bad idea. However, this is a serious undertaking in the age of AI. Companies and market research firms alike are spending millions of dollars building this synthetic data, creating digital clones of their client types. 

There are two main forces behind the hype for synthetic data. The first force is profit. It may be faster and cheaper to run ideas through a customized chatbot than run real market observations. Companies also tend to want to run many ideas through synthetic data, such as sixty concepts for a new drink. 

The second force is at best an under appreciation for the complexity of consumer decisions, at worst a disdain for consumers altogether. In this view, consumers are simplistic and fairly stable. It is not offensive to think that we can create the 'archetype' for a fifty-five years old woman, it is just part of the business plan. 

These forces — profit and a simplistic view of consumers — collide to shroud key problems behind synthetic data.

The first set of problems are found in traditional consumer surveys, they're now just accentuated. Do people themselves know why they're purchasing a product, what they want to see in the market? Are people reliable in expressing their expectations, and how wide is the gap between what they say they want, and what they actually open their wallets for?

The second set of problems are mostly about the limited nature of the clones. If some information is only found in the offline world or in people's brains, the clones do not have access to it. If people buy a product because of a belief or a quirky way they use the product at home, and they never shared it online or through a survey before, the clones will not have access to this information. The gap between the lived experiences of consumers and their digital selves is bound to be an issue for synthetic data. 

The third set of problems is about timing and updating. AI clones, even if connected to the internet, and even if starting with a great foundation of information about a key issue, need updating at a scale that may not be doable. New events happen every day. Things come out of fashion, culture shifts, ideas evolve. AI clones have no way of keeping up. The clones built on the consumers of yesterday might recommend you a collab with Labubu. 

So, synthetic data is unreliable because it doesn't account for the limitations of surveys as a methodology, introduces a clone restricted in information and without the ability to update itself at the speed general culture and relevant information evolve. 

Yet, the deeper challenge comes from the views that lead a business to consider using synthetic data in the first place: that getting close to and serve your customers isn’t the purpose of a business.

Business is a relationship, the whole point is to get closer to your customers. 

Synthetic data is selling the lie that you’ll reach the perfect understanding of your consumers by avoiding talking to them. Its strategy is an oxymoron; get closer to your customers by looking somewhere else. 

This is an issue, and it's not raising enough eyebrows. The industry went from “bots are a problem” to “let's listen to the bots” in a very short period of time. The risk is an existential threat to a business, as the relationship they might have been building with their customers prior to using AI is where much of its unique selling point is developed. Businesses co-create the value with their customers, customers then become invested, and the business receives dividends in the form of brand equity. 

It’s also how you differentiate yourself from competition. You get additional insights from real people, rather than aim to satisfy the same archetype of a consumer your industry may be targeting. You make decisions on what areas of value you will pursue. It’s all internal knowledge you can leverage. 

Synthetic data is a mirage not worth chasing. It tends to inflate the number of ideas tested (and thereby reduce the average quality of these ideas) and deflate the attention of businesses invest in their customers and prospects. 

It's a dream supported by profit maximization and a simplistic view of consumers. In the long run, synthetic data will be seen for what it is: a waste of energy before the real insights are found.

The best way of testing ideas is going to market. Run a pre-order. Get a sign up page live. Have the product on a marketplace. The second best is to test ideas by engaging consumers. Do surveys. Conduct observations in the wild. The worst way is to move a step backward from the market and design AI clones to query. 

If you're considering using synthetic data, make sure to consider its costs. Not just in the dollar amount saved from research, but in the diversion of focus it may have on your business and its real customers and prospects in the long run. 

Businesses need to remember why they exist and where their success resides: in the closest and most aligned relationship possible with their customers. 

Next
Next

Why AI will not replace writers, a manifesto