Blind data: Cheap methods of acquiring consumer data can be unrepresentative and unreliable
Data drives the modern economy. It is vital for strategic decision-making, market analysis, and product development, as well as underpinning AI recommendations.
But while data is widely available, it is often unreliable and biased. Knowing what existing and potential customers really think and how they act is becoming increasingly difficult.
Popular tools like online ad blockers, virtual private networks (VPNs), and incognito web browsers protect the privacy of consumers and make it harder to track their behaviour.
The result is that data which can be collected is often skewed towards less privacy-sensitive individuals, a cohort that is not representative of the wider population.
Regulations such as the EU’s General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) further underpin consumer rights when it comes to data privacy. Similar rules exist in other key markets, including China, India, Japan, and Brazil.
Why is so much consumer data unreliable?
Some third-party brokers skirt these rules, collecting information without consent from subjects and without offering them any compensation. However, the data they gather often suffers from the same issues of being unreliable and unrepresentative.
All this causes obvious problems – bad data delivers flawed insights. This can lead to poor business strategies, ranging from mispriced services and misdirected marketing campaigns to the development of products that people just aren’t interested in.
The problems around poor data quality are well-recognised, but how can companies address the issue?
Some may take matters into their own hands and run their own surveys, but they are not data market experts, might not do it well, and may well end up paying far more than they need to.
Leaving individual companies to work out what data they need and how to collect it is not an efficient answer to the problem.
Others may turn to consent-based models which provide compensation to those sharing their data. There are two main models in use, but each have their shortcomings.
‘Fixed compensation’ schemes offer a cheaper approach, as they provide uniform payouts to subjects. However, the relatively low level of compensation on offer means they tend not to capture more privacy-sensitive users.
This leaves companies back at square one with biased data that only reflects part of the customer base.
How to buy more reliable customer data
To get the best insights, companies need data that is representative, putting as much emphasis on quality as quantity.
In their search for more representative data, companies may turn to the ‘centralised optimisation model’, which tries to customise the level of compensation offered to subjects based on the level of their sensitivity to privacy issues.
This also has its flaws, as it can encourage consumers to inflate the amount of compensation they demand, creating an inefficient market in which companies overpay for data.
A better option is to develop a system that encourages subjects to get involved by compensating them at an acceptable level but does not encourage them to inflate their demands. This enables the collection of truly representative data at a fair price.
The mechanism that my colleagues and I have developed achieves this – providing for the transparent and consensual collection and trading of data, compensating subjects sufficiently to encourage them to participate, while remaining affordable for companies.
Under our model, once a request for data has been made to the platform, relevant subjects are sorted into pairs who look identical in terms of their data, but whose privacy concerns are different.
This process is called Random Sampling of Rolling Pairs.
It then compares the price they would demand to share their data under a system akin to a Vickrey price auction (also known as a second-price, sealed-bid auction).
A cheaper way to buy reliable data?
In our system, two parties privately disclose what price they demand for their data.
The lower bidder is then selected, but they are paid the higher price that was demanded by the other, losing bidder.
This process can be repeated until there is a large enough pool of subjects to satisfy the company seeking information.
The structure of this mechanism removes the incentive for people to inflate the price of their privacy constraints and ensures that people are compensated at a near-optimal cost from the data buyer’s perspective. At the same time, the data collected is unbiased and reliable.
This strikes a far better balance between cost, on the one hand, and the ability to collect accurate data that is in full compliance with GDPR and other regulations, on the other.
Can firms buy reliable data to train AI systems?
We have tested the model in simulations using real world data, showing that it can be done in practice.
This matters for individual companies, but it is also important for the future of AI systems that require high-quality data to become more reliable and to serve the best interests of businesses and consumers.
At the same time, regulators need to do more to develop and enforce appropriate regulations on data sellers and resellers to ensure there is a fair market mechanism.
By doing so, they can ensure that data markets become more efficient and more reliable.
That should ultimately deliver better results for the companies that rely on AI and for their end consumers.
Further reading:
How websites deceive users on data sharing
Four keys to using big data to unlock better strategy
Do online privacy regulations increase data sharing?
Ram Gopal is Professor of Information Systems Management and Head of the Gillmore Centre for Financial Technology. He teaches Generative AI and AI on the MSc Financial Technology and the MSc Business Analytics and Artificial Intelligence.
Learn more about harnessing AI to give your organisation a competitive advantage on the two-day Executive Education course AI Leadership programme at WBS London at The Shard.
Discover more about AI, fintech and finance. Receive our Core Insights newsletter via email or LinkedIn.
Select Warwick Business School as a preferred source of trusted insights in your Google searches.