Building Diverse Datasets: How Global Crowdworkers Reduce Bias in AI Models

A diverse dataset isn't a matter of image. It's a matter of whether your model works for all of your users. In the Gender Shades study (Buolamwini, Gebru, 2018), commercial gender classification systems misclassified darker-skinned women in up to 34.7% of cases, compared with at most 0.8% for lighter-skinned men. The cause was in the data: in two popular face datasets, 79.6% and 86.2% of the subjects had lighter skin.
And it's not an isolated case. A systematic review in The Lancet Digital Health (Wen et al., 2021) looked at open datasets for skin cancer diagnosis: out of more than 100,000 images, only 2,436 recorded skin color, and just 11 of those showed brown or dark skin. A model trained on data like that can't work reliably for a large share of the world's population.
Model bias almost always starts with who made it into the dataset and who didn't. Below, we cover where data bias comes from, how to plan for diversity up front, what crowdworkers from around the world bring, and how to check that your dataset has actually become more diverse.
Where Data Bias Comes From
"Bias" sounds intentional, but in data it almost always creeps in by accident, through how and from whom the data was collected. Here are the main sources of skew and what they look like in practice.
Source of bias | How it shows up | What to do about it |
|---|---|---|
Who's in the sample | Data comes from whoever is easiest to reach: employees, students, English-speaking internet users | Set quotas by group in advance and recruit people to fill them |
Devices and conditions | Every photo shot on an expensive phone in good light, every recording made in silence | Add devices and conditions to the coverage matrix |
Labeling | Annotators from one culture have their own idea of a "polite reply," an "aggressive gesture" or "home cooking" | Recruit annotators from the same groups as your users |
Context | An object is always shot in the same setting, so the model learns to recognize the background, not the object | Ask for the same object in different places and situations |
Historical data | The model learns from people's past decisions and repeats their mistakes | Collect fresh data for underrepresented cases instead of only cleaning old data |
All five have one thing in common: the skew stays invisible until you look at the data group by group. An overall accuracy of 95% can hide 70% for a small group, and that group will find out before you do — from their own experience using your product.
How to Plan for Diversity Before Collection Starts
Diversity doesn't happen on its own, even with plenty of workers. It has to be built into the collection plan, and that starts with one question: in what ways do your users differ that could affect the model?
Common dimensions of diversity:
- People: age, gender, skin tone, appearance features (glasses, beards, head coverings).
- Language: native language, dialect, accent, code-switching.
- Environment: home, street, store, city or village, lighting, noise.
- Devices: high-end and budget phones, webcams, microphones.
- Edge cases: things that are rare but critical for users: people with speech impairments, non-standard documents, unusual angles.
You don't need every dimension at once. Pick the ones that actually change what the model sees or hears. For gesture recognition, skin tone, age and lighting matter, but native language doesn't. For a voice assistant, it's the other way around.
Example quotas for 2,000 videos of hand gestures:
Dimension | Groups | Minimum per group |
|---|---|---|
Skin tone | Light / medium / dark | 25% |
Age | 18–30 / 31–50 / 51+ | 20% |
Lighting | Daylight / artificial / dim | 20% |
Workers | Cap per person | no more than 10 videos |
Set a minimum per group rather than an exact share. That way no group drops out of the dataset, and collection doesn't stall if one group fills faster than the others.
Note: look not only at each dimension separately but also at the key intersections. In Gender Shades, the largest error fell on the intersection of "women + dark skin." A dataset can be balanced by gender and, separately, by skin tone while containing almost no darker-skinned women.
What Crowdworkers from Around the World Bring — and What They Don't
Crowdsourcing is the fastest way to get beyond the "team members and their friends" circle. Workers live in different countries, speak different languages, use different phones and photograph their real homes and streets. For collecting photos, voice, video and text, that's exactly the diversity open datasets lack.
But crowds have their own skews, and you need to know them:
- A platform's audience isn't the world's population. According to a study of the Mechanical Turk workforce (Difallah et al., 2018), about 75% of workers were from the US, and workers overall skewed younger than the general population. On any platform, some countries and age groups are better represented than others.
- Offline groups are out of reach. People without a smartphone and a decent internet connection don't end up in an online crowd. If your model is meant for them, remote collection has to be supplemented with fieldwork.
- Active workers eat up the quota. Without a per-person cap, a handful of the fastest contributors will fill a noticeable part of the dataset with their own voice, kitchen and phone.
So "we collect through a crowd" doesn't mean "we have diverse data." The crowd gives you access to diversity, and four tools turn it into a balanced dataset:
- Geo-targeting: separate campaigns by country, so one large country doesn't crowd out the rest.
- Requirements in the task: "for people aged 50 and over," "shoot in evening light."
- A per-worker cap: so one person doesn't become the face of an entire group.
- Higher rates for rare groups: if a task for older participants fills slowly, the rate is usually the reason.
Warning: only ask for age, skin tone and other demographic attributes when you genuinely need them to balance the dataset. Make it optional, with a "prefer not to say" choice, and explain why you're asking. Data protection laws set special requirements for some of this information.
How to Check That Your Dataset Is More Diverse and Your Model Is Fairer
Diversity gets checked twice: first in the data, then in the model's behavior.
Data audit. After collection, count how many examples landed in each group and in the key intersections, and compare that with your quotas. Describe the dataset in a short document: who collected the data and where, which groups are represented and which are missing. Datasheets for Datasets (Gebru et al.) offers a ready-made structure for this kind of document.
Per-group metrics. Overall accuracy hides problems. Calculate metrics separately for each group in your coverage matrix and look at the gap between the best and worst group. If the gap is large, run a targeted top-up of data for the weak group. This requires a test set where every group is well represented; otherwise, the metric for a small group will be noise.
Dataset Diversity Is Becoming a Legal Requirement
For high-risk AI systems in the EU (hiring, credit scoring, education and more), Article 10 of the EU AI Act requires training, validation and test data to be relevant and sufficiently representative, to be examined for possible biases, and to account for the geographical, behavioral and functional context of use. After the Digital Omnibus amendments, these requirements for Annex III systems apply from 2 December 2027. For companies entering the EU market, a documented coverage matrix and per-group metrics help not only with quality but with compliance too.
What a Balanced Dataset Costs on RapidWorkers
On RapidWorkers, the employer sets the rate per task (minimum $0.05), and a 15% fee is added to the task cost. Geo-targeting and separate campaigns let you pay more where a group fills more slowly. Example for the same 2,000 gesture videos (rates are for illustration):
Campaign | Videos | Rate | Cost |
|---|---|---|---|
Main | 1,600 | $0.40 | $640.00 |
Participants aged 51+ | 400 | $0.60 | $240.00 |
All tasks | 2,000 | $880.00 | |
15% fee | $132.00 | ||
Total | $1,012.00 |
The difference between "whoever got there first" and a balanced dataset here is $92 including fees, for the higher rate for the rare group. That's usually cheaper than collecting more data and retraining the model after users start complaining.
Diverse Datasets FAQ
What is a diverse dataset?
It's a dataset that adequately represents every group of users and every condition the model will work in: age, skin tone, language, accent, lighting, devices. Which dimensions matter depends on the model's task.
Why does data make a model biased?
A model learns from what it sees. If a group is barely represented in the data, or only appears in certain conditions, the model will make more mistakes on it. In the Gender Shades study, the error gap between groups was more than fortyfold: 34.7% versus 0.8%.
How do you reduce AI bias at the data collection stage?
Define your diversity dimensions and a minimum for each group, cap the number of tasks per worker, use geo-targeting, and check model metrics separately for each group. For weak groups, run a targeted data top-up.
Does crowdsourcing make a dataset diverse automatically?
No. A crowd gives you access to people from different countries with different devices, but every platform has its own skews by country and age. Balance comes from quotas, geo-targeting and per-worker caps.
Does the law require diverse training data?
In the EU, for high-risk systems. Article 10 of the EU AI Act requires sufficiently representative data and checks for bias, and for Annex III systems these requirements apply from 2 December 2027. This isn't legal advice: check how it applies to your system with a lawyer.
Diversity Belongs in the Collection Plan, Not in a Fix Afterward
Model bias is cheaper to prevent than to fix. Define your diversity dimensions, set a minimum for every group, and evaluate the model by group, not just by overall accuracy. Crowdworkers around the world give you access to the right people and conditions, and quotas and caps turn that access into a balanced dataset.
To get started, launch a pilot campaign on RapidWorkers with a minimum for each group. For more on collecting specific types of data, see our guides to images and speech.
Ready to Get Started?
Join thousands of workers and employers already using RapidWorkers to get tasks done fast.
- No subscription
- Cancel anytime
- 24/7 availability
- Dispute resolution

