Geo-Targeted Data Collection: How to Get AI Data from Specific Countries

RapidWorkers Team09 Oct 2026 • 8 min read
Geo-Targeted Data Collection: How to Get AI Data from Specific Countries

Geo-targeted data collection solves a simple problem: making sure your data comes from the countries where your model will actually be used. It sounds obvious, but getting it wrong can cost tens of percentage points in accuracy. When researchers tested speech recognition systems from Amazon, Apple, Google, IBM and Microsoft, they found an average word error rate of 35% for Black Americans versus 19% for white Americans. Same country, same language, different dialects — and nearly twice as many errors.

The gap between countries is just as wide. In a Facebook AI study, accuracy in recognizing household objects differed by 15–20% between the US and Somalia or Burkina Faso. Models generalize poorly to what they haven't seen, and if your users are in Indonesia, data from the US won't make up for that.

Below, we cover when geo-targeting really matters, why a country filter often isn't enough, how to split your budget across countries and how to check that your data came from where it was supposed to.

When You Need Geo-Targeting and When You Don't

Not every task depends on geography. If a worker is labeling cars in existing images, their country doesn't matter. But if they're creating new data from their own surroundings, their country accounts for half the value of the result.

Data type

What changes from country to country

Do you need geo-targeting?

Language, accent, dialect, background noise

Yes, almost always

Products, interiors, streets, documents, signs

Yes, if the model serves specific markets

Gestures, everyday scenes, road conditions

Yes, for everyday and street scenes

Register, slang, local services, code-switching

Yes, if you need natural local language

Labeling existing data

Almost nothing, unless the task needs language skills or local context

Usually not

Country ≠ Language ≠ Accent

The main trap with geo-targeting: it filters by country, but what you often need is a language or an accent. Those are different things, and they line up less often than you'd think.

  1. One country can have hundreds of languages. According to Ethnologue, Nigeria has 520 living languages. A "Nigeria" filter will bring you Yoruba speakers, Hausa speakers and people who speak only English.
  2. One language lives in many countries. Spanish in Spain, Mexico and Argentina differs in pronunciation, vocabulary and forms of address. If your assistant is built for Mexico, a "Spanish-speaking countries" filter is too broad.
  3. Accent doesn't depend on country alone. The study in the introduction found a gap within a single country. Region, age and background shape speech as much as borders do.
  4. People move. A Tagalog speaker living in the UAE works for a speech dataset, but not for photos of Philippine stores.

That's why geo-targeting is the first filter, not the only one. The second filter goes into the task itself: "for native Yoruba speakers only," "shoot only in stores in your city." The third comes during review, covered below.

Note: ask workers to state their native language and region in every submission. Even if you don't need that data now, it lets you check your dataset's balance later and find which groups the model struggles with.

How to Split Data Across Countries

The first instinct is to split data in proportion to each country's share of users. That works poorly: a country with 5% of your users gets so little data that the model may never work well there. A more reliable approach: start with a guaranteed minimum for every country, then split the rest in proportion to user share.

Example: a voice assistant needs 3,000 voice recordings across four markets.

Country

User share

Proportional

Minimum + weighting

Nigeria

55%

1,650

1,170

Kenya

25%

750

750

Ghana

15%

450

610

Uganda

5%

150

470

Total

100%

3,000

3,000

In the second option, every country gets at least 400 recordings, and the remaining 1,400 are split by user share. Uganda gets three times as much data as it would under a proportional split, while Nigeria still makes up the largest part of the dataset. The right minimum is best tested in a pilot: train the model and look at its metrics for each country separately.

One Rate for Every Country or Different Rates?

Running a separate campaign per country lets you set different rates. Use that carefully:

  1. Raise the rate where workers are scarce. If a campaign in one country fills slowly, the rate is usually the reason.
  2. Don't cut the rate down to the bare minimum to match local wages. A task that pays too little attracts people who rush it, and your savings end up going to rejected submissions.
  3. Account for how hard the task is in each country. Photographing a chain supermarket in a capital city is easier than finding the same kind of store in a country where they're rare.
Pro tip: launch your pilot in all countries at once with the same rate. Fill speed and rejection rates by country will show you where to raise the rate and where to improve the instructions.

How to Check That Data Really Comes from the Right Country

Any country filter can, in theory, be bypassed, for example with a VPN. So it's worth checking the country against the data itself too. The good news: data gives away its origin if you know where to look.

Data type

Signs of the country

How to check

Photos of products and stores

Local brands, price tags in local currency, language on packaging

Spot-check manually, reverse image search

Voice

Language, accent, local names in free speech

Automatic language identification + review by a native speaker

Text

Slang, local services and references, date and number formats

Native-speaker review of a sample

Video and street scenes

Signs, road signs, license plates, left- or right-hand traffic

Manually check a few frames

Another signal is a worker's track record. If the same person sends photos of stores with different currencies or speaks with an accent that doesn't match their country, review all of their submissions.

Warning: don't collect EXIF geolocation or exact addresses just for verification. That's personal data, and storing it creates obligations for you under data protection laws. The country and city the worker provides, plus a content check, are enough.

Then there's the question of who does the checking. If no one on your team speaks Swahili, there's no one to check the Swahili recordings either. The fix is a second campaign with the same geo-targeting: other workers from the same country listen to the recordings and answer simple questions like "Is this Swahili?" and "Was the phrase read without mistakes?"

How to Set Up Country-Based Collection on RapidWorkers

On RapidWorkers, you set geo-targeting when you create a campaign: the task is visible only to workers in the countries you select. Here's a workflow that works:

  1. One country, one campaign. Each country gets its own task limit and rate, and one large country can't eat up the whole budget.
  2. Language and location requirements in the task. "For native Swahili speakers only," "shoot in stores in your city."
  3. Metadata with every submission. Native language and region or city, as stated by the worker.
  4. A pilot in all countries at once with the same rate, then adjust rates and instructions.
  5. Native-speaker review as a separate campaign with the same geo-targeting, if your team doesn't include speakers of the language.

The employer sets the rate per task, with a $0.05 minimum, and a 15% fee is added to the task cost. The minimum deposit is $20 and the minimum campaign is $0.30. Rejected submissions aren't paid, and if you end a campaign early, the unused budget goes back to your balance.

Example calculation for the same 3,000 voice recordings (rates are for illustration; let's assume the task filled more slowly in Uganda during the pilot, so the rate there was raised):

Country

Recordings

Rate

Task cost

Nigeria

1,170

$0.25

$292.50

Kenya

750

$0.25

$187.50

Ghana

610

$0.25

$152.50

Uganda

470

$0.30

$141.00

All tasks

3,000


$773.50

15% fee



$116.03

Total



$889.53

Native-speaker review of 10% of the recordings (300 tasks at $0.05) adds another $15.00 + $2.25 in fees = $17.25.

Geo-Targeted Data Collection FAQ

What is geo-targeting in AI data collection?

It's a setting that makes a data collection task visible only to workers in the countries you select. That way, your dataset reflects the languages, accents, products and conditions of the markets where your model will run.

Is a country filter enough to get the right language?

No. One country can have many languages (Nigeria has more than 500), and one language sounds different from country to country. Add a language requirement to the task and a native-speaker review on top of geo-targeting.

How much data should you collect from each country?

A reliable approach is a guaranteed minimum for every country plus the rest split in proportion to user share. How large that minimum should be will show up in the model's per-country metrics after a pilot.

How can you make sure a worker is really from the right country?

Check the content: local brands and currency in photos, accent and language in recordings, local references in texts. If something doesn't add up, review all of that worker's submissions. You don't need to collect precise geolocation for this.

Should you pay workers in different countries different rates?

Not necessarily. Start with the same rate in the pilot and raise it where the task fills slowly. Cutting rates sharply in lower-wage countries isn't worth it: the savings usually go to rejected submissions.

Geo-Targeting Is the First Filter, Not the Only One

A model works where its data looks like reality. Geo-targeting brings you the right people, the task pins down the language and location, quotas keep one country from crowding out the rest, and content checks confirm the data came from where it should. Together, these four steps help close an accuracy gap that changing the model's architecture won't fix.

If your model serves several markets, launch country-by-country pilot campaigns on RapidWorkers. We covered the overall process of collecting training data in our guide to AI training data collection.

geo-targeted data collectionai training datamultilingual data collectionspeech datalocalized datasetscrowdsourcingdata diversityai data collection

Ready to Get Started?

Join thousands of workers and employers already using RapidWorkers to get tasks done fast.

  • No subscription
  • Cancel anytime
  • 24/7 availability
  • Dispute resolution