Speech Data Collection: How to Build a Multilingual Speech Dataset

RapidWorkers Team04 Oct 2026 • 15 min read
Speech Data Collection: How to Build a Multilingual Speech Dataset

Speech Data Collection: How to Build a Multilingual Speech Dataset with Crowdworkers

An hour of speech data collected to your spec costs $150 to $400 or more from a vendor, and off-the-shelf datasets run $50–150 per hour (Source: SpeechData.ai). Yet speech data collection comes down to a simple action: a person picks up a phone and talks. The real questions are who talks, in what language, under what conditions, and with whose consent.

Speech models usually work well on American-accented English in a quiet room. The trouble starts beyond that: Vietnamese, Chilean Spanish, speech in a car, a sentence that starts in one language and ends in another. That's exactly the data open corpora lack — and exactly the data that's easiest to collect from real people in the right countries.

This guide covers how to build a multilingual speech dataset with microtasks: the types of speech data, how many hours and speakers you need, which recording parameters to set, what an hour of audio costs with crowdsourcing, and what consent you need to record someone's voice. For a broader overview of sourcing training data, see our article on AI training data collection.

Types of Speech Data

Not all speech is equally useful for training. A text-to-speech model learns from clean reading by a single voice, while speech recognition for a call center learns from messy conversations full of interruptions. Pick the type you need before you start collecting.

Speech type

How it's recorded

What it's used for

Pros and cons

Scripted reading

A person reads given sentences from a screen

Speech recognition (ASR), text-to-speech (TTS), rare words and terms

The text is known in advance, so no transcription is needed. But the speech sounds unnatural

Spontaneous speech

A person talks freely on a given topic

ASR for natural speech, voice assistants

Natural, but needs transcription

Two-person conversation

Two people chat about everyday topics

Conversation recognition, speaker diarization, voice agents

The most valuable data, but both speakers must consent

Short commands

A person says phrases like "turn on the lights"

Voice control, wake words

Fast and cheap, but narrow use

Code-switched speech

A person or a pair switches between languages

ASR for bilingual markets

Almost absent from open data, and harder to transcribe

Collecting speech from large numbers of people online has long proven itself. Since 2017, Mozilla's Common Voice has been building an open speech corpus with volunteers who read sentences and validate each other's recordings. A microtask platform works on the same principle, but with pay, your own topics and country targeting.

Note: for recognizing natural speech, scripted reading doesn't replace spontaneous speech. People read more evenly than they talk — without pauses, slips or "um." A model trained only on reading makes noticeably more mistakes on real conversations.

How Much Speech Data You Need: Hours and Speakers

The answer depends on whether you're training a model from scratch or fine-tuning an existing one — for example, Whisper or another open speech model. Almost every team today takes the second route.

Task

Volume guide

Narrow fine-tuning on a widely spoken language

Tens of hours

Fine-tuning for several accents and noisy conditions

Low hundreds of hours

A balanced dataset for most applied tasks

100–300 hours

Training a model from scratch

Thousands to tens of thousands of hours

Source: Spirelight. The same guide makes an important point: a hundred hours from many different speakers, devices and recording conditions often beats a thousand homogeneous hours.

That's good news for crowdsourcing. A microtask platform is a poor fit for one voice actor reading a hundred hours, but it handles "10 minutes each from 600 different people" very well.

How to plan your speaker sample

Count quotas for each group, not just hours. For speech data, these usually matter:

  1. Language and region. Spanish sounds different in Mexico, Argentina and Spain, and the model needs to hear every variety it will face.
  2. Gender and age. Children's and older voices are recognized worst because they're underrepresented in data. Children must not be recorded through open microtasks.
  3. Native and non-native speakers. English with Indian, Nigerian or Filipino accents are separate groups with their own quotas.
  4. Acoustic conditions. A quiet room, the street, a car, a kitchen with appliances running — in the proportions your product will encounter them.
  5. Devices. Microphones on cheap and expensive phones, headsets and laptops all sound different.
Pro Tip: split your test set by speaker, not by file. If one person's voice ends up in both training and test, your word error rate (WER) will look lower than it really is.

Audio Specs and Metadata

Phone recordings work for most speech tasks, as long as you define the parameters up front. Otherwise half the files will arrive compressed through a messaging app, and some will have music in the background.

Parameter

Typical requirement

Why it matters

Format

WAV or FLAC, or the original file from the phone's voice recorder

Messaging apps and social media re-compress audio with loss

Sample rate

At least 16 kHz; 44.1 or 48 kHz for speech synthesis

16 kHz is enough for recognition; synthesis needs more detail

Channels

Mono; for conversations, ideally a separate track per speaker

Separate tracks make transcription and diarization easier

Duration

A specific range, e.g. "at least 10 minutes of actual speech"

Otherwise the share of pauses and silence will vary widely

Environment

A quiet room or deliberately specified noise (street, car)

Random noise hurts; planned noise improves model robustness

Background sounds

No music, TV or other people's voices

Background speech ends up in the transcript and confuses the model

Processing

No effects, editing or speed changes

The model needs to learn from natural sound

Alongside the files, collect minimal metadata through the task form. Without it you can't check the balance of your sample:

  1. native language and the region where the person grew up;
  2. age group (18–25, 26–40 and so on), not an exact date of birth;
  3. gender, if it matters for the task;
  4. device and recording location.

Names, phone numbers and addresses aren't needed to train a model, so don't collect them.

Pro Tip: if you're already collecting video with speech, the audio track can double as speech data — as long as the consent covers it. We covered how to set recording requirements in our guide to video data collection for AI training.

How to Collect Multilingual and Code-Switched Speech

The golden rule of multilingual collection: one language or region per task. That way you can set different rates, track progress separately, and avoid ending up with 90% of recordings in the most common language.

On RapidWorkers, geo-targeting handles this: only workers from the countries you select can see the task. But country isn't the same as native language, so add two more filters:

  1. A native-language question in the task form. Simple self-reporting screens out most of the wrong workers.
  2. Automatic language identification during review. Off-the-shelf language ID models quickly flag recordings made in the wrong language.

Code-switched speech

In bilingual countries, people switch languages constantly within one conversation and even one sentence: Vietnamese with English technical terms, English with Tagalog or Hindi mixed in. Models trained on monolingual data make especially many mistakes on this kind of speech.

To collect it, say plainly in the task that switching languages is allowed and wanted. In front of a microphone, people instinctively try to speak "properly," in one language. Give topics where mixed speech comes up naturally: work, tech, school, online shopping.

Note: only someone fluent in both languages can transcribe code-switched speech. Set transcription up as a separate task with the same geo-targeting, and agree up front on how to tag the language of each segment.

How to Launch Speech Data Collection with Microtasks

The process is similar to any data collection, but speech has its own nuances at every step.

  1. Define the speech type and the transcription you need. For scripted reading, the text already exists. For spontaneous speech and conversations, decide right away who will transcribe the recordings and how.
  2. Prepare topics or sentences. For spontaneous speech, 10–20 everyday topics for people to choose from. For reading, a set of sentences with the terms you need, split into batches so different people read different text.
  3. Choose countries for each language. One task per language or region.
  4. Record an example yourself. 30 seconds of a correct recording and 30 seconds with a typical mistake — for example, a TV playing in the background.
  5. Run a pilot with 10–20 people. Listen to every recording and check the file format: this is usually where compressed files from messaging apps show up.
  6. Fix the task and scale in batches. After each batch, check the balance by gender, age and region, and fill the gaps.

Here's what a spontaneous speech task might look like:

Why: the recordings will be used to train a speech recognition model and stored for no more than 3 years.
What to do: record 10 minutes of speech in Vietnamese, talking about 2–3 topics from the list.
Topics: how your day went; your work or studies; your last online purchase; your favorite dish.
Allowed: using English words if that's how you normally talk.
Conditions: a quiet room, phone 20–30 cm from your face.
Don't: read text, play music or TV, record other people.
How to submit: the original file from your phone's voice recorder + a short form (native language, region, age group).
Examples: [good recording] [recording with a mistake]
Pro Tip: for conversations, ask the worker to record a chat with an adult they know rather than a stranger. The conversation sounds more natural, and it's easier to get the second participant's consent. More on consent below.

How Much an Hour of Speech Data Costs

Here's how the market looks: open corpora are free but often don't allow commercial use; off-the-shelf datasets cost $50–100 per hour for popular languages and $60–150 for rare ones; custom vendor collection costs $150–400 and up.

With crowdsourcing, the math works differently: you pay per task, and an hour of audio is made up of several tasks. On RapidWorkers, the employer sets the rate (minimum $0.05), a 15% fee is added on top, and you don't pay for rejected submissions. Here's a sample calculation per approved hour of audio:

Scenario

Recording

Transcription

Total per audio hour incl. 15% fee

Scripted reading: 5 minutes of speech per task, $0.75

12 tasks × $0.75 = $9

Not needed, text is known

≈ $10

Spontaneous speech: 10 minutes per task, $2.00

6 × $2.00 = $12

6 tasks at $4.00 = $24

≈ $41

Bilingual conversation: 10 minutes per task, $3.00

6 × $3.00 = $18

6 tasks at $6.00 = $36

≈ $62

The rates in the table are examples. They're based on the worker's real time: 10 minutes of speech plus setup and upload takes about 20 minutes, and transcription usually takes several times longer than the recording itself. Rare languages and narrow country targeting call for higher rates.

Note: comparing these numbers with vendor prices isn't apples to apples. Vendors' $150–400 covers project management, quality control, consent paperwork and annotation. With crowdsourcing, your team does that work, and its cost is your team's time. Even so, for pilots and mid-sized volumes, crowdsourcing is usually much cheaper.

A planning example: 100 hours of spontaneous speech with transcription means 600 speakers at 10 minutes each and roughly $4,100 at the rates in the table. A pilot with 20 speakers without transcription costs about $46 (20 × $2.00 + 15%).

How to Check Audio Quality

Audio is easier to check than video: scripts catch most defects before a person ever listens. Here are the checks, from cheapest to most expensive:

  1. Technical parameters. Format, sample rate, duration, clipping and loudness.
  2. Share of speech. A voice activity detector shows how much real speech a recording contains. "10 minutes" with six minutes of silence doesn't meet the task.
  3. Noise. A signal-to-noise estimate separates quiet recordings from ones made on a busy street, if you didn't ask for those.
  4. Language. Automatic language identification catches recordings in the wrong language. For code-switched speech, check that both languages are present.
  5. One person, one speaker. Speaker verification models catch cases where the same person submitted recordings from several accounts. For speaker diversity, this matters more than it seems.
  6. Duplicates. Audio fingerprinting finds identical or re-edited recordings.
  7. Manual review. A person listens to clips from the beginning, middle and end: is the speech natural, is the speaker reading text, are there other voices?

Transcripts need checking too. A simple method is to run the recording through an off-the-shelf speech recognition model and compare its output with the worker's transcript. A large mismatch doesn't always mean the person made a mistake, but it shows which files to re-listen to first.

Pro Tip: reject a recording only for breaking a rule from the task, and use the same wording every time: "music in the background," "less than 10 minutes of speech," "file from a messaging app." Workers learn faster, and you see which part of the instructions needs clarifying.

Consent: A Voice Is Personal Data

A voice recording is personal data, and a voiceprint that can identify someone is biometric data. In the EU, this falls under the GDPR. In the US, the strictest law is Illinois's BIPA: it explicitly lists voiceprints as biometric identifiers and requires written consent, including electronic consent. In lawsuits under the law, damages start at $1,000 per negligent violation and $5,000 per intentional or reckless one, per person (Source: Enzuzo, 2026).

Speech data has a quirk that's often overlooked: a conversation involves two people, but only one of them accepts the task. The second speaker isn't registered on the platform and hasn't signed anything. So build these into every speech task:

  1. The purpose and retention period in the first line of the task. What the recordings will be used for and how long they'll be kept.
  2. The worker's written consent. A separate checkbox in the task form. A spoken phrase in the recording isn't enough for laws like BIPA.
  3. The second speaker's consent. For conversations, require both participants to be adults and both to consent. The most reliable option is a separate consent form the worker uploads with the recording.
  4. No bystanders. A ban on recording children, phone calls without warning, and people on the street.
  5. No personal details in the speech. Ask speakers not to say full names, addresses, phone or card numbers in the recording.
  6. Deletion on request. Keep the link between each file and its task ID so you can delete a specific person's recordings.
Warning: this section is a general overview, not legal advice. If you're collecting voices from EU or US residents for a commercial product, have a lawyer review your consent wording before launch.

Speech Data Collection FAQ

How much speech data is needed to train ASR?

Fine-tuning an existing model usually takes tens to low hundreds of hours, and 100–300 balanced hours cover most applied tasks. Training from scratch takes thousands of hours. Diversity of speakers, devices and recording conditions matters more than the raw number of hours.

How many hours of audio do you need to fine-tune Whisper?

There's no universal number. For a widely spoken language and a narrow domain, noticeable gains often show up at tens of hours. New accents, a rare language or noisy conditions need more. The most reliable approach is to collect a pilot batch, measure WER on a test set, and keep adding data while it keeps dropping.

How much does a speech dataset cost per hour?

From vendors, off-the-shelf datasets cost $50–150 per hour and custom collection $150–400 and up. With crowdsourcing, in our examples an hour of scripted reading costs about $10, spontaneous speech with transcription about $41, and a bilingual conversation about $62 including the fee. Your team handles quality control and consent paperwork.

How do you record speech data for AI?

You need a quiet room (or deliberately specified noise), a phone 20–30 cm from the speaker's face, and a standard voice recorder app. The original file is uploaded directly, not forwarded through a messaging app. 16 kHz is enough for speech recognition; 44.1–48 kHz is better for speech synthesis.

Is it legal to collect voice data for AI training?

Yes, if you have written consent from every person in the recording — including the second speaker in conversations — and you follow local personal data laws. State the purpose and retention period in the task and allow recordings to be deleted on request.

Why collect your own data instead of using open speech datasets?

Open corpora like Common Voice are great for getting started, but they consist mostly of read speech and don't always cover the accents, topics and conditions you need. Your own collection closes exactly the gaps where your model makes mistakes.

A Good Speech Dataset Is About People, Not Hours

Speech models fail where the data lacked the right voices: accents, ages, conditions, switches between languages. Those gaps are easiest to close by recording many different people in the right countries.

The key takeaways:

  1. fine-tuning usually takes tens to hundreds of hours, as long as they come from many different speakers;
  2. one language or region per task, and explicitly allow code-switching in the task when you need it;
  3. require original files and collect minimal metadata to check your sample balance;
  4. calculate the cost per hour of approved audio, transcription included;
  5. you need consent from every voice in the recording, not just the worker's.

If you want to test this approach in your language, run a small pilot in the AI Tasks section on RapidWorkers: pick your countries, set a rate, and get your first recordings to review.

speech data collectionmultilingual speech datasetaudio data collection for aiasr training dataspeech dataset for machine learningvoice data collectioncode-switchingai training data

Ready to Get Started?

Join thousands of workers and employers already using RapidWorkers to get tasks done fast.

  • No subscription
  • Cancel anytime
  • 24/7 availability
  • Dispute resolution