Speech Data for AI

Speech Data Collection
for AI Training

Collect voice recordings from native speakers in any country. Get audio with transcripts in one task for ASR, TTS and voice AI.

Native speakers worldwideAudio and transcripts in one taskReview every recording

190+

Countries

13,487

Tasks completed

12,000+

Registered workers

~4 Hours

Avg time to results

How It Works

How Speech Data Collection Works

Launch a speech task yourself and get voice recordings from real people without long negotiations.

Person creating a speech task with phrases to record
01 · Create

Describe what people should record

Add a script to read, a topic to talk about or a list of voice commands. Ask for a transcript with each recording if you need one.

  • Scripts, topics or voice commands
  • Transcripts on request
  • Your own instructions and examples
World map with selected target countries for speakers
02 · Target

Choose countries and speakers

Select the countries you need, set the number of recordings and describe speaker requirements such as language, accent or age in your task.

  • Any number of countries in one task
  • Speaker requirements in your instructions
  • You set the number of recordings
Native speaker recording voice on a smartphone
03 · Record

Native speakers record on their own devices

Your task becomes available to workers in the selected countries. They record audio on their phones and submit it together with the transcript.

  • Native speakers in your target markets
  • Recorded on real phones
  • Audio and transcript in one submission
Reviewing submitted audio recordings with approve or reject actions
04 · Review

Approve only the recordings that fit

Listen to each submission, check the transcript and reject anything that does not meet your requirements. Keep a clean dataset for training.

  • Listen to every recording
  • Approve or reject each submission
  • Pay only for what you accept
Speech Data Types

Types of Speech Data You Can Collect

Set up a task for the exact type of recordings your model needs. Workers follow your instructions and record on their own phones.

Scripted Speech

Every recording matches your script

Workers read sentences, phrases or numbers from your script. You know exactly what was said in every recording, so the data is ready for training without guesswork.

  • Your own sentences, phrases or numbers
  • Same script read by many different speakers
  • Text is known in advance for every recording
Create a speech task
Woman at home reading from a script while recording on her smartphone
Spontaneous Speech

Natural speech the way people really talk

Workers speak freely on a topic you set, like describing their day or giving an opinion. You get real phrasing, pauses and fillers that scripted data cannot give you.

  • Topics and questions you define
  • Natural pauses, fillers and real phrasing
  • Transcript added to every recording
Create a speech task
Man in a café talking naturally while recording a voice message on his phone
Voice Commands

Commands and wake words from thousands of voices

Workers repeat short commands and trigger phrases in their own voice and manner. Collect the variety your voice assistant needs to respond to anyone.

  • Wake words, commands and short phrases
  • Many repetitions from different speakers
  • Different ages, genders and voices
Create a speech task
Person in a kitchen speaking a short command toward a smartphone on the counter
Dialogues

Real conversations between two people

Workers record a conversation with a friend or family member, following roles or a free topic. Both participants agree to the recording.

  • Role-based or free conversations
  • Real back-and-forth speech
  • Consent from both participants
Create a speech task
Two friends at home talking while a smartphone between them records the conversation
Accents and Dialects

Cover the accents your users actually have

Collect the same phrases from speakers in different countries and regions. Your model learns to understand people beyond one standard accent.

  • Speakers from the countries you choose
  • Same phrases across different regions
  • Native speakers of each language
Create a speech task
People in different countries recording voice messages on their smartphones
Noisy Environments

Speech recorded in real-world conditions

Workers record on the street, in a car, in a café or at home with background sounds. Your model learns to hear speech where people actually use it.

  • Street, car, café or home settings
  • Real background noise, not added later
  • Location type noted for each recording
Create a speech task
Woman on a busy city street recording a voice message on her smartphone
Task Example

Recordings With Transcripts
in One Task

Ask workers to send both the audio and the text of what they said. You get ready training pairs without a separate transcription step.

Your setup

What you set up

Task settings

Choose Audio collection, set how many files each worker sends and add instructions with phrases to read and recording rules.

  • Audio collection task type
  • Number of files per submission
  • Instructions with your script

Your view

What you receive

Submission

Each submission arrives as an audio file with the transcript attached, ready to review and approve.

  • Audio and text in one submission
  • Worker country on every submission
  • Approve or reject each pair
Targeting

Target the Right Speakers

Choose where your speakers come from and tell them exactly how to record. Every submission comes from a different person.

Countries and regions

Accept workers from all countries or pick specific ones by continent or individually. Only people from the selected countries can take your task.

Speaker requirements

Describe who should record in your instructions: native language, accent, age group or gender. Reject submissions that do not match.

Example recordings

Upload a sample recording, a script or a short video so workers hear exactly what you expect before they start.

One speaker, one submission

Each worker can complete your task only once. Set the number of recordings you need and get them from that many different speakers.

Quality Control

Quality Control You Stay in Charge Of

You decide what counts as a good recording and review every submission before it joins your dataset.

Person typing step-by-step recording instructions on a laptop

Clear instructions

Describe exactly what workers must record: phrases, length, tone and recording conditions.

Laptop showing an audio waveform with headphones on a wooden desk

Example recordings

Upload sample audio or a short video so workers hear what a good result sounds like.

Person with headphones reviewing audio recordings with approve and reject marks

Approve or reject

Listen to each submission and accept only recordings that meet your requirements.

Woman with headphones checking audio and text on a laptop

Second review

Launch a separate task where other workers check recordings and transcripts for you.

Transparency

Collected With Consent

Every recording comes from a real person who chose to take your task and knew what it was for. No scraped audio and no unknown sources.

Workers know the purpose

Your task description explains what people will record and that it will be used to train AI. Workers read it before they accept.

Voluntary participation

Nobody is assigned to your task. Each worker decides to take it and submits their recording on their own.

Known source of every file

Each submission comes from a specific worker in the country you selected, recorded specifically for your task.

Pricing

Example Budgets for Speech Data

You set the pay per recording. Here is what typical tasks cost with the 15% platform fee included.

Smartphone on a desk showing a voice recording screen with an audio waveform

20 recordings

Test run

Check quality and instructions on a small batch before you scale.

$4.60

$0.20 per recording

Create task
Printed list of phrases next to a smartphone recording audio

100 recordings

Scripted speech

Phrases read aloud from your script by 100 different speakers.

$23

$0.20 per recording

Create task
Headphones and a smartphone with an audio waveform on a dark wooden table

300 recordings

Spontaneous speech

Free speech on your topic with a transcript in every submission.

$103.50

$0.30 per recording

Create task
Several smartphones on a desk showing audio waveforms and country flags

1,000 recordings

Accent coverage

The same phrases from speakers across many countries and regions.

$287.50

$0.25 per recording

Create task

Prices are examples. You choose the pay rate, starting from $0.10 per task.

FAQ

Frequently Asked Questions

Speech data collection is gathering voice recordings from real people to train and test AI models such as speech recognition, text-to-speech and voice assistants. Good datasets include many different speakers, accents and recording conditions.

You set the pay for each recording, starting from $0.10, plus a 15% platform fee. A campaign starts from $0.30, so you can run a small test before scaling.

Workers can start as soon as your campaign is live. Speed depends on the countries you choose, your pay rate and how complex the task is. A small test run is the quickest way to check timing and quality.

Yes. Ask workers in your submission instructions to type exactly what they said, and each submission will include both the audio and the text.

Each worker can complete your task only once, so every submission comes from a different speaker. If you need several phrases from one person, set the number of files per submission.

You can accept workers from all countries or choose specific continents and individual countries when you create the task.

Describe the speakers you need in your task instructions, for example native Hindi speakers aged 25–40. You review every submission and can reject recordings that do not match.

Any you need. Describe the file format, recording length and conditions in your submission instructions, and workers follow them when recording.

Yes. Upload up to 30 reference files, such as a sample recording, a script or a short instructional video, so workers hear what a good result sounds like.

You listen to every submission and approve or reject it. For larger projects you can also launch a separate task where other workers check recordings and transcripts.

Yes. Extend your existing task instead of creating a new one. This keeps all recordings in one place and still lets each worker submit only once.

If you need a different file type, a special review flow or help with your instructions, contact us through the form below and we will set it up with you.

Contact

Planning a Large Speech Data Project?

Need a custom file type, a special review flow or help with your instructions? Send a message and we will set it up with you.