Outsourcing Data Labeling: How to Get Accurate Annotations Fast and at Low Cost

RapidWorkers Team05 Oct 2026 • 11 min read
Outsourcing Data Labeling: How to Get Accurate Annotations Fast and at Low Cost

Even the most famous datasets contain labeling mistakes. A team of MIT researchers found an average of at least 3.3% wrong labels across the test sets of 10 popular datasets, and at least 6% in the ImageNet validation set (Source: Northcutt et al., 2021). If the biggest academic projects get it wrong, data labeling outsourcing without a quality system is almost guaranteed to produce more errors.

At the same time, having your own ML engineers label data is the most expensive option. An hour of an engineer's time costs as much as hundreds of outsourced labeled images, and repetitive work drains a team faster than it drains the budget. So the question usually isn't whether to outsource, but how to do it without losing accuracy.

This guide covers which labeling you can outsource, how an agency compares with freelancers and crowdsourcing, how to write instructions, how to keep accuracy high, and what it all costs. For a broader look at sourcing AI training data, see our article on AI training data collection.

What Labeling You Can Outsource — and What You Shouldn't

Labeling tasks vary enormously in difficulty. Almost any attentive person can mark whether a photo contains a cat. Only a doctor can mark a tumor on an MRI. That determines who should do the work and how.

Labeling type

What the annotator does

Difficulty

Who to give it to

Classification

Picks a category: "barrel sauna / no barrel sauna"

Low

Crowdsourcing

Bounding boxes

Draws a rectangle around objects

Low to medium

Crowdsourcing, if you have an annotation tool

Polygons and segmentation

Traces an exact outline or paints pixels

Medium to high

A trained team or an agency

Text: sentiment, topic, spam

Picks a label for a phrase or message

Low

Crowdsourcing with the right language

Audio transcription

Turns speech into text

Medium

Native speakers via crowdsourcing

Rating model responses

Compares two answers and picks the better one

Low to high

Crowdsourcing for everyday topics; experts for code and medicine

Expert annotation

Medical scans, legal documents

High

Domain specialists

Don't send data to open outsourcing if it can't be shown to outsiders: customer conversations, documents, medical records with personal data. Either anonymize it before sharing, or label it in-house or with a vendor under a confidentiality agreement.

Note: a complex task can often be split into simple ones. Instead of "label every object in this street photo," first ask "are there pedestrians in this photo?" and then "draw a box around each pedestrian." Each step is easier to explain, check and pay for.

Agency, Freelancers or Crowdsourcing: Which to Choose

There are three main outsourcing models. They differ less in price than in who owns quality and how fast you can start.

Factor

Labeling agency

Freelancers

Crowdsourcing via microtasks

Time to start

Weeks: brief, contract, pilot

Days: finding and vetting people

Hours: you can launch a task right away

Minimum budget

Usually thousands of dollars

Depends on the hourly rate

From $20 (the minimum deposit on RapidWorkers)

Scale

Large, but by agreement

Limited by the number of people

Grows quickly with many workers

Quality control

Handled by the agency, often with a contractual accuracy level

On your side

On your side: gold tasks, overlap, review

Complex and expert tasks

Yes, if the agency has specialists

Yes, if you find the right people

A poor fit

Best for

Long projects with clear requirements

Small complex tasks

Simple high-volume labeling, pilots, multilingual data

The key difference is who writes the rules and approves the work. An agency takes that on and includes it in the price. According to labeling provider CVAT, outsourcing to an agency typically costs more than 1.5 times as much as an in-house labeling team (Source: CVAT). With crowdsourcing, you set the rules and the review process yourself: it's cheaper and faster, but it takes a few hours of your time up front.

Pro Tip: you can combine models. Send simple high-volume labeling to crowdsourcing, disputed cases to an in-house expert, and expert work to a specialist vendor. That way expensive specialists only spend time where they're truly needed.

Annotation Guidelines: Your Main Accuracy Tool

Most labeling errors aren't carelessness — they're different readings of the rules. Does a half-hidden sauna count? What about a sauna on a billboard? If the instructions don't say, each annotator decides on their own, and each is right in their own way.

Good annotation guidelines have five parts:

  1. Why the labeling is needed. One sentence about what the model will do. An annotator who understands the goal resolves some disputed cases on their own.
  2. Precise class definitions. Not "sauna," but "a free-standing wooden sauna shaped like a barrel lying on its side."
  3. Examples for each class. At least two or three correct examples.
  4. Edge cases. A "case → decision" table: a partly hidden object, an image on a screen or poster, a blurry photo, several objects side by side.
  5. An "unsure" option. Let annotators flag a disputed case. It's more honest than a random guess and shows which rules need adding.

Here's an example of short instructions for a microtask:

Why: we're training a model to find barrel saunas in photos of country properties.
Task: for each photo, choose "Barrel sauna," "No" or "Unsure."
A barrel sauna is a wooden structure shaped like a barrel lying on its side. Examples: [3 photos].
Counts: the sauna is at least half visible; any season.
Doesn't count: a sauna in a drawing, ad or on a screen; a regular log sauna; a water barrel.
Unsure: choose this if the photo is too blurry or you can't tell what the structure is.
Pro Tip: treat the guidelines as a living document. Add every new edge case from review to the decisions table and number the versions. That way you always know which rules each batch was labeled under.

How to Keep Labeling Accurate

Good instructions reduce errors but don't eliminate them. Four control mechanisms take it from there. They aren't mutually exclusive and are usually used together.

Gold-standard tasks

Mix 5–10% of items with known correct answers into the task stream. Annotators don't know which ones they are. Accuracy on gold tasks shows accuracy on everything else. It's the cheapest way to spot careless workers right away.

Overlap (consensus)

Two or three people label the same item, and the final label is decided by majority vote. For classification, three answers are usually enough. Overlap multiplies the cost by two to three, but on simple labeling it stays low while accuracy improves noticeably. If annotators disagree, the item goes to an expert for review.

Inter-annotator agreement

How often annotators agree (inter-annotator agreement) reflects the quality of the instructions more than the quality of the people. It's usually measured with Cohen's kappa, and a common benchmark for production labeling is 0.8 or higher (Source: Label Your Data). If agreement is low across all annotators at once, the problem isn't the team — it's vague rules.

Expert spot checks

During the pilot, an expert reviews everything. After that, they check 10–20% of random items from each batch plus every item marked "unsure."

Mechanism

What it catches

Cost

Gold-standard tasks

Careless and random-answering annotators

+5–10% of volume

Overlap

One-off errors by individual annotators

×2–3 the cost

Agreement

Unclear rules in the instructions

Free when you use overlap

Expert review

Systematic errors and disputed cases

Your specialist's time

Note: reject work only for breaking a rule that's in the instructions. If an annotator got a case wrong that the rules didn't cover, approve the work and update the instructions. Your reputation as an employer shapes who takes your future tasks.

How Much Data Labeling Costs and How Fast You Can Get It

Typical market benchmarks from labeling providers: a bounding box costs around $0.02–0.04 per object, and a polygon or segmentation label starts at $0.06. Annotator hourly rates range from $4–6 in lower-cost regions to $30 or more for domain experts in North America and Europe. A typical mid-size agency project starts at $5,000–10,000 (benchmarks from the same Label Your Data guide).

With crowdsourcing there's no minimum engagement: on RapidWorkers, the employer sets the rate per task (from $0.05), a 15% fee is added on top, and you don't pay for rejected submissions. It works best to bundle 1–5 minutes of work into one task. Sample calculations:

Task

One task

Volume and overlap

Total incl. 15% fee

Per item

Image classification

10 photos, ~1 minute, $0.10

10,000 photos, ×3

$345

≈ $0.035

Bounding boxes

5 photos, ~3 minutes, $0.30

2,000 photos, ×2

$276

≈ $0.14

Text sentiment

20 messages, ~2 minutes, $0.20

5,000 messages, ×3

$172.50

≈ $0.035

Audio transcription

1 minute of audio, ~5 minutes, $0.40

1,000 clips, no overlap

$460

≈ $0.46

The rates in the table are examples: the harder and longer the task, the higher the pay should be. The math is simple: number of tasks × overlap × rate + 15%. For classification, for example: 1,000 tasks × 3 × $0.10 = $300 + 15% = $345.

Crowdsourcing speed depends on three things: the rate, the number of available workers in the countries you select, and task difficulty. Simple labeling with no country restrictions runs in parallel across many people, so batches often finish in hours rather than weeks.

Warning: the most expensive mistake is cutting overlap and gold tasks. Labeling without them is two to three times cheaper, but nobody catches the errors, and they go straight into your model. Finding and fixing them later costs more.

How to Speed Up Labeling with Models

In 2026, labeling is rarely done from scratch. More often, a model applies preliminary labels and people check and fix them. This approach is called human-in-the-loop, and it pairs well with microtasks.

Three workflows that work:

  1. Pre-labeling + review. A model draws boxes or assigns classes, and a worker answers "correct / incorrect" and fixes mistakes. Reviewing is several times faster than labeling from scratch.
  2. Active learning. Only the items the model is unsure about go to people. According to Label Your Data, this can cut manual labeling volume by 30–70%.
  3. Disagreement review. If an annotator's label conflicts with a confident model prediction, the item goes back for a second check. This catches both human errors and model errors.

A pilot plan

Run a small labeling round before the big one. It takes a day and costs a few dollars to a few dozen.

  1. Take 100–200 items and label them yourself — these become your gold answers.
  2. Launch a task on the same items with ×3 overlap.
  3. Compare accuracy against your answers, agreement between annotators, and time per task.
  4. Go through every disagreement and add the edge cases to the instructions.
  5. If accuracy and agreement meet your bar, keep some of the pilot items as gold tasks and launch the main volume in batches.
Pro Tip: if you collect data first and label it later, build the metadata you'll need for labeling into the collection step. We covered how to do that for video and speech in our guides to video data collection and speech data collection.

Data Labeling Outsourcing FAQ

How much does data labeling cost?

From providers, a bounding box usually costs around $0.02–0.04 per object, segmentation starts at $0.06 per label, and mid-size projects start at $5,000–10,000. With crowdsourcing there's no minimum engagement: in our examples, classification with triple overlap comes to about $0.035 per image, and bounding boxes with double overlap to about $0.14.

Is outsourcing data labeling worth it?

For simple high-volume labeling, almost always: ML engineers' time is too expensive to spend on repetitive work. Expert labeling and confidential data are better kept in-house or with a specialist vendor under contract.

How do you ensure data annotation quality?

Four tools work together: detailed guidelines with edge cases, gold-standard tasks with known answers, overlap with majority voting, and expert spot checks. Start with a pilot on 100–200 items.

What is a good inter-annotator agreement?

A common benchmark for production labeling is a kappa of 0.8 or higher. If it's lower across all annotators at once, the problem is most likely the guidelines, not the people.

Crowdsourcing vs a data annotation company: which is better?

Crowdsourcing wins on simple high-volume labeling, pilots and multilingual data: you can start within hours and there's no minimum engagement. A company wins on long, complex and confidential projects where you need a contract with guaranteed accuracy.

How do you write annotation guidelines?

Explain what the labeling is for, give precise class definitions with examples, describe edge cases as "case → decision," and add an "unsure" option. For a microtask, the guidelines should fit on one screen.

Accurate Labeling Comes from a System, Not a Vendor

People often expect that outsourcing labeling is just a matter of finding "good annotators." In practice, accuracy comes from instructions, gold tasks and overlap — and those work with any outsourcing model.

The key takeaways:

  1. send simple high-volume labeling to crowdsourcing, and expert or confidential work to specialists;
  2. split complex tasks into simple steps;
  3. the heart of good guidelines is edge cases and an "unsure" option;
  4. don't skimp on gold tasks and overlap — they cost less than errors in your model;
  5. start with a pilot on 100–200 items and your own gold answers.

If you want to test this approach on your own data, run a small pilot in the AI Tasks section on RapidWorkers: set your instructions, rate and overlap, and get your first results the same day.

data labeling outsourcingoutsource data annotationdata annotation servicesdata labeling costcrowdsourced data labelingannotation guidelinesinter-annotator agreementai training data

Ready to Get Started?

Join thousands of workers and employers already using RapidWorkers to get tasks done fast.

  • No subscription
  • Cancel anytime
  • 24/7 availability
  • Dispute resolution