Outsourcing Data Labeling: How to Get Accurate Annotations Fast and at Low Cost

Even the most famous datasets contain labeling mistakes. A team of MIT researchers found an average of at least 3.3% wrong labels across the test sets of 10 popular datasets, and at least 6% in the ImageNet validation set (Source: Northcutt et al., 2021). If the biggest academic projects get it wrong, data labeling outsourcing without a quality system is almost guaranteed to produce more errors.
At the same time, having your own ML engineers label data is the most expensive option. An hour of an engineer's time costs as much as hundreds of outsourced labeled images, and repetitive work drains a team faster than it drains the budget. So the question usually isn't whether to outsource, but how to do it without losing accuracy.
This guide covers which labeling you can outsource, how an agency compares with freelancers and crowdsourcing, how to write instructions, how to keep accuracy high, and what it all costs. For a broader look at sourcing AI training data, see our article on AI training data collection.
What Labeling You Can Outsource — and What You Shouldn't
Labeling tasks vary enormously in difficulty. Almost any attentive person can mark whether a photo contains a cat. Only a doctor can mark a tumor on an MRI. That determines who should do the work and how.
Labeling type | What the annotator does | Difficulty | Who to give it to |
|---|---|---|---|
Classification | Picks a category: "barrel sauna / no barrel sauna" | Low | Crowdsourcing |
Bounding boxes | Draws a rectangle around objects | Low to medium | Crowdsourcing, if you have an annotation tool |
Polygons and segmentation | Traces an exact outline or paints pixels | Medium to high | A trained team or an agency |
Text: sentiment, topic, spam | Picks a label for a phrase or message | Low | Crowdsourcing with the right language |
Audio transcription | Turns speech into text | Medium | Native speakers via crowdsourcing |
Rating model responses | Compares two answers and picks the better one | Low to high | Crowdsourcing for everyday topics; experts for code and medicine |
Expert annotation | Medical scans, legal documents | High | Domain specialists |
Don't send data to open outsourcing if it can't be shown to outsiders: customer conversations, documents, medical records with personal data. Either anonymize it before sharing, or label it in-house or with a vendor under a confidentiality agreement.
Note: a complex task can often be split into simple ones. Instead of "label every object in this street photo," first ask "are there pedestrians in this photo?" and then "draw a box around each pedestrian." Each step is easier to explain, check and pay for.
Agency, Freelancers or Crowdsourcing: Which to Choose
There are three main outsourcing models. They differ less in price than in who owns quality and how fast you can start.
Factor | Labeling agency | Freelancers | Crowdsourcing via microtasks |
|---|---|---|---|
Time to start | Weeks: brief, contract, pilot | Days: finding and vetting people | Hours: you can launch a task right away |
Minimum budget | Usually thousands of dollars | Depends on the hourly rate | From $20 (the minimum deposit on RapidWorkers) |
Scale | Large, but by agreement | Limited by the number of people | Grows quickly with many workers |
Quality control | Handled by the agency, often with a contractual accuracy level | On your side | On your side: gold tasks, overlap, review |
Complex and expert tasks | Yes, if the agency has specialists | Yes, if you find the right people | A poor fit |
Best for | Long projects with clear requirements | Small complex tasks | Simple high-volume labeling, pilots, multilingual data |
The key difference is who writes the rules and approves the work. An agency takes that on and includes it in the price. According to labeling provider CVAT, outsourcing to an agency typically costs more than 1.5 times as much as an in-house labeling team (Source: CVAT). With crowdsourcing, you set the rules and the review process yourself: it's cheaper and faster, but it takes a few hours of your time up front.
Pro Tip: you can combine models. Send simple high-volume labeling to crowdsourcing, disputed cases to an in-house expert, and expert work to a specialist vendor. That way expensive specialists only spend time where they're truly needed.
Annotation Guidelines: Your Main Accuracy Tool
Most labeling errors aren't carelessness — they're different readings of the rules. Does a half-hidden sauna count? What about a sauna on a billboard? If the instructions don't say, each annotator decides on their own, and each is right in their own way.
Good annotation guidelines have five parts:
- Why the labeling is needed. One sentence about what the model will do. An annotator who understands the goal resolves some disputed cases on their own.
- Precise class definitions. Not "sauna," but "a free-standing wooden sauna shaped like a barrel lying on its side."
- Examples for each class. At least two or three correct examples.
- Edge cases. A "case → decision" table: a partly hidden object, an image on a screen or poster, a blurry photo, several objects side by side.
- An "unsure" option. Let annotators flag a disputed case. It's more honest than a random guess and shows which rules need adding.
Here's an example of short instructions for a microtask:
Pro Tip: treat the guidelines as a living document. Add every new edge case from review to the decisions table and number the versions. That way you always know which rules each batch was labeled under.
How to Keep Labeling Accurate
Good instructions reduce errors but don't eliminate them. Four control mechanisms take it from there. They aren't mutually exclusive and are usually used together.
Gold-standard tasks
Mix 5–10% of items with known correct answers into the task stream. Annotators don't know which ones they are. Accuracy on gold tasks shows accuracy on everything else. It's the cheapest way to spot careless workers right away.
Overlap (consensus)
Two or three people label the same item, and the final label is decided by majority vote. For classification, three answers are usually enough. Overlap multiplies the cost by two to three, but on simple labeling it stays low while accuracy improves noticeably. If annotators disagree, the item goes to an expert for review.
Inter-annotator agreement
How often annotators agree (inter-annotator agreement) reflects the quality of the instructions more than the quality of the people. It's usually measured with Cohen's kappa, and a common benchmark for production labeling is 0.8 or higher (Source: Label Your Data). If agreement is low across all annotators at once, the problem isn't the team — it's vague rules.
Expert spot checks
During the pilot, an expert reviews everything. After that, they check 10–20% of random items from each batch plus every item marked "unsure."
Mechanism | What it catches | Cost |
|---|---|---|
Gold-standard tasks | Careless and random-answering annotators | +5–10% of volume |
Overlap | One-off errors by individual annotators | ×2–3 the cost |
Agreement | Unclear rules in the instructions | Free when you use overlap |
Expert review | Systematic errors and disputed cases | Your specialist's time |
Note: reject work only for breaking a rule that's in the instructions. If an annotator got a case wrong that the rules didn't cover, approve the work and update the instructions. Your reputation as an employer shapes who takes your future tasks.
How Much Data Labeling Costs and How Fast You Can Get It
Typical market benchmarks from labeling providers: a bounding box costs around $0.02–0.04 per object, and a polygon or segmentation label starts at $0.06. Annotator hourly rates range from $4–6 in lower-cost regions to $30 or more for domain experts in North America and Europe. A typical mid-size agency project starts at $5,000–10,000 (benchmarks from the same Label Your Data guide).
With crowdsourcing there's no minimum engagement: on RapidWorkers, the employer sets the rate per task (from $0.05), a 15% fee is added on top, and you don't pay for rejected submissions. It works best to bundle 1–5 minutes of work into one task. Sample calculations:
Task | One task | Volume and overlap | Total incl. 15% fee | Per item |
|---|---|---|---|---|
Image classification | 10 photos, ~1 minute, $0.10 | 10,000 photos, ×3 | $345 | ≈ $0.035 |
Bounding boxes | 5 photos, ~3 minutes, $0.30 | 2,000 photos, ×2 | $276 | ≈ $0.14 |
Text sentiment | 20 messages, ~2 minutes, $0.20 | 5,000 messages, ×3 | $172.50 | ≈ $0.035 |
Audio transcription | 1 minute of audio, ~5 minutes, $0.40 | 1,000 clips, no overlap | $460 | ≈ $0.46 |
The rates in the table are examples: the harder and longer the task, the higher the pay should be. The math is simple: number of tasks × overlap × rate + 15%. For classification, for example: 1,000 tasks × 3 × $0.10 = $300 + 15% = $345.
Crowdsourcing speed depends on three things: the rate, the number of available workers in the countries you select, and task difficulty. Simple labeling with no country restrictions runs in parallel across many people, so batches often finish in hours rather than weeks.
Warning: the most expensive mistake is cutting overlap and gold tasks. Labeling without them is two to three times cheaper, but nobody catches the errors, and they go straight into your model. Finding and fixing them later costs more.
How to Speed Up Labeling with Models
In 2026, labeling is rarely done from scratch. More often, a model applies preliminary labels and people check and fix them. This approach is called human-in-the-loop, and it pairs well with microtasks.
Three workflows that work:
- Pre-labeling + review. A model draws boxes or assigns classes, and a worker answers "correct / incorrect" and fixes mistakes. Reviewing is several times faster than labeling from scratch.
- Active learning. Only the items the model is unsure about go to people. According to Label Your Data, this can cut manual labeling volume by 30–70%.
- Disagreement review. If an annotator's label conflicts with a confident model prediction, the item goes back for a second check. This catches both human errors and model errors.
A pilot plan
Run a small labeling round before the big one. It takes a day and costs a few dollars to a few dozen.
- Take 100–200 items and label them yourself — these become your gold answers.
- Launch a task on the same items with ×3 overlap.
- Compare accuracy against your answers, agreement between annotators, and time per task.
- Go through every disagreement and add the edge cases to the instructions.
- If accuracy and agreement meet your bar, keep some of the pilot items as gold tasks and launch the main volume in batches.
Pro Tip: if you collect data first and label it later, build the metadata you'll need for labeling into the collection step. We covered how to do that for video and speech in our guides to video data collection and speech data collection.
Data Labeling Outsourcing FAQ
How much does data labeling cost?
From providers, a bounding box usually costs around $0.02–0.04 per object, segmentation starts at $0.06 per label, and mid-size projects start at $5,000–10,000. With crowdsourcing there's no minimum engagement: in our examples, classification with triple overlap comes to about $0.035 per image, and bounding boxes with double overlap to about $0.14.
Is outsourcing data labeling worth it?
For simple high-volume labeling, almost always: ML engineers' time is too expensive to spend on repetitive work. Expert labeling and confidential data are better kept in-house or with a specialist vendor under contract.
How do you ensure data annotation quality?
Four tools work together: detailed guidelines with edge cases, gold-standard tasks with known answers, overlap with majority voting, and expert spot checks. Start with a pilot on 100–200 items.
What is a good inter-annotator agreement?
A common benchmark for production labeling is a kappa of 0.8 or higher. If it's lower across all annotators at once, the problem is most likely the guidelines, not the people.
Crowdsourcing vs a data annotation company: which is better?
Crowdsourcing wins on simple high-volume labeling, pilots and multilingual data: you can start within hours and there's no minimum engagement. A company wins on long, complex and confidential projects where you need a contract with guaranteed accuracy.
How do you write annotation guidelines?
Explain what the labeling is for, give precise class definitions with examples, describe edge cases as "case → decision," and add an "unsure" option. For a microtask, the guidelines should fit on one screen.
Accurate Labeling Comes from a System, Not a Vendor
People often expect that outsourcing labeling is just a matter of finding "good annotators." In practice, accuracy comes from instructions, gold tasks and overlap — and those work with any outsourcing model.
The key takeaways:
- send simple high-volume labeling to crowdsourcing, and expert or confidential work to specialists;
- split complex tasks into simple steps;
- the heart of good guidelines is edge cases and an "unsure" option;
- don't skimp on gold tasks and overlap — they cost less than errors in your model;
- start with a pilot on 100–200 items and your own gold answers.
If you want to test this approach on your own data, run a small pilot in the AI Tasks section on RapidWorkers: set your instructions, rate and overlap, and get your first results the same day.
Ready to Get Started?
Join thousands of workers and employers already using RapidWorkers to get tasks done fast.
- No subscription
- Cancel anytime
- 24/7 availability
- Dispute resolution

