OCR Data for AI

OCR Data Collection
From Real Documents

Photos of receipts, forms, signs and labels from real phones — with the text typed out next to every image. Ground truth for OCR and document AI, in any language.

Photo + transcriptionAny scriptReview every file
Printed forms and documents photographed for OCR training
Forms
Handwritten text on paper for OCR training
Notes
Street signs and neon shop text in the wild
Signs
Person photographing an item with a smartphone
Phone shot
City street in mixed light for real-world capture conditions
In the wild

190+

Countries

13,487

Tasks completed

12,000+

Registered workers

~4 Hours

Avg time to results

What You Can Collect

Documents and Text You Can Collect

Pick the type of text your OCR model needs to read. Contributors photograph real items around them, or fill in templates you provide.

  • Receipts

    Store and café receipts, crumpled, faded or folded.

  • Invoices

    Printed invoices and bills in many layouts.

  • Printed forms

    Applications, checklists and tables, filled in by hand.

  • Handwritten notes

    Notes, lists and letters in many handwriting styles.

  • Menus

    Restaurant and café menus in local languages.

  • Product labels

    Ingredients, nutrition facts and packaging text.

  • Price tags

    Shelf labels and price tags in real stores.

  • Street signs

    Road signs, shop names and public notices.

  • Book and magazine pages

    Printed pages with columns, tables and captions.

  • Whiteboards

    Handwritten text and diagrams on boards.

  • Screens

    Photos of text on monitors, terminals and displays.

  • Template forms

    Your own blank forms, filled with fictional data.

Real Conditions

Text Captured the Way Users Capture It

Clean scans are easy. Your model needs to read text that is tilted, crumpled, shadowed or shot in bad light. Ask for exactly the conditions you need.

  • 01

    Tilted and skewed

    Photos taken at an angle, not straight from above.

  • 02

    Crumpled and folded

    Paper with creases, folds and torn edges.

  • 03

    Low light

    Text shot in the evening, in dim rooms or under lamps.

  • 04

    Glare and shadows

    Glossy paper, flash reflections and hand shadows.

  • 05

    Motion blur

    Quick shots taken on the go.

  • 06

    Different phones

    Old and new phones, different cameras and resolutions.

How It Works

Photos With Ground Truth Text in One Submission

Each contributor uploads the photo and types the text they see. You get image and transcription pairs ready for training and testing.

  1. 01

    Set up the task

    Choose Image collection, describe what to photograph and how to type the text: line by line, keeping numbers and symbols exactly as printed.

  2. 02

    Workers capture and type

    Contributors in the countries you choose photograph the item and type its text in the same submission.

  3. 03

    Review and approve

    Compare each photo with its text, approve accurate pairs and reject the rest. Launch a second task to double-check transcriptions if you need higher accuracy.

Printed menu or document photographed for OCR training
Typed transcription
THALI HOUSE
Veg Thali180
Masala Chai40
TOTAL220

Country: India · Language: Hindi, English · Phone camera

Languages and Scripts

Text in Any Language and Script

Choose countries where people read and write the scripts your model must support. Native writers type the text correctly, including local characters and symbols.

  • Hello

    Latin

  • Привет

    Cyrillic

  • مرحبا

    Arabic

  • नमस्ते

    Devanagari

  • 你好

    Chinese

  • こんにちは

    Japanese

  • สวัสดี

    Thai

  • হ্যালো

    Bengali

Privacy

Keep Personal Data Out of Your Dataset

Documents often contain names, addresses and card numbers. Set clear rules in every task so your dataset stays clean.

  • Ask contributors to cover or blur names, addresses, phone and card numbers before taking the photo.
  • Ask for their own receipts and documents only, never other people's.
  • Reject any submission that shows personal data you did not ask for.
  • Do not request ID cards, passports or bank documents.
Use Cases

What Teams Build With This Data

  • Receipt scanning apps

    Read totals, items and dates from real receipts.

  • Invoice and document processing

    Extract fields from printed business documents.

  • Form digitization

    Turn handwritten forms into structured data.

  • Translation apps

    Read menus, signs and labels through the phone camera.

  • Retail price recognition

    Read shelf labels and price tags in stores.

  • Handwriting recognition

    Teach models to read many styles of handwriting.

Pricing

Estimate Your OCR Data Cost

You set the pay per task. Typing out long text takes more time, so set a higher rate for long documents. The 15% platform fee is included in the total.

min

How long it takes a worker to complete one task.

$

Minimum pay is $0.10 per task. Need more documents later? Extend the same task instead of creating a new one.

FAQ

Frequently Asked Questions

OCR training data is images of text paired with the exact text they contain. Models learn to read printed and handwritten text from these pairs. Real photos taken on phones help models work in everyday conditions.

Yes. Ask workers to type the text line by line in the same submission, keeping numbers, prices and symbols exactly as printed. You get image and transcription pairs.

You review every submission and approve only accurate pairs. For higher accuracy, launch a second task where other workers compare each photo with its text and mark mistakes.

Any language and script used by workers in the countries you select, including Latin, Cyrillic, Arabic, Devanagari, Chinese, Thai and many others.

Ask contributors to cover names, addresses and card numbers, and to photograph only their own documents. Do not request ID cards or bank documents. Template forms filled with fictional data are the safest option.

Yes. Upload blank templates as reference files and ask workers to print or copy them, fill them in by hand with fictional data and photograph the result.

You decide. Set the number of files per submission, for example five different receipts. Each worker can complete your task only once.

You set the pay for each task, starting from $0.10, plus a 15% platform fee. Long documents with full transcription need a higher pay rate than single photos. A campaign starts from $0.30.

Yes. Extend your existing task instead of creating a new one. Workers who already took part still cannot submit again.

If you need special document types, a review flow of your own or help writing transcription rules, contact us through the form below and we will set it up with you.

Contact

Planning a Large OCR Data Project?

Tell us about your document types, languages and volume. We will help you write capture and transcription rules.

  • Custom document types
  • Transcription rules
  • Template forms with fictional data