Image Data Collection for Computer Vision

Image Data Collection for Computer Vision: How to Build a Custom Dataset
A computer vision model only knows the world through the images it was trained on. If the training set was full of bright offices, new smartphones and products from American store shelves, accuracy will drop in a dim warehouse, on a cheap camera or in a shop in Lagos, and switching architectures won't fix that. That's why teams shipping CV models to production sooner or later reach the same conclusion: they need a custom dataset built for their own use case.
The market backs this up. According to Grand View Research, the computer vision market was worth $23.6 billion in 2025 and is projected to reach $101.5 billion by 2033 (a 20.1% CAGR). The more models move into real-world conditions, the higher the demand for images that reflect those conditions.
This guide covers image data collection for computer vision from scratch: when open datasets aren't enough, what images different tasks need, how to write a dataset spec, how many images you need, how to run collection through microtasks, and how to check quality before labeling. We covered the overall process in our guide to AI training data collection; here the focus is on images only.
Open Datasets vs a Custom Dataset: When Off-the-Shelf Isn't Enough
Open datasets are a good starting point. COCO has about 330,000 images across 80 object categories, and ImageNet has more than 14 million images. They're great for pre-training and for benchmarking architectures. But for production they have three limitations.
Domain shift. Large open datasets were scraped from the internet, and they contain noticeably more photos from Western, high-income countries. In the paper "Does Object Recognition Work for Everyone?", Facebook AI researchers tested five commercial recognition services (Azure, Clarifai, Google Cloud Vision, Amazon Rekognition and IBM Watson) on photos taken in homes across 54 countries. Accuracy for households earning under $50 a month was roughly 10% lower than for households earning over $3,500. The gap between the United States and Somalia or Burkina Faso was 15–20%. The reason is simple: soap, toothbrushes and kitchens look different from country to country and appear in different surroundings.
Your classes aren't there. If your model needs to spot defects on your products, recognize your packaging on a shelf or read receipts in a specific format, open datasets simply don't contain those images.
Licensing. "Publicly available" doesn't mean "cleared for commercial use." The authors of "Can I use this publicly available dataset to build commercial AI software?" reviewed six popular datasets, including ImageNet, MS COCO, Cityscapes and FFHQ, and found that none of them explicitly allows commercializing models trained on it. For every one, there was a usage scenario with a risk of license violation.
Open dataset | Custom dataset | |
|---|---|---|
Cost | Free | You pay for collection and labeling |
Fit for your task | General; your classes may be missing | Exact: your objects, conditions, devices |
Geography and diversity | Skewed toward Western countries | You set it: countries, lighting, cameras |
Usage rights | Often restricted or unclear | Rights and consent are handled during collection |
Time to start | Immediately | From a few days (pilot) |
The setup most teams use: pre-train on open data, then fine-tune and evaluate on a custom set shot in real usage conditions.
What Images Different Computer Vision Tasks Need
Before you write a task for workers, decide what the model will do. That determines what goes in the frame, how many objects each photo contains and how the data will be labeled later.
Task | What the photo should show | Example collection format |
|---|---|---|
Classification | One main object or scene, close up | 5 photos of the inside of your fridge in normal lighting |
Object detection | Several objects in a natural setting, at different scales | A full store shelf shot from 1–2 m away |
Segmentation | Clear object boundaries, varied backgrounds | Houseplants against different backgrounds |
OCR and documents | Text at different angles and in different light, with glare and creases | Receipts, labels, storefront signs |
Gestures, poses, hands | People in frame, different poses, skin tones, backgrounds | A hand showing a specific gesture |
Retail and products | Packaging from different sides, on the shelf and in hand | A product from the front, side and back |
Scenes and environments | Interiors, streets, weather, time of day | A building entrance by day and by night |
If the model needs to understand motion or a sequence of actions, still photos won't be enough. In that case it's better to collect video and extract frames. We explained how to set that up in our article on video data collection for AI training.
The Dataset Spec: What to Lock Down Before Collection Starts
A spec is a one- or two-page document that becomes the basis for worker instructions and acceptance criteria. The more precise it is, the fewer photos you'll have to reject. It should cover the following.
- Subject and framing. What must be in the frame, how much of the frame it fills, and whether the edges of the object can be cropped.
- Angles and distance. For example: front, 45° angle, top-down; from 30 cm and from 1 m.
- Lighting. Daylight, artificial light, dusk. If the model will run in poor lighting, the dataset needs poor lighting too.
- Background and setting. A plain background makes labeling easier, but a model trained only on plain backgrounds performs poorly in real settings. You usually want natural backgrounds.
- Resolution and format. Minimum size on the short side (for example, 1080 px), JPEG or PNG, no filters, frames or captions. Screenshots and photos of screens are usually not allowed.
- Devices. If the model will run in a mobile app, photos should be taken on a range of smartphones, including budget ones.
- Metadata. What workers submit along with the photo: country, type of location, phone model, object class. Decide up front whether you need EXIF data and geolocation, and if you don't, strip them on intake.
- Restrictions. Other people's faces, documents, license plates, screens showing personal data, images from the internet.
Example wording: "Photograph any carton of milk in your home. We need 4 photos: front, side, top and in your hand. The carton should fill at least half the frame. Use your phone camera with no filters. No faces or documents in the frame."
How Many Images You Need and How to Get Real Diversity
There's no universal number, but there are benchmarks. Ultralytics' training tips for YOLO detection models recommend at least 1,500 images and at least 10,000 labeled instances per class, plus 0–10% "background" images with no target objects to reduce false positives. That's a benchmark for strong production results. For a pilot or for fine-tuning a pre-trained model, teams usually start with far less and then collect more data wherever the model makes mistakes.
Diversity matters more than the raw count. The same guidance lists what should vary: time of day, season, weather, lighting, angle and source (different cameras). A thousand near-identical shots from one person usually give a model less than three hundred photos from a hundred different people in different conditions.
To make diversity a requirement rather than a wish, build a coverage matrix and set quotas. For example, for 2,000 photos of products on shelves:
Parameter | Distribution |
|---|---|
Countries | 4 countries, 500 photos each |
Lighting | 60% store lighting, 25% daylight, 15% dim |
Distance | 50% full shelf, 50% close-up |
Workers | No more than 3 tasks (12 photos) per person |
Background images | 5% shelves without the target products |
A per-worker cap is a simple, reliable way to get different cameras, locations and shooting styles. It also prevents data leakage between splits: put all of one person's photos in either the training set or the test set, otherwise your test metrics will be inflated.
How to Collect Images Through Microtasks on RapidWorkers
Crowdsourcing fits image collection well: workers already have smartphones, they live in different countries, and they photograph real homes, stores and streets rather than studio sets. RapidWorkers has an AI Tasks section for this kind of work (currently in beta). Here's how the process works.
- Turn the spec into a task. Short steps, examples of good and bad photos, a list of restrictions and one line on what the data is for. A worker should understand the task in a minute.
- Decide what counts as one task. A set works best: for example, 4 photos of one object from different angles. That makes review easier and keeps the per-person cap under control.
- Choose countries. Geo-targeting shows the task only to workers in the countries you select. For the coverage matrix above, you can run a separate campaign for each country with its own limit.
- Set the rate. The employer sets the rate per task, with a $0.05 minimum. Harder shoots (going to a store, finding a specific product, shooting in specific light) should pay more than a photo of an object on a desk, or the task will be picked up slowly.
- Top up and launch a pilot. The minimum deposit is $20 by card or crypto, and the minimum campaign is $0.30. A 15% platform fee is added to the task cost.
- Review and approve submissions. Rejected submissions aren't paid. If you end a campaign early, the unused budget goes back to your balance (a $0.20 cancellation fee may apply).
Example calculation. You need 2,000 photos of packaging. One task is 4 photos of one package, so you need 500 tasks. At $0.20 per task:
Item | Amount |
|---|---|
Pilot: 50 tasks × $0.20 | $10.00 |
15% fee | $1.50 |
Pilot total | $11.50 |
Main collection: 500 tasks × $0.20 | $100.00 |
15% fee | $15.00 |
Total for 2,000 photos | $115.00 |
This is an example rate for illustration. The final price depends on how hard the shoot is and how fast you need the data. If you're on a tight deadline, you can pin the task to the top of the feed for $10.
Checking Images on Intake
Check photos before they go to labeling. Otherwise you pay twice for every extra or faulty image: once for collection and once for labeling. The most practical approach combines automated filters with manual review.
Automated checks (a few lines of Python with OpenCV or Pillow):
- resolution below the minimum in your spec;
- blur (for example, using variance of the Laplacian) and frames that are too dark or overexposed;
- duplicates and near-duplicates by perceptual hash, including across different workers;
- screenshots and images from the internet (no camera EXIF, unusual aspect ratios, found by reverse image search);
- faces and text in the frame where they shouldn't be.
Manual review covers what automation can't catch: whether it's the right object, whether the angles and lighting match the spec, and whether the metadata matches what's in the photo. Review every submission at the start, then switch to a sample from each worker. When you reject a submission, give a specific reason from a ready-made list ("photo is blurry", "side view missing") so workers pick up the requirements faster.
Faces, Personal Data and Consent
The simplest option is to collect photos with no people in them at all. If your task needs faces, hands or bodies, the requirements get stricter. According to guidance from the UK ICO, a photo of a person isn't biometric data in itself, but it becomes biometric data when it undergoes specific technical processing that allows a person to be uniquely identified. Illinois' BIPA follows similar logic: photographs themselves are excluded, but face geometry derived from them is protected. Statutory damages are $1,000 for negligent and $5,000 for intentional or reckless violations per person.
In practice, that comes down to three rules. Workers photograph only themselves or objects, not bystanders. The task states plainly what the photos will be used for, and explicit consent is collected. Incidental faces, license plates and documents in the background are blurred, or the photo is rejected.
After Collection: Labeling and Targeted Top-Ups
Collected images only become a dataset once they're labeled with class tags, bounding boxes or masks. If the class is already defined at collection time ("photograph a carton of milk"), part of the work for classification is already done. Detection and segmentation need a separate step. We explained how to run it without losing accuracy in our article on outsourcing data labeling.
A dataset isn't built in one pass. After the first training run, look at where the model fails: in the dark, on a specific class, in one of the countries. Then run a targeted top-up for exactly those images. A few hundred deliberately collected hard examples often do more than another thousand random ones. Microtasks are handy here: a new campaign with a refined task takes a few minutes to launch.
Checklist: An Image Dataset in 7 Steps
- Define the model's task and the list of classes.
- Write the spec: framing, angles, lighting, background, resolution, restrictions.
- Build a coverage matrix and set a per-worker cap.
- Run a pilot of 30–50 tasks and review every submission.
- Fix the task based on the pilot's common mistakes.
- Launch the main collection with automated filters and sample-based review.
- Label, train, find the model's weak spots and collect more data for them.
A custom image dataset is something competitors can't download: it reflects your objects, your users and the real conditions your model works in. To get started, create a campaign on RapidWorkers and run a small pilot.
Image Data Collection FAQ
What is image data collection for computer vision?
It's the process of gathering photos that reflect the real conditions a model will work in: the right objects, angles, lighting, devices and countries. After collection, the images are checked and labeled, and only then used for training.
How many images do you need to train a model?
For object detection, Ultralytics recommends at least 1,500 images and at least 10,000 labeled instances per class. For a pilot or fine-tuning a pre-trained model, teams usually start with less and add data where the model makes mistakes.
Can you train a commercial model on open datasets?
Not always. Many popular datasets don't explicitly allow commercial use of models trained on them. Check the license with a lawyer before using one in a product.
How much does it cost to collect an image dataset with microtasks?
On RapidWorkers, the employer sets the rate, starting at $0.05 per task, plus a 15% fee. In the example above, 2,000 photos (500 tasks at $0.20) cost $115.
What's the difference between image collection and image labeling?
Collection means getting new photos. Labeling means adding the tags, boxes or masks the model learns from. You can run both through microtasks as separate campaigns.
How do you collect photos from specific countries?
Use geo-targeting: the task is shown only to workers in the countries you select. For even coverage, run a separate campaign for each country.
Ready to Get Started?
Join thousands of workers and employers already using RapidWorkers to get tasks done fast.
- No subscription
- Cancel anytime
- 24/7 availability
- Dispute resolution

