Text Data Collection for LLM Fine-Tuning

Text data collection for LLM fine-tuning has gotten harder than it was a couple of years ago, and the reason is ironic: language models themselves. In 2023, EPFL researchers estimated that 33–46% of workers on Amazon Mechanical Turk used LLMs on a task that asked them to write a summary in their own words. And in July 2026, Amazon announced that Mechanical Turk would stop accepting new customers from July 30.
Demand for human-written text hasn't gone anywhere, though. If you want a support chatbot to understand how your customers actually phrase their questions, or an assistant that speaks natural Polish or Tagalog instead of translated English, you need examples from real people. The good news is that fine-tuning doesn't need that many of them. Below, we cover which texts to collect, how many, how to write the task, and how to avoid ending up with ChatGPT output instead of human data.
What Kind of Text Data LLM Fine-Tuning Needs
Fine-tuning doesn't teach a model language from scratch. It shows the model, through examples, how to behave in your specific use case: what questions it will get, what tone to use, what it shouldn't say. So the data format depends on what exactly you want to change.
Format | What a person writes | What it's for |
|---|---|---|
Prompts | Questions and requests written the way real users write them | Covering every type of request, including typos and casual phrasing |
Prompt–response pairs | A prompt and a reference answer to it | Teaching the model the style, format and content of responses |
Multi-turn dialogues | A 3–10 message conversation based on a scenario | Keeping context, asking clarifying questions, declining politely |
Paraphrases | 5–10 different ways to ask the same thing | Robustness to unusual phrasing; intent classification data |
Response ratings | Picking the better of two answers, with a short explanation | Preference data for methods like RLHF and DPO that tune model behavior |
In practice, most teams start with prompt–response pairs. But one person doesn't have to write both halves. A setup that works well: workers in different countries write the prompts, and your team or vetted experts write the reference answers. Prompts need diversity; answers need product knowledge and a consistent style.
Note: this article is about written text. If your assistant works with voice, see our guide to speech data collection, which has its own requirements for recording, speakers and volume.
Why You Can't Just Generate the Data with Another Model
Synthetic data is cheaper and faster, and for some tasks it's enough. But it has three weak spots that show up precisely when you fine-tune for real users.
- Models write like models. Generated prompts are grammatical, polite and similar to each other. Real people write "didnt get the code what do i do", mix up terms, switch languages and ask two questions in one message.
- Errors compound. A paper in Nature (Shumailov et al., 2024) showed that when each new generation of models is trained on text generated by the previous one, models lose rare cases and degrade. The effect is called model collapse.
- No local context. A model doesn't know how people in Kenya ask about M-Pesa transfers or what words Filipino shoppers use in a store chat. People who live there know that best.
The approach that works for most teams: the core of the dataset is written by people, and synthetic data is added for volume and checked against real examples.
How Many Examples You Need for Fine-Tuning
Fewer than most people think. OpenAI's documentation sets the minimum at 10 examples, and noticeable improvements typically come from 50–100. In the LIMA study (Zhou et al., 2023), a 65-billion-parameter LLaMa model was fine-tuned on just 1,000 carefully curated prompt–response pairs, and in 43% of cases its answers were as good as or better than GPT-4's. The authors' conclusion: the quality and diversity of examples matter more than their number.
Volume benchmarks depend on the goal (industry practice, not hard rules):
Fine-tuning goal | Where to start | What matters most |
|---|---|---|
Tone, style, response format | 50–200 examples | A consistent style across reference answers |
Customer support for your product | A few hundred dialogues | Covering every type of request |
Responses in a language the model handles poorly | 1,000+ examples, often more | Native speakers and natural, conversational language |
Intent classification | 20–50 phrasings per intent | Variety of phrasing |
Pro tip: don't order thousands of examples up front. Collect 100–200, fine-tune the model and see where it fails. A second batch aimed at those failures usually does more than the same budget spent blind at the start.
Scenario Cards: How to Get Varied Texts Instead of a Hundred Similar Ones
Ask a hundred people to "write a question to an online store's support team," and half will write about delivery while another third will write about returns. To cover every case, give each worker their own scenario card: a short description of a situation that they turn into text in their own words.
A card is built from a few parameters. By combining them, you set the dataset's distribution in advance:
- User goal: check order status, change the address, complain, get a refund.
- State of mind: calm, in a hurry, annoyed, confused.
- Channel: website chat, messaging app, email. People write shorter messages with no punctuation in messaging apps.
- Language and register: formal, casual, mixed languages (Taglish, for example).
- Complexity: one question, two questions at once, missing information (no order number).
Five parameters with 3–4 values each produce hundreds of distinct cards. It's easiest to generate them with a script and hand out one per task. Using a model here is perfectly fine: it writes the situation, and a person writes the actual text.
What to Include in the Task for Text Writers
A writing task is different from a photo or voice recording task: the person isn't following instructions so much as making something up. That's why the situation and constraints matter more than step-by-step instructions. A good task includes:
- A scenario card: who you are, what happened, what you want to achieve.
- Format and length: "one message, up to 40 words" or "a 4–6 message dialogue, writing both sides."
- Language and style: "write the way you'd write on WhatsApp, typos and abbreviations included, if that's how you normally write."
- One example of a good response and one bad one. No more: when there are many examples, people start copying them.
- A clear ban on AI tools, plus an explanation of why you need human-written text specifically.
- A ban on personal data: no real names, phone numbers, addresses or card numbers, only made-up ones.
Example task: "You ordered sneakers from an online store a week ago, and they arrived in the wrong size. You're in a hurry and a bit annoyed. Write your first message to the support chat the way you'd actually write it, in your native language, up to 40 words. Don't use ChatGPT or other AI tools: we need to see how real people write. Don't include real names or numbers."
How to Filter Out AI-Written Text
This is the biggest risk when you crowdsource text, and there's no simple fix: you can't treat AI text detectors as the final judge. OpenAI's own classifier caught only 26% of AI-written text and wrongly flagged 9% of human-written text. It was shut down in July 2023 because of its low accuracy. Detectors struggle most with short texts and non-English languages, which is exactly what we're dealing with here.
What works is a mix of prevention and review. In a follow-up study by the same authors, about 30% of workers used LLMs when no restrictions were in place. Simply asking them not to use AI and disabling copy-paste cut that roughly in half. But the researchers also noticed a downside: LLM-assisted texts were higher quality but more uniform. For a fine-tuning dataset, uniformity is worse than typos.
Signal | What it may mean | What to do |
|---|---|---|
Perfect grammar in a "messaging app" text, bullet points, em dashes | The text may have been written by a model | Check this worker's other submissions |
Different workers opening the same way ("Hello! I would like to clarify…") | A shared source: the same model | Run near-duplicate detection across the whole set |
Text goes over the length limit or doesn't match the scenario | The worker pasted the task into a chatbot | Reject with a specific reason |
Several long submissions from one worker sent almost at once | The text was pasted, not written | Review all of their submissions manually |
You can also design the task so that AI isn't much help. Short, casual texts are easier to write yourself than to coax out of a model. Local details (service names, slang, mixed languages) come harder to models than to native speakers.
Warning: don't reject submissions based on a detector's verdict alone. A skilled writer working in a second language often looks like a model to a detector: Stanford researchers showed that detectors consistently flag text by non-native English speakers as AI-generated. Rely on a combination of signals and on the worker's track record.
How to Get Enough Variety Across Writers
Even honest texts from one person resemble each other: everyone has favorite words and turns of phrase. So variety has to be built in at the campaign level:
- a per-worker task cap (for example, no more than 5 dialogues);
- separate geo-targeted campaigns by country if the model serves several markets;
- a different scenario card for each task;
- a near-duplicate check at the end (using embeddings or MinHash) to remove examples that are too similar.
Collecting Text Through Microtasks on RapidWorkers: Process and Cost
Writing short texts splits naturally into microtasks: one card, one prompt or one dialogue. On RapidWorkers, these tasks go in the AI Tasks section (currently in beta), and geo-targeting lets you show a task only to people in the countries you choose. That matters when you need native speakers of a specific language.
The employer sets the rate per task, with a $0.05 minimum. A 15% platform fee is added to the task cost. The minimum deposit is $20 by card or crypto, and the minimum campaign is $0.30. Rejected submissions aren't paid, and if you end a campaign early, the unused budget goes back to your balance (a $0.20 cancellation fee may apply). Here's an example calculation with illustrative rates:
What you collect | Volume and rate | Tasks | 15% fee | Total |
|---|---|---|---|---|
Pilot: prompts | 50 × $0.08 | $4.00 | $0.60 | $4.60 |
Prompts up to 40 words | 1,000 × $0.08 | $80.00 | $12.00 | $92.00 |
Dialogues of 4–6 messages | 300 × $0.30 | $90.00 | $13.50 | $103.50 |
Dialogues cost more because they take longer and require keeping the logic of the conversation in your head. If the rate is too low for the amount of work, the temptation to hand the task to a chatbot grows, so saving on the rate often ends up costing you quality.
Personal Data in the Texts
People often write about themselves even when the scenario is fictional: they drop in their own name, city or phone number. If that text ends up in training, the model may memorize it. So before training, run the dataset through personal data detection (names, phone numbers, emails, addresses, ID numbers) and replace what you find with placeholders. And don't give workers your real customer conversations as samples: those contain personal data too.
How to Store the Collected Data
Most fine-tuning tools accept data in JSONL format: one example per line, each containing a list of messages with roles. That's also the format used in OpenAI's documentation. Store metadata next to the text: the scenario card, language, country and an anonymous worker ID. You'll need it to check the dataset's balance and to split it into training and test sets without leakage (all texts by one writer should end up in the same split).
The meta field is usually stripped before you upload the data to a training tool, but it should stay in your master copy of the dataset.
LLM Fine-Tuning Data Collection FAQ
What data do you need to fine-tune an LLM?
Most often, prompt–response pairs or dialogues in a role-based message format. The prompts should look like what your real users write, and the responses should look like how you want the model to answer.
How many examples do you need for fine-tuning?
OpenAI's minimum is 10 examples, and noticeable improvements usually come from 50–100. A language the model handles poorly or a complex domain takes more. The LIMA study showed that 1,000 high-quality, diverse examples can deliver strong results.
Can you fine-tune an LLM on synthetic data?
Yes, but only as a supplement. Synthetic data is uniform and doesn't reflect how real people write, and repeatedly training on generated data leads to model collapse. The core of the dataset and the test set should be written by people.
How can you tell if a text was written by a person or by ChatGPT?
There's no reliable detector, so look at a combination of signals: the same phrasing across different workers, perfect grammar where it would be unnatural, several long submissions sent almost at once. Prevention works better than detection: short, casual texts, local details and a clear ban on AI in the task.
How much does it cost to build a fine-tuning dataset?
On RapidWorkers, the employer sets the rate, starting at $0.05 per task, plus a 15% fee. In the example above, 1,000 prompts at $0.08 cost $92, and 300 dialogues at $0.30 cost $103.50.
What Really Determines the Quality of a Fine-Tuning Dataset
The size of a fine-tuning dataset matters less than it seems. Results come down to three things: the texts are written by real people, not a model; the scenarios cover what the model will actually face; and there are enough writers that no single style dominates. Start with a 50-task pilot to test the task itself, then collect 100–200 examples, fine-tune the model and add data wherever it makes mistakes.
If you need texts from native speakers in different languages and countries, launch a pilot campaign on RapidWorkers. We covered the overall process of AI data collection and its other formats in our guide to AI training data collection.
Ready to Get Started?
Join thousands of workers and employers already using RapidWorkers to get tasks done fast.
- No subscription
- Cancel anytime
- 24/7 availability
- Dispute resolution

