NewLLM feedback tasks are live in 190+ countries

LLM Training Data
From Real Users

Get prompts, ideal answers and human feedback on your model's responses from everyday people in the countries and languages you choose. Set the task, set the pay and approve only what fits.

No subscription · Pay per approved answer · 12,000+ workers online

190+

Countries

13,487

Tasks completed

12,000+

Registered workers

~4 Hours

Avg time to results

How It Works

How LLM Data Collection Works

Launch a prompt writing or answer evaluation task yourself and start receiving human feedback, without long negotiations.

Person at a laptop writing LLM evaluation task instructions
01 · Create

Add your prompts and model answers

Paste the prompts or model responses you want people to work with into the task. Explain what to write, compare or rate, and add an example of a good submission.

  • Prompts or model responses
  • Clear rating rules
  • Example of a good answer
Laptop with a world map on screen for country and language targeting
02 · Target

Choose countries and languages

Select the countries your contributors should come from and state the language they should work in. See how your model performs for real users in each market.

  • Any number of countries in one task
  • Native speakers by country
  • You set the total volume
Person comparing two model answers side by side on a screen
03 · Evaluate

Real users write, compare and rate

Your task becomes available to workers in the selected countries. Each person completes it once, so you get independent opinions from many different people.

  • Independent opinions
  • Short reasons for every choice
  • New person in every submission
Person reviewing human feedback text on a computer monitor
04 · Review

Approve only the feedback that fits

Read every submission, approve the answers and ratings that follow your rules and reject the rest. Keep your dataset clean from the start.

  • Read every submission
  • Approve or reject each one
  • Clean data for training
LLM Data Types

Types of LLM Training Data You Can Collect

Set up a task for the human input your language model needs. Contributors follow your instructions and work in their own language.

Prompts and Answers

Prompts with the answers people expect

Contributors write realistic prompts and the answer they would want to get. Use the pairs for supervised fine-tuning and to set a quality bar for your model.

  • Prompt and answer in one submission
  • Topics and tone you define
  • Written without AI tools
Create an LLM task
Person typing a question and ideal answer on a laptop
Response Comparison

Which answer is better, and why

Workers see two responses to the same prompt and choose the better one with a short reason. Collect human preference data for RLHF and model selection.

  • A versus B choices
  • Short reason for every pick
  • Many opinions per pair
Create an LLM task
Two text blocks side by side on a screen for response comparison
Response Rating

Scores for helpfulness and clarity

Contributors rate model answers on a scale you set, such as helpfulness, clarity or tone. Track how your model improves from version to version.

  • Your own rating scale
  • Several criteria per answer
  • Comments on weak points
Create an LLM task
Person thinking while reading a model answer to rate it
Fact Checking

Find what the model got wrong

Workers read model answers on everyday topics and mark mistakes, made-up facts or confusing parts. Spot weak areas before your users do.

  • Everyday topics and facts
  • Mistakes marked and explained
  • Clear yes or no checks
Create an LLM task
Person checking facts with notes while reviewing a model answer
Multilingual

How your model sounds in every language

Native speakers in different countries check if answers are correct, natural and polite in their language. Find translation errors and awkward phrasing early.

  • Native speakers by country
  • Grammar, tone and wording
  • Local cultural context
Create an LLM task
People in different countries using phones for multilingual evaluation
Real User Prompts

Questions people actually ask

Contributors write the questions and requests they would type into a chatbot in daily life. Build realistic test sets that match how real users talk.

  • Everyday requests and questions
  • Natural wording and typos
  • Topics by country and language
Create an LLM task
Person typing a everyday chatbot question on a phone
Task Example

Model Answers Reviewed
in One Task

Put your prompt and two model responses into the task. Every worker picks the better one and explains why, so you get preference data with reasons.

Your setup

What you set up

Task settings

Choose Text collection, paste the prompt and two model responses into the instructions and explain how to choose the better one.

  • Text collection task type
  • Prompt and responses in instructions
  • Rules for choosing and explaining

Your view

What you receive

Submission

Each submission arrives with the worker's choice and a short reason, ready to review and approve.

  • Choice and reason in one submission
  • Worker country on every submission
  • Approve or reject each answer
Targeting

Target the Right Contributors

Choose where your contributors come from and which language they work in. Every submission comes from a different person.

Countries and languages

Accept workers from all countries or pick specific ones by continent or individually. Test your model with native speakers in each market.

Contributor requirements

Describe who should take part in your instructions, such as language level, age group or field of interest. Reject submissions that do not match.

Examples and rules

Add a sample of a good comparison or rating, so contributors follow the same logic and your data stays consistent.

Independent opinions

Each worker can complete your task only once. Many different people rate the same answers, so you see real agreement and disagreement.

Quality Control

Quality Control You Stay in Charge Of

You decide what counts as a good answer or rating and review every submission before it joins your dataset.

Person writing notes at a laptop to define rating rules

Clear rules

Define rating scales, criteria and what a good reason looks like.

Example evaluation text shown on a screen for contributors

Example submissions

Show good and bad examples so contributors know exactly what you expect.

Person reviewing LLM feedback submissions on a laptop

Approve or reject

Read each submission and accept only answers and ratings that follow your rules.

Checklist on paper used as control questions for quality checks

Control questions

Add a few answers with an obvious correct choice to spot careless submissions.

Transparency

Human Feedback You Can Trust

Every answer and rating comes from a real person who chose to take your task and knew what it was for.

Contributors know the purpose

Your task description explains that the prompts, answers and ratings will be used to train or evaluate AI.

Voluntary participation

Nobody is assigned to your task. Each worker decides to take it and submits their work on their own.

Known source of every answer

Each submission comes from a specific worker in the country you selected, written specifically for your task.

Pricing

Estimate Your LLM Data Cost

You set the pay per task. Adjust the numbers to see what your project will cost, with the 15% platform fee included.

min

How long it takes a worker to complete one task.

$

Minimum pay is $0.10 per task. Need more feedback later? Extend the same task instead of creating a new one.

FAQ

Frequently Asked Questions

LLM training data is the text and human feedback used to train and evaluate large language models: prompts, ideal answers, comparisons between responses and ratings. Data from real people helps models answer the way users expect.

It is data where people compare two or more model answers and choose the better one. These choices are used in RLHF and similar methods to teach models which answers people prefer.

Our contributors are everyday users from many countries, not certified experts. They are best for judging clarity, helpfulness, tone and language. You can describe the skills you need in your instructions and reject submissions that do not meet them.

You paste the prompts and responses into your task instructions. Workers read them, then submit their choice, rating or comment as text.

You set the pay for each task, starting from $0.10, plus a 15% platform fee. Writing tasks usually need a higher pay rate than quick comparisons. A campaign starts from $0.30.

Yes. Every worker completes your task once, so the same answers get independent opinions from as many people as you need.

Any language spoken by workers in the countries you select. Choose the countries and state the required language in your instructions.

Ask workers in your instructions to use their own words without AI tools. Review submissions carefully, reject anything that looks generated, and add control questions to spot careless work.

Yes. Extend your existing task instead of creating a new one. Workers who already took part still cannot submit again.

If you need a special format, a review flow of your own or help writing your rating rules, contact us through the form below and we will set it up with you.

Contact

Planning a Large LLM Data Project?

Need many languages, high volume or help designing your rating rules? Send a message and we will set it up with you.