The work
Design evaluation tasks that distinguish genuine capability from pattern matching; write adversarial and edge-case prompts in your specialism; review model-generated research code, experiments and claimed results; grade reasoning on systems, theory and algorithms problems.
Who it suits
- PhD in computer science, statistics or a related field with publications in machine learning, NLP, systems or theory.
- Strong engineering skills; experience training or evaluating large models a plus.
How it works
- Apply with your CV and the hours you can give. 5 to 20 hours a week suits most people; the work is remote and on your own schedule.
- We verify your credentials, licence or degree and your identity before any paid work. Every piece of work traces to a named, verified expert.
- A paid calibration task on the project's own guidelines comes first.
- You are paid at the rate on this page for every hour of work that passes review. Rates rise with the difficulty of the project and your track record.

