O
AI Agent Task Evaluator — Freelance (Remote — Worldwide)
AI AgentsEvaluationQAAutomation
About the role
Frontier labs are training agents that browse, click and execute multi-step tasks. Outlier needs freelance evaluators to run these agents through realistic workflows and grade every step.
Responsibilities
Execute agent trajectories in sandboxed environments; grade tool calls and reasoning steps; document failures with reproducible notes; suggest harder test scenarios.
Requirements
Comfortable with browsers, spreadsheets and basic developer tools; systematic thinker; strong written English. Coding experience is a plus, not required.
Benefits
Fully remote, work whenever you want. Weekly payouts. No fixed hours or minimum commitment.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.