SWE-Bench Task Auditor
Mercor (client confidential) · Remote — United States · Remote
- Pay
- $70–90/hr
- Commitment
- hourly
- Hours / week
- ~40
- Source
- mercor
About this role
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback. Basic Qualifications • 3+ years professional software engineering • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles) • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++) Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar repository benchmarks • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.) • Prior code-review or task-grading experience
Eligible applicant countries
This role accepts applicants from:
- USA
Skills & domains
- ai-training
- rlhf
- sme
- annotation
