mercor
Software Engineer - Benchmark Auditor
About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: SWE-Bench Task Auditor Type: Contract Compensation: $70–$90/hour Location: Remote Role Responsibilities • Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks . • Assess repository-level tasks, reference patches, test harnesses, and grading integrity. • Provide clear, rubric-based written feedback to improve AI model training . • Audit reference patches, test runners, and Docker isolation to detect answer leakage and reward hacking. • Work independently and asynchronously to meet deadlines and enhance AI model performance . Qualifications Must-Have • 3+ years professional software engineering experience. • Real open-source contribution or maintainer experience (merged PRs, committer/maintainer roles). • Strong ability to audit reference patches, test runners, and Docker isolation. • Fluency across common ecosystems ( Python and at least one of Java / Go / TypeScript / C++ ). Preferred • Familiarity with SWE-Bench (Verified) or similar repository benchmarks. • Maintainer history on major Python OSS ( Django , Flask , scikit-learn , sympy , pytest , etc.). • Prior code-review or task-grading experience. Application Process (Takes 20–30 mins to complete) • Upload resume • AI interview based on your resume • Submit form Resources & Support • For details about the interview process and platform information, please check: • For any help or support, reach out to: PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity. Originally posted on Himalayas
