Job Description
ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT THE ROLE
You’ll work with large enterprises to capture their data and transform it into high-fidelity RL environments for capability evaluations and training datasets for frontier labs. We focus on pushing the frontier of world-building, verifier engineering, and more alongside our partners.
Your goal will be to automate the process of building evals for real work in the economy.
WHAT YOU'LL DO
- Ship models for workflow extraction, classification, and grading.
- Engineer autonomous task refinement processes which distill data taste into pipelines.
- Deliver data to customers and deploy into real engagements.
- Help define the future of agentic transformation for enterprises around the world.
- Deeply learn about the intricacies of enterprises through building evaluations for all aspects of work.
- Build end-to-end environments for labs & enterprises by platformizing sandbox app clones, load real data into the sandboxes, build prompts from real workflows, and write verifiers leveraging enterprise expertise & golden outputs.
- Systematize the production of environments to scale throughput while maintaining high-quality worlds and verifiers.
WHAT WE'RE LOOKING FOR
- Prior experience shipping environments – you’ve contributed to an OSS framework, built environments at previous companies, or worked on agentic evaluations.
- Strong full-stack engineering skills – you’ll be responsible for everything from infrastructure to app code to analytics
- Bias to action – this team is focused on shipping evals, not just philosophizing about them.
- Curiosity – being biased towards understanding and digging deep into model behavior and actually looking at the data.
- Sweat the details that make a simulation indistinguishable from the real thing and have systems-level thinking skills that allow you to scale up quality.
NICE TO HAVE
- Experience with Temporal, Modal, or similar orchestration/compute services
- Experience with synthetic data generation for frontier models.Past work auditing and scrutinizing industry-standard evaluations
BENEFITS
- Semi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance