Mercor Logo

Mercor

Mercor Research Fellowship — APEX

Posted 3 Days Ago
In-Office or Remote
Hiring Remotely in New York City, NY, USA
40K-80K Annually
Entry level
In-Office or Remote
Hiring Remotely in New York City, NY, USA
40K-80K Annually
Entry level
Design, build, and validate new AI benchmarks and evaluation methodologies. Fellows create task specifications and grading rubrics, collaborate with domain experts, run frontier models, analyze failures, test for contamination and gaming, and publish findings through papers, datasets, leaderboards, or internal methodologies. The fellowship provides mentorship, compute, expert labor, and access to enterprise evaluation problems.
The summary above was generated by AI
About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Fellowship

Mercor’s APEX benchmark family measures whether frontier AI models can actually do economically valuable work: multi-hour agentic tasks in investment banking and corporate law, real professional accounting workflows, real-world software engineering, and graduate-level science. Every APEX benchmark is built and validated with Mercor’s network of domain experts — not written from a textbook.

The Mercor Research Fellowship funds people to build the next generation of benchmarks and evaluation techniques. You pitch a benchmark or eval methodology you want to build — a new domain, a harder task format, a better way to measure agentic reliability — and if selected, you get the time, compute, expert labor, and mentorship to design, implement, and release it end to end.

You’ll work directly with the APEX research team, get access to real enterprise evaluation problems from Mercor’s Fortune 500 and frontier-lab partners, and see your benchmark shape how the industry measures AI capability.

Program Details

  • Duration: 3–6 months, rolling admission

  • Commitment: minimum 30 hours/week; full-time preferred

  • Location: remote, or in-person at Mercor’s San Francisco office

  • Admission: apply with a specific benchmark or eval technique you want to build — the fellowship is funded around your pitch, not a generic research rotation

What You’ll Do

  • Propose and scope a new benchmark or evaluation technique in a domain APEX doesn’t yet cover well, or a meaningfully harder version of one it does.

  • Design task specifications and grading rubrics in partnership with Mercor’s network of vetted domain experts — lawyers, accountants, engineers, scientists, and consultants.

  • Build and validate the benchmark: pilot tasks, calibrate scoring, and stress-test for contamination and gameable shortcuts.

  • Run frontier models against your benchmark and analyze where and why they fail.

  • Publish your results — as a paper, an open dataset, a new leaderboard on APEX, or a methodology the APEX team adopts internally.

  • Partner with Mercor’s research and engineering teams to fold what you learn back into APEX’s public benchmark family.

Focus Areas

  • Long-horizon, multi-app agentic tasks in professional services (law, finance, consulting) — extending APEX-Agents

  • Real-world software engineering evaluation beyond issue resolution — extending APEX-SWE

  • Professional accounting and finance workflows — extending APEX-Accounting

  • AI-for-Science evals: research-level mathematics, biology, materials science, and theoretical physics

  • Novel evaluation methodology: contamination resistance, rubric design, human-vs-model grading agreement, cost-adjusted scoring

  • Strong pitches outside this list are welcome — we fund the best ideas, not the closest fit to a template.

What We’re Looking For

  • Genuine interest in evaluation as a research discipline — not just a stepping stone to a model-building role.

  • Background in CS, ML, statistics, or an adjacent field (measurement, psychometrics, HCI, social science); no requirement to have published in ML venues.

  • A specific, well-scoped idea for a benchmark or eval technique you want to build — the fellowship is built around your pitch.

  • Comfortable in a startup environment: fast iteration, direct access to real customer problems, less hand-holding than an academic lab.

  • Able to commit at least 20 hours/week for the duration of the fellowship — during a leave, over a summer, or a flexible stretch of a PhD.

  • Bonus: experience with agentic evaluation, RL environments, or domain expertise in law, finance, medicine, or a scientific field.

Compensation & Benefits

  • 3 month stipend of $40,000 or 6 month stipend of $80,000

  • Unlimited API credits, plus a dedicated budget for GPU compute and paid expert/human-data time

  • Weekly 1:1 mentorship with a member of the APEX research team, plus regular access to the broader research org

  • Access to frontier model APIs, Mercor’s internal evaluation infrastructure, and — where appropriate — real enterprise evaluation problems from Mercor’s customers

  • Optional desk in Mercor’s San Francisco office for fellows who want to be in person

  • Introductions to Mercor’s network of researchers across frontier labs and academia

  • Standout fellows are considered for a full-time offer on the APEX research team at the end of the fellowship

Similar Jobs

2 Hours Ago
Remote or Hybrid
United States
60K-101K Annually
Mid level
60K-101K Annually
Mid level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Manage assigned client accounts to ensure satisfaction and renewals. Coach clients on SailPoint/IdentityIQ identity and access solutions, monitor usage and risks, provide strategic updates, identify expansion opportunities, and drive resolutions to customer issues.
Top Skills: IdentityiqSailpoint
2 Hours Ago
Remote
United States
166K-193K Annually
Senior level
166K-193K Annually
Senior level
Healthtech • Social Impact • Software • Telehealth
Guide growth strategy through financial analysis of complex deals, pricing, contract economics, and strategic forecasting. Partner with Sales, Marketing, and Revenue Operations to evaluate trade-offs, translate financial data into actionable insights, and support executive decision-making. This individual contributor role requires navigating ambiguity and balancing short-term efficiency with long-term business impact in a fast-growing healthcare company.
3 Hours Ago
Easy Apply
Remote or Hybrid
New York, NY, USA
Easy Apply
250K-330K Annually
Senior level
250K-330K Annually
Senior level
Artificial Intelligence • Cloud • Software
Lead Vercel’s consumption forecasting strategy and end-to-end machine learning systems across infrastructure products. Develop advanced time-series, probabilistic, hierarchical, causal, and predictive models for operational and financial planning. Build scalable tooling for backtesting, monitoring, drift detection, retraining, and explainability. Partner with Finance, Infrastructure, Product, and GTM leadership on revenue planning, capacity optimization, pricing, and adoption scenarios. Set technical standards, influence organizational ML practices, and mentor senior data scientists and ML engineers.
Top Skills: AirflowBayesian ModelingCausal InferenceDbtDeep LearningDelta LakeFeature StoresMlops ToolingPythonSnowflakeSQL

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account