← 返回岗位列表美国IT/互联网fulltime

AI Red Team Specialist - LLM

雇主

mercor

地点

Remote · 美国

待遇

$USD 60 - USD 90 / hourly

工作模式

远程

截止日期

11月26日

🤖 AI 简历匹配评估

检测你的简历与该岗位的匹配度,免费

免费评估

岗位摘要

About the jobMercor connects elite creative and technical talent with leading AI research labs.

岗位职责

About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Position: LLM Red Team Specialist — Failure Modes & Edge Cases
Type:Contract
Compensation:$60–$90/hour
Location:Remote
Commitment:35 hours/week
Role Responsibilities
Probe models to explore how frontier AI models handle coding, ML, and analysis tasks. Identify spots where models quietly fail.
Design challenges by turning identified weaknesses into well-crafted tasks that are challenging for models but fair to grade.
Document findings clearly with reproducible evidence and steps.
Strengthen tasks by collaborating with task authors to close loopholes, shortcuts, and grading gaps.
Work as a team by sharing insights with researchers and fellow experts to continually improve the benchmark.
Qualifications
Must-Have
MSc or PhD in a STEM field or equivalent practical experience.
1+ years of experience in research, research-engineering, security, or AI-evaluation roles.
Ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems.
Proficiency in Python and Git for scripting probes and analyses.
Familiarity with LLM capabilities, limitations, and evaluation techniques.
Ability to engage reliably for approximately 35 hours/week.
Preferred
Experience in AI training, model evaluation, or benchmark/task authoring.
High attention to detail, creativity, and strong written communication skills.
Application Process(Takes 20–30 mins to complete)
Upload resume
AI interview based on your resume
Submit form
Resources & Support
For details about the interview process and platform information, please check:
For any help or support, reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas

申请条件

- MSc or PhD in a STEM field or equivalent practical experience
- 1+ years of experience in research, research-engineering, security, or AI-evaluation roles
- Ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems
- Proficiency in Python and Git for scripting probes and analyses
- Familiarity with LLM capabilities, limitations, and evaluation techniques
- Ability to engage reliably for approximately 35 hours/week
- Experience in AI training, model evaluation, or benchmark/task authoring (preferred)
- High attention to detail, creativity, and strong written communication skills (preferred)

工作地要求

United States

雇主简介

Mercor connects elite creative and technical talent with leading AI research labs, providing AI safety evaluation services.

对这个岗位感兴趣?

该岗位暂未开放在线申请,顾问可为您推荐同类岗位或申请指导

咨询不收取任何费用,顾问将为您推荐合适的岗位与申请方式

申请海外岗位,英文简历符合当地格式规范吗?

AI 自动评估你与该岗位的匹配度,3 分钟出结果

免费评估简历匹配度

数据来源:Himalayas

岗位信息来源于公开渠道,版权归原作者所有