
- Must already know RL/evals/post-training
- 4% commission (~$500k TC)
- Data-company background ideal (e.g. Handshake)
- Build coding RL evals/tasks + RealSuite benchmark
Specific is hiring its first SF engineer — a founding research engineer to build out the pipeline of coding RL evals and tasks. They’re mid-flight on pilots/contracts with top frontier labs and want to scale the engineering muscle in SF proactively. Day one, this person is working with customers and on RealSuite — running evals of frontier-lab models on the benchmark and on representative tasks. By ~3 months they’re building RL tasks directly with frontier labs; by ~6 months, ideally leading a team of engineers.
The profile is a research-engineering blend — with data, the key is genuinely understanding what an eval is, what a harness is, what an RL task is: real research intuition plus the engineering to build it. The hard non-negotiable: no zero-baseline candidates — they must already understand RL / coding RL / evals / benchmarks / post-training (they won’t teach that from scratch). The second non-negotiable: young, high-energy, hungry — the two traits that worked for their India team were (a) prior data-company experience and (b) being extremely hungry, wanting to do and learn as much as possible. The ideal hire comes from another data company; absent that, clear self-driven evidence of understanding post-training counts (a personal benchmark, a side project, a post-training/quant experiment). Culture is the tiebreaker: very high agency, works really hard, earnest, high integrity, wants to learn — they’d be the 5th person in the office.
What You’ll Work On
- Build coding RL evals/tasks and the pipeline of task creation
- Work on RealSuite — evaluate frontier-lab models on the benchmark and representative tasks
- Build RL tasks directly with frontier labs (by ~month 3)
- Work with customers (labs) from day one; quality-control and deliver to spec
- Grow into leading an engineering team (~month 6)
Requirements
- Already understands RL / coding RL / evals / harnesses / benchmarks / post-training — no zero-baseline candidates
- Research-engineering blend: real research intuition plus the engineering to build tasks/evals
- Ideal: prior experience at a data company; absent that, concrete self-driven evidence (a built benchmark, a post-training/quant side project)
Background
- Target data/RL companies: Handshake (Janak hears many want to leave), Prime Intellect, Mercor, Trajectory, Bespoke, “AptQuery,” and similar
- Bonus: research at a strong lab/school (Scale-adjacent, Berkeley, etc.); wants to start a company someday
- Strong signal: talent from outside the SF/CA ecosystem who’s deeply curious about AI/RL/post-training and wants to come to SF (their best hire came from New York)
Location and Visa
- A hungry, high-agency research engineer who already lives in RL/evals/post-training and wants to build the task pipeline for frontier labs
- Someone from a data/RL company (or with self-built benchmark/post-training evidence) ready to own it as the 5th person in the room
- Earnest, high-integrity, here-to-learn — not chasing quick cash
- Energized by huge upside (4% project commission → ~$500k) and a fast path to leading a team