- Research + data-production ML
- RL environments / post-training data for frontier labs
- ML-research spike + strong engineering (both)
- Strong data taste
Inbound demand for post-training data quality exceeds what Emulated can currently fulfill, so this hire is about scaling the talent to meet demand. The shape they've found works best — and the one to source for — is a domain/research expert with a strong ML-or-engineering spike: someone who both does research and produces data, because in Emulated's thesis the two are inseparable (their research lead builds post-training data extensively). For this early hire they're explicitly merging the research and data-production archetypes into one person.
The technical scope is broad: post-training models as effectively as possible (from provisioning infra through curating high-quality post-training data from their generation processes), and — critically — scaling data generation superlinearly with human intervention. A core competency is data taste: good judgment about what problems pervade a domain and what makes a realistic, economically valuable task/environment a lab will actually buy. They index on research/ML horsepower first. Above all they screen for high slope.
What You'll Own
- Post-training models end to end — infra provisioning through curating high-quality post-training data
- Producing RL environments / long-horizon tasks across domains (incl. research-loop, science, chip, and shock-physics-type environments)
- Scaling data generation superlinearly via self-improving developer processes, not more bodies
- Exercising and developing data taste — realistic, economically valuable tasks/environments labs will buy
- Blending research and data production as one function (the company's core thesis)
Requirements
- A genuine spike in ML/AI research (post-training, RL, evals, interpretability, or similar) plus strong software-engineering ability — the dual competency is the whole shape
- Can post-train models and build/curate high-quality post-training data (the two are inseparable here)
- Data taste, or the research horsepower to develop it fast
- Fast problem-solving / code comprehension (their technical round tests how quickly you understand code and problem-solve, not rote LeetCode; ML variant uses PyTorch / debugging ML pipelines)
Background
- Domain/research expertise in a specific field + ML/eng core (neuroscience, physics, chemistry, systems, etc. all fit — breadth is the culture)
- Pedigree is a soft signal, not an index: Palantir / Anthropic / OpenAI welcome; competitor data cos (Surge, Mercor, Fleet) interesting but less defined
Who Will Thrive Here
- A researcher-engineer who treats data as a research problem and wants to build the model that makes data
- High-slope polymath with a real spike who figures anything out
- Someone with (or fast to develop) taste for lab-valuable tasks/environments
- Deeply committed, results-driven, obsessive builder with public artifacts
- Excited by the bitter-lesson, superlinear-scaling thesis and frontier-lab customers
No benefits listed yet.