Loading your workspace…
Rounds/Retell/Research Scientist - Audio·#28934
Live · accepting submissions
Log in to see accurate information
You're viewing this role as a guest. Sign in to see your application status, referral credits, and personalized match details.
Retell logo
AI·Series A·Redwood City, California

Research Scientist - Audio

at RetellAI
Location
Redwood City, California
Salary
$225,000 - $400,000
Type
Full-Time
About Retell
Retell AI is building the modern CX platform for enterprise contact centers, using first-principles voice AI to automate sales, support, and logistics calls at scale. Thousands of companies—including CVS/Aetna, American Airlines, Lenovo, and Grab—use Retell's AI voice agents to replace large teams of human agents, with the platform powering everything from frontline agents to QA analysts and managers. Backed by Y Combinator (W24), Alt Capital, and Carya Venture Partners, Retell claims $80M ARR as of August 2026 (public data: $60M+ as of April 2026), up from $5M at the start of 2025, and a valuation north of $1.5B. The team is 50 strong, with alumni from Stripe (payments lead), Google AI research, and Facebook product. The company has been recognized as a top AI app by a16z, #3 Fastest-Growing Software Company (G2 2026), and a member of the Nasdaq & Wing VC Enterprise Tech 30 (2026). The vision: a fully AI-powered contact center where intelligent agents continuously execute, monitor, and improve every customer interaction—no more basic automation that needs constant human tuning. The market believes: Retell is the fastest-growing voice AI company in the category, with real revenue, enterprise traction, and a reputation for technical depth.
Series A • 51 – 200
Stage & size
AI
Industry
2023
Founded
About This Role

The mandate: Advance the model capabilities behind human-like voice agents operating in real-world conditions. Explore new approaches across LLMs and audio models, design novel evaluation frameworks, and prototype systems that improve reasoning, latency, and conversational quality — with your work landing in production. This is a research role at a company with no research org yet. The body of the JD calls it "Founding," which is the honest framing: this person defines what ML research means at Retell. What You'll Own

  • Research and experimentation — new techniques across LLMs and audio models targeting reasoning, latency, and conversational quality in real-time systems

  • Model training — build and iterate on models and pipelines fast; innovate on training paradigms and inference

  • Evaluation and benchmarking — design novel eval frameworks, datasets, and metrics for complex real-world voice tasks. This is arguably the highest-leverage piece; conversational quality is subjective and unsolved.

  • Research-to-production — work directly with engineering to get findings deployed

  • Human feedback loops — methods for incorporating human evaluation into model improvement, especially on subjective conversational quality

  • Frontier tracking — bring new ideas into Retell's product and infrastructure Requirements Hard gates:

  • Master's in CS, ML, AI, or related field required. PhD preferred. Equivalent research-level engineering experience considered — but treat the degree as the default filter and only bypass it for candidates with a genuinely research-grade portfolio.

  • Advanced ML research experience: LLM pre-training or post-training, transcription/ASR model training, TTS, or multimodal systems. Industry or academia both count.

  • PyTorch fluency, model architecture depth, and comfort with the underlying math. The loop tests all three directly.

  • On-site in the Bay Area — relocation fully covered Profile:

  • Can go from open-ended problem to working prototype without a spec

  • Translates research into systems that survive production

  • Communicates complex ideas cross-functionally Strong bonuses:

  • First-author or co-author publications at NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP, or equivalent. Interspeech and ICASSP are the highest-signal venues for this specific req.

  • Competition awards

  • Speech-specific depth: ASR, TTS, speaker diarization, voice conversion, streaming/low-latency inference

  • Real-time or streaming model experience — latency is a first-class constraint here, not an afterthought

  • RLHF, preference modeling, or eval design for subjective quality Location and visa:

  • On-site, Redwood City. 100% relocation provided — this is on this req and not the others. Lead with it for out-of-market candidates.

  • Sponsorship: Yes — H-1B, TN, L-1, E-3, F-1 (OPT/CPT). Note: no O-1 listed here, unlike the FDE req. Worth confirming with Retell, since O-1 is the most relevant visa for publishing researchers. Anti-patterns

  • ML engineers who fine-tune off-the-shelf models and call it research. The PyTorch and ML-theory rounds will expose this immediately.

  • Pure academics with no interest in production. Half this job is getting research deployed.

  • Researchers who need a large team, established infrastructure, and a defined agenda

  • Applied scientists whose work is entirely feature engineering and classical ML

  • Anyone requiring remote — no exception surfaced

  • Candidates who only want to publish. Retell isn't a lab; publication isn't the deliverable. Who Will Thrive Here

Someone with real research credentials who is tired of the queue at a big lab. They want their model in front of 50M calls next month, not in a paper next year. They're equally happy deriving the loss function and debugging the inference server at 2am. Ex-big-lab researchers who want founding-level scope, and strong PhDs from speech/audio groups who want production stakes, are the two clearest profiles.

Job Details
Experience
3+ Years
Salary
$225,000 - $400,000
Visa Sponsorship
Yes
Employment Type
Full-Time
Work Arrangement
In office
Work Intensity
9-9-5
Benefits & Perks
$70/day Doordash Credit For Unlimited Meals And Snacks
$200/month Wellness Reimbursement
$75/month Phone Bill Reimbursement
$50/month Internet Reimbursement
$300/month Commuter Reimbursement
100% Coverage For Medical Dental And Vision Insurance
Green Flags
First/co-author at top-tier ML venues (NeurIPS, ICML, ICLR, ACL, Interspeech)
Built and shipped ML models for real-time audio or LLMs in production
Experience with human-in-the-loop evaluation or novel benchmarking
Red Flags
No advanced ML research experience (LLMs, audio, TTS, or multimodal)
Lacks hands-on PyTorch/model implementation depth
Pure academic with no evidence of shipping research to production
Ideal Companies
G
Google DeepMind
O
OpenAI
A
Anthropic
M
Meta AI
M
Microsoft Research
A
Amazon Alexa AI
A
Apple ML Research

Required Candidate Q&A

Question 1
Are you able to work in-person in Redwood City, CA (100% in-office, full relocation provided)?
Question 2
What is your current visa status, and do you require sponsorship to work in the US?
Question 3
Walk through your strongest evidence for advanced ML research experience (LLMs, audio, TTS, or multimodal).
Question 4
Describe a research idea you took from concept to production—what was the impact?
Question 5
Have you published at top-tier ML venues or won notable ML/AI competitions? If so, which ones?
Question 6
How do you approach designing evaluation frameworks for complex, real-world ML tasks?
Question 7
Why Retell—what about real-time voice AI or this team pulls you?

Candidate scorecard

· 9 criteria
PhD or Master's in CS/ML/AI or equivalent research-level engineering experience
Strong ML research background (LLM pre-training/post-training, transcription, TTS, or multimodal)
Deep technical foundation in PyTorch, model architectures, and ML math
Rounds · Confidential to recruiting partners · Last refreshed Aug 21, 2026