Machine Learning Scientist
Rime
Official active full-time posting for Machine Learning Scientist at Rime, located in United States.
Role facts
| Employer | Rime |
|---|---|
| Location | United States |
| Work mode | Remote |
| Remote scope | unknown |
| Employment | full-time |
| Function | Research & data |
| Track | industry |
| Seniority | unspecified |
| Compensation | unknown |
| Posted | 2026-05-15 |
| Work authorization | unknown |
| Official posting | Open official posting ↗ |
Responsibilities
- Design, train, and evaluate speech synthesis models, autoregressive and non-autoregressive.
- Drive research on full-duplex and half-duplex multi-modal architectures, including unified S2S systems.
- Choose and iterate on speech representations: neural codecs, semantic tokens, mel features, continuous latents.
- Build rigorous evaluation, objective and perceptual. Hold the bar on quality and prosodic control.
- Collaborate with our linguists on TTS frontend behavior so modeling and frontend choices reinforce each other.
Required skills
- Deep familiarity with the speech synthesis literature, contemporary and historical — Tacotron, FastSpeech, VITS, VALL-E, the codec-LM lineage. Opinions on what worked and why.
- Hands-on training with neural codecs (EnCodec, DAC, Mimi, etc.) and multiple representation choices.
- Experience with full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S).
- Strong attention to detail on data quality. You notice when an annotation pipeline is silently degrading or when an eval set has leakage.
- Willing to roll up your sleeves on unglamorous data and training work — paired with the agency to build pipelines so the team isn't stuck doing it by hand.
- Working knowledge of TTS frontend (G2P, normalization, prosody) and experience working with linguists.
- Strong PyTorch fundamentals. Comfortable with training loops, distributed training, model internals.
- PhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant.
Preferred skills
- Multilingual TTS experience.
- Background in prosody or paralinguistics.
- Published work in speech, audio, or core ML venues.
- Experience taking research models to production: quantization, distillation, streaming inference.
- WHY JOIN RIME
- Category-defining voice AI infrastructure, not incremental research deltas.
- Direct collaboration with founders, including a CEO with a Stanford computational linguistics PhD.
- Real impact on company trajectory.
- Meaningful equity upside.
- High ownership, high standards, low bureaucracy.
Tools, models and methods
- PyTorch
- TTS
Related
Sources & changes
View sources and updates
Source details
Statuscited source ↗
canonical namecited source ↗
subtypecited source ↗
short descriptioncited source ↗
Official URLcited source ↗
primary geographycited source ↗
datescited source ↗
external idscited source ↗
exact titlecited source ↗
employer namecited source ↗
posting datecited source ↗
posting statuscited source ↗
locationcited source ↗
workplace typecited source ↗
work authorizationcited source ↗
employment typecited source ↗
senioritycited source ↗
compensationcited source ↗
responsibilitiescited source ↗
required skillscited source ↗
preferred skillscited source ↗
named tools models methods domainscited source ↗
education requirementscited source ↗
experience requirementscited source ↗
official application sourcecited source ↗
departmentcited source ↗
teamcited source ↗
role familycited source ↗
employer subtypecited source ↗
selection basiscited source ↗
Recordcited source ↗
Update history
- Status unknown → active
- Official URL unknown → jobs.ashbyhq.com/rime/b76243da-3ce0-463e-9f29-2b6d1b78666e
- Record maintenance · 12 fields updated
Is this your role? Claim this record →·See something wrong? Report a correction →