Software Engineer, ML Serving
Rime
Official active full-time posting for Software Engineer, ML Serving at Rime, located in United States.
Role facts
| Employer | Rime |
|---|---|
| Location | United States |
| Work mode | Hybrid / flexible |
| Remote scope | unknown |
| Employment | full-time |
| Function | Research & data |
| Track | industry |
| Seniority | unspecified |
| Compensation | unknown |
| Posted | 2026-06-18 |
| Work authorization | unknown |
| Official posting | Open official posting ↗ |
Responsibilities
- Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.
- Model optimization from a single-node to disaggregated fleet serving.
- Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.
- Continuous integration and deployment workflows for the model serving pipeline.
- Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.
- Resource provision, cost management across our GPU fleet.
Required skills
- Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.
- Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).
- Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).
- IaC experience — Terraform, Packer, or comparable. You should have opinions about how to do this right.
- On-call is part of the job. You treat production reliability as a shared responsibility.
Preferred skills
- Experience with multinode training (DDP, FSDP, etc.).
- Experience with gRPC or other bidirectional binary streaming protocols.
- Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).
- Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.
- Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).
- Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).
- SRE, DevOps, or platform engineering background at a startup.
- Experience at an early-stage company.
- WHY JOIN RIME
- Build the serving infrastructure behind a category-defining voice AI company from the ground up.
- You will bring in experience that no one else currently has at the company: you can help us set the vision.
- Direct collaboration with the inference, platform, and ML teams — no handoff culture.
- The systems you build determine what experiences our customers can deploy at scale.
- Meaningful equity upside at an early stage.
- High ownership, high standards, low bureaucracy.
- SF / Bay Area.
Tools, models and methods
- cloud infrastructure
- Kubernetes
- Docker
- AWS
- GCP
- TTS
Related
Sources & changes
View sources and updates
Source details
exact titlecited source ↗
canonical namecited source ↗
subtypecited source ↗
Statuscited source ↗
short descriptioncited source ↗
Official URLcited source ↗
primary geographycited source ↗
datescited source ↗
external idscited source ↗
employer namecited source ↗
posting datecited source ↗
posting statuscited source ↗
locationcited source ↗
workplace typecited source ↗
work authorizationcited source ↗
employment typecited source ↗
senioritycited source ↗
compensationcited source ↗
responsibilitiescited source ↗
required skillscited source ↗
preferred skillscited source ↗
named tools models methods domainscited source ↗
experience requirementscited source ↗
official application sourcecited source ↗
departmentcited source ↗
teamcited source ↗
role familycited source ↗
employer subtypecited source ↗
selection basiscited source ↗
Recordcited source ↗
Update history
- Status unknown → active
- Official URL unknown → jobs.ashbyhq.com/rime/3ce06fc8-1896-4fa3-99d4-54b615480f59
- Record maintenance · 12 fields updated
Is this your role? Claim this record →·See something wrong? Report a correction →