DataJobs.io
← Back to all jobs

Job Description

Build and optimize a speech-to-text core that powers mobile translation performance on iPhone-class devices.

Responsibilities

  • Ingest, clean, segment, label, and version multilingual audio and transcript data, with focus on code-switching and borrowed-word behavior across the target language set.
  • Fine-tune and compress large ASR models to meet iPhone memory, latency, and battery constraints while preserving transcription quality.
  • Apply parameter-efficient and size-reduction approaches such as LoRA/QLoRA, quantization, distillation, and/or other techniques as appropriate.
  • Design packaging and on-demand language selection so language-specific weights can be downloaded based on context (for example, selecting Chinese ASR weights when an operator interviews a Chinese speaker).
  • Build and evaluate decision logic for borrowed English terms: transcribe as-is versus output as transliteration or a native equivalent in the source language.
  • Create reproducible evaluation pipelines using metrics such as word error rate and character error rate, plus latency and robustness checks (accent, noise, speaking rate, and code-switching).
  • Communicate results against defined success criteria for each language and deployment target.
  • Write clear model cards, dataset documentation, and evaluation write-ups for technical and non-technical stakeholders, including what the model does, comparisons to alternatives, and risks/limitations.

Requirements

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a closely related field.
  • Strong data-engineering background building production pipelines for large, messy, or unstructured audio/text datasets.
  • Hands-on experience fine-tuning or adapting speech/audio models using parameter-efficient methods (LoRA, QLoRA, adapters) and/or compression techniques (quantization, distillation, pruning) for constrained hardware.
  • Practical experience developing and evaluating ASR/speech-to-text systems across multiple languages, including error analysis under real-world conditions (accents, noise, code-switching).
  • Strong Python and SQL skills; experience with PyTorch, Hugging Face Transformers/PEFT, torchaudio, librosa, or comparable tooling.
  • Experience deploying and monitoring production ML systems, including secure handling of sensitive audio, transcripts, and derived data in regulated environments.
  • Ability to explain model behavior, tradeoffs, and limitations to both technical and non-technical stakeholders.

Technologies

  • Python, SQL
  • PyTorch
  • Hugging Face Transformers, PEFT
  • torchaudio, librosa
  • LoRA, QLoRA
  • Quantization, distillation, pruning
  • iPhone, iOS
  • Swift, AVFoundation

Benefits

  • Generous and flexible time-off policy
  • Flexible work schedules and telework options, including remote work availability for eligible projects
  • Career development: mentorship program, Dev University technical and management training, DevLab hands-on learning, tuition reimbursement, and paid training opportunities
  • Industry-leading benefits including choice of two health plans with dental and vision, flexible spending account, commuter benefits, life insurance, and more
  • 401K matching with a 5% matching contribution
  • Regular team and company social events (annual party, happy hours, fitness challenges, and more)
  • Community engagement: company-wide support activities, employer match for donations, and time off for volunteer efforts

Compensation: USD 80,000 - 160,000 per year.

Location: Reston, VA (onsite).

What this role is (and isn’t):

  • Owns the speech-to-text model: data, training/adaptation, model size and latency on-device, and transcription accuracy across languages.
  • Does not own iOS application development, translation (source-language to target-language), or the Swift/AVFoundation integration layer. Those are handled by a separate mobile engineering function this role will collaborate with.

Preferred (not required)

  • Prior exposure to mobile/on-device ML deployment constraints, even without owning the mobile codebase directly.
  • Experience with agentic or multi-step workflow orchestration involving model outputs, retrieval, or human review.

Similar Jobs