This position is no longer accepting applications
Closed on September 9, 2026.
This role is filled — get an email when new Data Processing roles open on DataJobs.io:
Audio AI Engineer
Ai Engineer
Artificial Intelligence
Asr
Audio
Audio Ai
Data Processing
Engineer
Machine Learning Engineer
Model Compression
On Device Ai
On Device Ml
Speech Recognition
Voice Recognition
View similar jobs
Get alerted when similar jobs are posted — set up a New Data Processing jobs on DataJobs.io alert.
See other roles at Dev Technology.
Job Description
Build and optimize a speech-to-text core that powers mobile translation performance on iPhone-class devices.
Responsibilities
- Ingest, clean, segment, label, and version multilingual audio and transcript data, with focus on code-switching and borrowed-word behavior across the target language set.
- Fine-tune and compress large ASR models to meet iPhone memory, latency, and battery constraints while preserving transcription quality.
- Apply parameter-efficient and size-reduction approaches such as LoRA/QLoRA, quantization, distillation, and/or other techniques as appropriate.
- Design packaging and on-demand language selection so language-specific weights can be downloaded based on context (for example, selecting Chinese ASR weights when an operator interviews a Chinese speaker).
- Build and evaluate decision logic for borrowed English terms: transcribe as-is versus output as transliteration or a native equivalent in the source language.
- Create reproducible evaluation pipelines using metrics such as word error rate and character error rate, plus latency and robustness checks (accent, noise, speaking rate, and code-switching).
- Communicate results against defined success criteria for each language and deployment target.
- Write clear model cards, dataset documentation, and evaluation write-ups for technical and non-technical stakeholders, including what the model does, comparisons to alternatives, and risks/limitations.
Requirements
- Bachelor's degree in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a closely related field.
- Strong data-engineering background building production pipelines for large, messy, or unstructured audio/text datasets.
- Hands-on experience fine-tuning or adapting speech/audio models using parameter-efficient methods (LoRA, QLoRA, adapters) and/or compression techniques (quantization, distillation, pruning) for constrained hardware.
- Practical experience developing and evaluating ASR/speech-to-text systems across multiple languages, including error analysis under real-world conditions (accents, noise, code-switching).
- Strong Python and SQL skills; experience with PyTorch, Hugging Face Transformers/PEFT, torchaudio, librosa, or comparable tooling.
- Experience deploying and monitoring production ML systems, including secure handling of sensitive audio, transcripts, and derived data in regulated environments.
- Ability to explain model behavior, tradeoffs, and limitations to both technical and non-technical stakeholders.
Technologies
- Python, SQL
- PyTorch
- Hugging Face Transformers, PEFT
- torchaudio, librosa
- LoRA, QLoRA
- Quantization, distillation, pruning
- iPhone, iOS
- Swift, AVFoundation
Benefits
- Generous and flexible time-off policy
- Flexible work schedules and telework options, including remote work availability for eligible projects
- Career development: mentorship program, Dev University technical and management training, DevLab hands-on learning, tuition reimbursement, and paid training opportunities
- Industry-leading benefits including choice of two health plans with dental and vision, flexible spending account, commuter benefits, life insurance, and more
- 401K matching with a 5% matching contribution
- Regular team and company social events (annual party, happy hours, fitness challenges, and more)
- Community engagement: company-wide support activities, employer match for donations, and time off for volunteer efforts
Compensation: USD 80,000 - 160,000 per year.
Location: Reston, VA (onsite).
What this role is (and isn’t):
- Owns the speech-to-text model: data, training/adaptation, model size and latency on-device, and transcription accuracy across languages.
- Does not own iOS application development, translation (source-language to target-language), or the Swift/AVFoundation integration layer. Those are handled by a separate mobile engineering function this role will collaborate with.
Preferred (not required)
- Prior exposure to mobile/on-device ML deployment constraints, even without owning the mobile codebase directly.
- Experience with agentic or multi-step workflow orchestration involving model outputs, retrieval, or human review.