Staff Data Engineer
Job Description
LiveView Technologies is building physical AI systems, and this role sits at the center of the company’s data engine. As a Staff Data Engineer in Seattle, you will own the end-to-end data flywheel that turns raw edge telemetry and video into labeled training data, frozen evaluation and benchmark sets, and governed dataset outputs that support model training and evaluation.
This is a senior individual-contributor position with technical leadership responsibilities, partnering closely with AI/ML research and MLOps to ensure datasets stay consistent, versioned, and ready for measurable comparisons over time.
What you’ll do
- Own the full loop that converts raw edge telemetry and video into labeled training data, frozen evaluation sets, and the next iteration of model outputs.
- Design and operate pipelines that register raw source data, standardize it into a single well-defined schema, and join and aggregate curated datasets so teams use one consistent store and reader instead of copying and reformatting per use case.
- Append labels and semantic annotations without rewriting source data, then ensure datasets are versioned, quality-checked, and served for downstream training and evaluation, in partnership with annotation and data-operations teams.
- Own frozen, versioned validation and benchmark datasets so model comparisons remain valid over time, with the scrubbing and review discipline needed before sets are shared externally.
- Set up schema and content versioning that allows producers to evolve datasets without breaking consumers by supporting opt-in versions, append-without-rewrite for new fields, and controlled reader/writer indirection during rollout.
- Build the read/write libraries and integrations researchers rely on, including support for PyTorch/Lightning dataloaders, a record-level CRUDL API, and Spark/analytics access and self-service.
- Implement governance directly in the pipeline so classification of clips, frames, labels, and embeddings is machine-enforced, including scrubbing and anonymization during load jobs, plus lineage and provenance for dataset versions and annotation campaigns.
- Define data-engineering standards for the flywheel schema conventions, dataset contracts, and quality gates, mentoring other engineers as the function grows.
Required experience
- 8+ years building and operating large-scale production data pipelines and data-lake or lakehouse systems, including ingestion, ETL/ELT, partitioning and storage-format decisions, and reader/writer library design.
- Experience building pipelines for model training and evaluation, including labeled data and evaluation/benchmark sets, with an understanding of how data quality and versioning affect model results.
- Strong background with medallion-style layered architectures and modern table/lake formats such as Iceberg, Delta, Parquet (or comparable), including schema evolution and dataset versioning.
- Hands-on experience working with large multimodal data including video, image, and sensor/telemetry, with storage and access patterns that make data queryable at scale.
- Practical knowledge of the data side of ML frameworks, including PyTorch/Lightning dataloaders, plus strong Python and Spark skills.
- Practical experience enforcing data governance in pipelines covering classification, access control, lineage and provenance, and retention for privacy-sensitive data.
- Demonstrated track record setting data-engineering direction and leveling up engineers through technical leadership (formal management not required).
- Bachelor’s or Master’s in Computer Science, Engineering, or a related field (or equivalent practical experience).
Tools and technologies
- PyTorch, Lightning, Spark, Python
- Iceberg, Delta, Parquet
- Kafka, EMR, Lambda, MCP
- PyTorch/Lightning dataloaders, Encord, Labelbox, CRUDL API
Compensation
- Salary range: USD 171,900 - 221,000 per year, determined by location, job-related experience, and education/training.
- Additional total earning potential via a bonus structure tied to meeting goals.
- Employee equity program with ownership from day one.
Benefits
- Comprehensive health, dental, and vision coverage
- Retirement benefits (401k match up to 4%)
- Flexible PTO
Preferred qualifications
- Streaming or near-real-time ingestion from edge/IoT sources into a data lake (for example, Kafka, Lambda, EMR, or similar).
- Append-without-rewrite and hash-indexed dataset approaches on open table formats, plus dataset and feature versioning systems.
- Generative-AI data work, including fine-tuning and evaluation dataset curation for LLMs and VLMs.
- Exposing datasets to AI agents through MCP-style query interfaces, including semantic schema and plain-language documentation for retrieval.
- Computer-vision and video annotation tooling and workflows (for example, Encord, Labelbox, or similar).