Foundation Model Data Engineer
Job Description
Sciforium is seeking a Foundation Model Data Engineer to lead dataset strategy, creation, and curation for foundation model training.
Responsibilities
- Own end-to-end creation of pre-training datasets for LLMs, including defining the mix of web data, code, books, and technical papers to improve downstream model performance.
- Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
- Lead development of post-training datasets, including SFT instructions, multi-turn dialogues, and preference modeling datasets for RLHF/DPO.
- Drive acquisition and processing of vision and video data, supporting multimodal alignment and managing video compression and temporal consistency.
- Build high-throughput Python data processing scripts using multiprocessing and multithreading to ingest and transform data at massive scale.
- Perform deep statistical analysis of training corpora to detect biases, knowledge gaps, and quality regressions, keeping the dataset “diet” mathematically balanced.
- (Added Value) Design pipelines to generate high-reasoning synthetic data to close gaps in natural datasets using existing models for labeling and refinement.
Requirements
- 5+ years of industry experience in Data Science or Machine Learning, with a track record of building and managing datasets for foundation models.
- Expert-level performance engineering skills for large-scale data tasks, including multiprocessing, multithreading, and efficient memory management.
- Proven experience working with petabyte-scale datasets used to train production-grade LLMs or Large Vision Models.
- Hands-on experience building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data.
- Experience building datasets for RLHF, DPO, and multi-turn instruction following, including management of human labeling workflows and quality gold-sets.
- Mastery of data-at-scale frameworks and formats such as Spark, Ray, WebDataset, and Parquet.
Technologies
- Python
- multiprocessing
- multithreading
- Spark
- Ray
- WebDataset
- Parquet
- Common Crawl
- RLHF
- DPO
- SFT
Benefits
- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity
Nice-to-Haves
- Experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines).
- Familiarity with large-scale crawling of multimodal data and the associated video processing, codecs, and compression challenges.
- Experience designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
- A Master’s or PhD in a quantitative field with a focus on data-centric AI or information retrieval.
Location: San Francisco, CA (onsite)
Compensation: USD 155,000 - 210,000 per yearly