DataJobs.io
← Back to all jobs

Job Description

Sciforium is seeking a Foundation Model Data Engineer to lead dataset strategy, creation, and curation for foundation model training.

Responsibilities

  • Own end-to-end creation of pre-training datasets for LLMs, including defining the mix of web data, code, books, and technical papers to improve downstream model performance.
  • Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
  • Lead development of post-training datasets, including SFT instructions, multi-turn dialogues, and preference modeling datasets for RLHF/DPO.
  • Drive acquisition and processing of vision and video data, supporting multimodal alignment and managing video compression and temporal consistency.
  • Build high-throughput Python data processing scripts using multiprocessing and multithreading to ingest and transform data at massive scale.
  • Perform deep statistical analysis of training corpora to detect biases, knowledge gaps, and quality regressions, keeping the dataset “diet” mathematically balanced.
  • (Added Value) Design pipelines to generate high-reasoning synthetic data to close gaps in natural datasets using existing models for labeling and refinement.

Requirements

  • 5+ years of industry experience in Data Science or Machine Learning, with a track record of building and managing datasets for foundation models.
  • Expert-level performance engineering skills for large-scale data tasks, including multiprocessing, multithreading, and efficient memory management.
  • Proven experience working with petabyte-scale datasets used to train production-grade LLMs or Large Vision Models.
  • Hands-on experience building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data.
  • Experience building datasets for RLHF, DPO, and multi-turn instruction following, including management of human labeling workflows and quality gold-sets.
  • Mastery of data-at-scale frameworks and formats such as Spark, Ray, WebDataset, and Parquet.

Technologies

  • Python
  • multiprocessing
  • multithreading
  • Spark
  • Ray
  • WebDataset
  • Parquet
  • Common Crawl
  • RLHF
  • DPO
  • SFT

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity

Nice-to-Haves

  • Experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines).
  • Familiarity with large-scale crawling of multimodal data and the associated video processing, codecs, and compression challenges.
  • Experience designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
  • A Master’s or PhD in a quantitative field with a focus on data-centric AI or information retrieval.

Location: San Francisco, CA (onsite)

Compensation: USD 155,000 - 210,000 per yearly

Similar Jobs