DataJobs.io
← Back to all jobs

Job Description

Vulcan Elements is building the data foundation that enables operational analytics and AI/ML at scale. In this onsite role in Durham, NC, you will design and deliver the data architecture and pipelines that move manufacturing and lab outputs into a secure analytics and AI-ready Lakehouse environment, with attention to CUI and ITAR data handling requirements. If you value engineering clarity and reliability, you will help shape standards the team can grow on.

What you’ll do

  • Design and own Vulcan’s data architecture from operational data stores through ETL pipelines into the analytics and AI layer
  • Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, balancing scalability, compliance requirements, operational burden, and cost
  • Review, refine, and implement data architecture design documents, ensuring designs account for CUI and ITAR data handling
  • Make and document key platform and design decisions so future team members can follow the reasoning
  • Ensure the architecture scales from pilot plant to full-scale facility without fundamental redesign
  • Apply engineering best practices across the data stack, including version control, testing, observability, and documentation, as the team grows
  • Design and build ETL/ELT pipelines that move operational data into the data Lakehouse with contextual enrichment for analytics and AI/ML workloads
  • Build dependable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems
  • Collaborate with engineering, operations, and IT to understand data flows and dependencies, then translate requirements into pipeline and architecture decisions
  • Identify and eliminate manual workflows by replacing them with monitored, reliable pipelines
  • Diagnose and resolve data quality issues across the stack, and add monitoring so issues surface early
  • Define data models that support operational queries, analytical workloads, and future AI/ML applications
  • Own data contextualization standards so every data point includes the metadata needed to make it meaningful
  • Contribute to schema design and payload definitions for operational data stores to improve consistency and legibility
  • Support reporting and visibility tooling that gives operations and leadership insight into process and quality data
  • Write clear technical documentation for architecture decisions, data models, pipeline designs, and operational runbooks

What you bring

  • 8+ years of experience in data engineering, data infrastructure, or a closely related technical role delivering production systems
  • Experience designing and building data lakes, Lakehouses, or analytical data stores, including making and defending platform selection tradeoffs
  • Strong ETL/ELT skills, especially enriching and contextualizing data
  • Deep fluency in data modeling for operational and analytical workloads
  • Experience with relational databases such as PostgreSQL or SQL Server, including writing and debugging SQL
  • Comfort working in a fast-moving environment with a small team, making decisions with incomplete information and documenting them clearly
  • Strong communication skills across technical and non-technical stakeholders
  • Must be a U.S. Person due to required access to U.S. export-controlled information or facilities

Tools you may use

PostgreSQL, SQL Server, InfluxDB, TimescaleDB, Delta Lake, Apache Iceberg, Airflow, Prefect, dbt, Python, SQL, AWS, Azure, GCP, MQTT, CUI, ITAR, EAR

Desired background

  • Time-series database experience (InfluxDB, TimescaleDB, or similar) in industrial/IoT settings
  • Industrial data familiarity such as historian data, process tags, and OT/IT integration
  • Experience with a Unified Namespace or MQTT-based data architecture
  • Familiarity with Lakehouse platforms and open table formats like Delta Lake and Apache Iceberg
  • ETL orchestration experience with Airflow, Prefect, dbt, or similar tools
  • Scripting comfort for pipeline development and data quality tooling (Python, SQL, or similar)
  • Cloud platform familiarity (AWS, Azure, or GCP) and evaluating on-prem vs. cloud tradeoffs
  • Experience working in controlled information environments, including CUI and export-controlled data handling under ITAR or EAR
  • Experience in manufacturing, industrial, or operations-heavy environments

Location

This role begins in Durham, NC onsite and is expected to move to Benson, NC upon completion of a new facility.

Similar Jobs