DataJobs.io
← Back to all jobs

Job Description

Lead Data Engineer at INSPYR Solutions in Doraville, GA (remote) responsible for architecting and optimizing distributed data pipelines, Lakehouse architecture, ML workflows, data quality, and production reliability.

Responsibilities

  • Design, build, and optimize scalable distributed data pipelines using Apache Spark within a high-volume, mission-critical environment.
  • Design and sustain an enterprise Lakehouse with Delta Lake, ensuring ACID transactions, data lineage, auditability, and governance.
  • Create automated ingestion frameworks for batch, streaming, and event-driven data across multiple cloud services and integration points.
  • Enable machine learning workflows by preparing feature-ready datasets and establishing reproducible deployment patterns for ML models.
  • Lead platform wide data quality, access control, and data cataloging initiatives.
  • Apply cost optimization, cluster tuning, and performance engineering techniques to maximize efficiency.
  • Collaborate with Finance, BI, Operations, and ML teams to translate business needs into scalable data solutions.
  • Own production reliability, troubleshoot incidents, and perform root cause analysis for data and ML pipelines.

Requirements

  • 7+ years of experience in advanced data engineering with distributed compute technologies.
  • Expert Spark engineering experience, including performance tuning, cluster configuration, partition strategies, and large dataset optimization.
  • Hands-on experience with Lakehouse architectures featuring ACID transactions, schema evolution, and governance frameworks.
  • Advanced Python and SQL proficiency for large-scale data transformations.
  • Experience supporting machine learning pipelines or model operationalization.
  • Proven track record architecting cloud-native data platforms on Azure, AWS, or GCP.
  • Strong ability to integrate diverse, complex data sources at enterprise scale.
  • Demonstrated capacity to own mission-critical production systems.
  • Experience with distributed streaming frameworks such as Kafka, Event Hubs, or equivalent.
  • Experience building or supporting ML platforms, feature stores, or experiment tracking systems.
  • Background in data security, compliance controls, or audit-ready governance.
  • Experience automating data operations with CI/CD and infrastructure as code.

Technologies

  • Databricks, Python, SQL, Apache Spark, Delta Lake, MLflow, Notebooks
  • Hugging Face Transformers, LangChain, LlamaIndex, Anthropic Claude, Meta LLaMA, Google Gemini
  • Kafka, Event Hubs, REST APIs, ADLS, S3, GCS
  • Git, GitHub, GitLab, Azure Repos, Databricks Repos, GitHub Actions, Azure DevOps
  • Unity Catalog, RBAC

Benefits

  • Comprehensive medical benefits
  • Competitive pay
  • 401(k) retirement plan

Core Tools

  • Databricks (Spark, Delta Lake, MLflow, Notebooks)
  • Python & SQL
  • Apache Spark (via Databricks)
  • Delta Lake for Lakehouse architecture

Cloud Platforms

  • Azure, AWS, or GCP
  • Cloud storage options: ADLS, S3, GCS

Data Integration

  • Kafka or Event Hubs for streaming data
  • Auto Loader for Databricks file ingestion
  • REST APIs
  • AI/ML workflows
  • MLflow for model tracking and deployment
  • Hugging Face Transformers
  • LangChain and LlamaIndex for LLM integration
  • LLMs: Anthropic Claude, Meta LLaMA, Google Gemini

DevOps

  • Git ecosystem: GitHub, GitLab, Azure Repos
  • Databricks Repos
  • CI/CD: GitHub Actions, Azure DevOps

Security & Governance

  • Unity Catalog
  • RBAC

Similar Jobs