DataJobs.io
← Back to all jobs

This position is no longer accepting applications

Closed on August 12, 2026.

This role is filled — get an email when new SQL roles open on DataJobs.io:

Job Description

Lead Data Engineer at INSPYR Solutions in Doraville, GA (remote) responsible for architecting and optimizing distributed data pipelines, Lakehouse architecture, ML workflows, data quality, and production reliability.

Responsibilities

  • Design, build, and optimize scalable distributed data pipelines using Apache Spark within a high-volume, mission-critical environment.
  • Design and sustain an enterprise Lakehouse with Delta Lake, ensuring ACID transactions, data lineage, auditability, and governance.
  • Create automated ingestion frameworks for batch, streaming, and event-driven data across multiple cloud services and integration points.
  • Enable machine learning workflows by preparing feature-ready datasets and establishing reproducible deployment patterns for ML models.
  • Lead platform wide data quality, access control, and data cataloging initiatives.
  • Apply cost optimization, cluster tuning, and performance engineering techniques to maximize efficiency.
  • Collaborate with Finance, BI, Operations, and ML teams to translate business needs into scalable data solutions.
  • Own production reliability, troubleshoot incidents, and perform root cause analysis for data and ML pipelines.

Requirements

  • 7+ years of experience in advanced data engineering with distributed compute technologies.
  • Expert Spark engineering experience, including performance tuning, cluster configuration, partition strategies, and large dataset optimization.
  • Hands-on experience with Lakehouse architectures featuring ACID transactions, schema evolution, and governance frameworks.
  • Advanced Python and SQL proficiency for large-scale data transformations.
  • Experience supporting machine learning pipelines or model operationalization.
  • Proven track record architecting cloud-native data platforms on Azure, AWS, or GCP.
  • Strong ability to integrate diverse, complex data sources at enterprise scale.
  • Demonstrated capacity to own mission-critical production systems.
  • Experience with distributed streaming frameworks such as Kafka, Event Hubs, or equivalent.
  • Experience building or supporting ML platforms, feature stores, or experiment tracking systems.
  • Background in data security, compliance controls, or audit-ready governance.
  • Experience automating data operations with CI/CD and infrastructure as code.

Technologies

  • Databricks, Python, SQL, Apache Spark, Delta Lake, MLflow, Notebooks
  • Hugging Face Transformers, LangChain, LlamaIndex, Anthropic Claude, Meta LLaMA, Google Gemini
  • Kafka, Event Hubs, REST APIs, ADLS, S3, GCS
  • Git, GitHub, GitLab, Azure Repos, Databricks Repos, GitHub Actions, Azure DevOps
  • Unity Catalog, RBAC

Benefits

  • Comprehensive medical benefits
  • Competitive pay
  • 401(k) retirement plan

Core Tools

  • Databricks (Spark, Delta Lake, MLflow, Notebooks)
  • Python & SQL
  • Apache Spark (via Databricks)
  • Delta Lake for Lakehouse architecture

Cloud Platforms

  • Azure, AWS, or GCP
  • Cloud storage options: ADLS, S3, GCS

Data Integration

  • Kafka or Event Hubs for streaming data
  • Auto Loader for Databricks file ingestion
  • REST APIs
  • AI/ML workflows
  • MLflow for model tracking and deployment
  • Hugging Face Transformers
  • LangChain and LlamaIndex for LLM integration
  • LLMs: Anthropic Claude, Meta LLaMA, Google Gemini

DevOps

  • Git ecosystem: GitHub, GitLab, Azure Repos
  • Databricks Repos
  • CI/CD: GitHub Actions, Azure DevOps

Security & Governance

  • Unity Catalog
  • RBAC

Similar Jobs