DataJobs.io
← Back to all jobs

Job Description

Infinitive Inc is hiring a Data Engineer to help design, build, and scale next-generation, event-driven data platforms in McLean, VA (onsite). The role blends high-volume streaming with durable workflow orchestration, with a strong focus on schema discipline and reliable data pipelines spanning batch and near-real-time processing.

In this position, you will work on distributed systems using Apache Kafka for streaming data movement and Temporal for long-running pipeline coordination. You will also design robust data contracts through schema versioning and automated validation, then connect that rigor to analytical models in cloud data warehouses and lakehouses.

What you’ll do

  • Architect, deploy, and maintain high-volume distributed data streams using Apache Kafka, including producers, consumers, Kafka Connect, and Schema Registry.
  • Define and enforce schema design standards, including versioning strategies and automated schema validation (Avro, Protobuf, JSON Schema) to keep strict data contracts across microservices, streaming consumers, and lakehouse storage.
  • Build durable execution workflows with Temporal to coordinate long-running pipelines, handle cross-system ETL tasks, and apply compensation patterns such as the Saga pattern.
  • Create end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark / PySpark.
  • Design and optimize analytical data models (dimensional/star schema) in modern cloud data platforms such as Snowflake, BigQuery, Databricks, and Redshift.
  • Implement automated testing, continuous schema validation, data drift detection, and observability for both streaming and batch workflows.
  • Collaborate with software engineers, machine learning engineers, and analysts to define standard schema definitions, data contracts, and production-grade CI/CD release patterns.

What you bring

  • 4+ years of professional experience in data engineering, backend distributed systems, or software engineering.
  • Hands-on experience with Temporal (or Cadence), including durable workflows concepts such as activities, retries, signals, queries, and long-running orchestration.
  • Deep expertise in Apache Kafka, including message partitioning, consumer groups, offset management, and topic design.
  • Practical experience with schema definition frameworks: Apache Avro, Protocol Buffers/gRPC, or JSON Schema.
  • Experience managing schema evolution and compatibility modes (backward, forward, full), plus schema registries such as Confluent Schema Registry or AWS Glue Schema Registry.
  • Ability to enforce data validation rules, contract testing, and data quality checks using tools such as Great Expectations, Pandera, Pydantic, or dbt tests.
  • Strong programming proficiency in Python (Go or Java is a plus), with clean code, design patterns, and unit/integration testing standards.
  • Hands-on development with Apache Spark (PySpark/Spark SQL) for large-scale data processing.
  • Strong relational database experience, dimensional data modeling, and query performance tuning.

Tools you’ll work with

Apache Kafka, Kafka Connect, Schema Registry, Temporal, Cadence, Apache Avro, Protocol Buffers, gRPC, JSON Schema, Confluent Schema Registry, AWS Glue Schema Registry, Great Expectations, Pandera, Pydantic, dbt, Python, SQL, Apache Spark, PySpark, Spark SQL, Snowflake, BigQuery, Databricks, Redshift.

Compensation

Infinitive Inc is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00 per year.

Similar Jobs