DataJobs.io
← Back to all jobs

Job Description

In this role, you will design and build scalable data pipelines and data solutions with an emphasis on data governance, validation, and quality assurance. You will collaborate in an agile environment to develop, maintain, and troubleshoot data solutions focused on performance, security, and reliability.

Key Responsibilities

  • Design and build data pipelines required for optimal data processing across a variety of data sources.
  • Collaborate with cross-functional teams to identify data requirements and align with business objectives for projects or initiatives.
  • Independently analyze, design, and troubleshoot data flows based on business needs.
  • Translate business requirements into technical specifications.
  • Adjust data collection processes, including indexing and query optimizations, to support optimal performance.
  • Build Extract, Transform, and Load (ETL) pipelines to enable efficient data collection and extraction.
  • Profile data sources to support successful pipeline builds.
  • Define success and failure thresholds for data collection pipelines.
  • Implement data governance policies and procedures for data handling, including data retention, to help maintain consistency, integrity, accuracy, and reliability across the data lifecycle.
  • Redact Personally Identifiable Information (PII) and Protected Health Information (PHI) to support compliance with data privacy and security standards.
  • Follow data security measures to protect data from unauthorized access, use, disclosure, alteration, or destruction.
  • Ensure data compliance with applicable laws, regulations, and industry standards.
  • Implement data validation and integrity checks, identify data quality issues, and address items that could impact data pipeline or model performance.
  • Define data annotation and labeling processes to support data quality, working independently where needed.
  • Design and implement automation of data validation and governance.
  • Independently design, develop, and optimize automated and scalable data pipeline architectures using ETL, creating reusable data products.
  • Implement appropriate data storage solutions for processed data to enable scalable, optimized access and analysis.
  • Manage day-to-day operation of data pipeline flow and storage.
  • Write runnable code and perform testing and debugging of data solutions.
  • Independently manage work by monitoring timelines and deliverables to keep initiatives on track and aligned with requirements.
  • Prioritize work proactively and adapt to changes in resources or timelines, suggesting adjustments to maintain efficiency.
  • Collaborate across teams to align expectations and achieve shared objectives, supporting effective partnerships with stakeholders.
  • Identify and address standard and non-standard issues according to standard practices, escalating more complex items appropriately.
  • Troubleshoot errors by analyzing data and information from multiple sources, including standard and non-standard issues.
  • Contribute to knowledge sharing and best practices and support a culture of continuous learning.
  • Develop ideas and recommend updates to improve the efficiency and effectiveness of team processes, protocols, and workflows.
  • Seek input on alternative approaches and methods to improve how work is executed.

Technologies

  • Extract, Transform, and Load (ETL)
  • Indexing
  • Query optimizations

Work Setup

Location: Cleveland, OH (onsite)

Collaboration: Agile development environment with other engineers.

Similar Jobs