GroundWork Renewables, Inc. seeks a Data Engineer to design, build, and maintain internal data infrastructure, pipelines, and tools to ensure laboratory measurement data is reliably accessible for the lab team. This onsite role in Albuquerque, NM offers a salary range of USD 90,000 to 100,000 per year and requires collaboration across engineering, laboratory, and operations to uphold data reliability, accessibility, and compliance.
Responsibilities
- Technical subject matter expertise: design, implement, and maintain relational and time-series databases for lab instrument data, environmental measurements, and operational records; develop and manage ETL/ELT pipelines to ingest, transform, and store data from IoT sensors, measurement hardware, and remote sensing platforms; build and deploy internal data access tools and applications using modern frameworks such as Streamlit, FastAPI, or React to enable lab staff to query and visualize data.
- Data quality assurance and control: develop and enforce QA/QC protocols to validate incoming data from lab instruments and field sensors in line with regulatory and accreditation standards (for example ISO 17025); implement automated checks, flagging routines, statistical validation, and audit trails to detect anomalies, missing data, and calibration drift; maintain defensible data records that satisfy chain-of-custody and traceability requirements.
- Database architecture and optimization: architect and optimize database schemas for performance, scalability, and ease of access; evaluate and recommend appropriate database technologies (SQL, NoSQL, time-series) based on data volume, query patterns, and reporting needs.
- Stakeholder collaboration: partner with lab engineers, metrology staff, and operations to translate data access requirements into technical solutions; serve as the primary point of contact for internal data availability and reporting needs.
- Internal tool and application development: design and build internal data access tools, dashboards, and reporting interfaces using modern full-stack frameworks (React, FastAPI, Streamlit, Plotly Dash); leverage AI-assisted development environments to accelerate development cycles while ensuring maintainability, security, and governance for lab data; enable non-technical lab staff to explore, filter, and export data.
- Data governance and documentation: maintain data dictionaries, schema documentation, and data lineage records consistent with laboratory quality management systems; contribute to SOPs and data management plans; stay current with emerging data engineering technologies, AI tooling, and laboratory informatics practices to continuously improve the data infrastructure.
Requirements
- Minimum of 3 years of experience in data engineering, database engineering, ETL/ELT pipeline development, or a related technical discipline.
- Experience designing and operating production data pipelines and infrastructure.
- Bachelor's degree in computer science, software engineering, information systems, data science, or a related field.
- Proficiency in SQL and experience with relational databases (PostgreSQL, MySQL, or similar); familiarity with time-series or NoSQL databases is a plus.
- Proficiency in Python (pandas, SQLAlchemy, FastAPI, or similar) for data engineering, scripting, and backend service development.
- Hands-on experience designing and operating ETL/ELT data pipelines and workflow orchestration tools (e.g., Apache Airflow, Dagster, Prefect, or similar), including scheduling, dependency management, and pipeline monitoring.
- Experience building web applications or data dashboards using tools such as Streamlit, Dash, FastAPI, React, or modern AI-assisted development environments; ability to deliver functional, user-facing tools rapidly using AI pair-programming workflows.
- Experience implementing QA/QC workflows for instrument or sensor data, including anomaly detection, validation rules, statistical flagging, and audit logging; familiarity with laboratory quality management standards (ISO 17025, GLP, or similar) is a strong plus.
- Excellent communication skills; ability to translate complex technical data concepts for non-technical stakeholders including lab engineers and business analysts.
- Familiarity with version control (Git), CI/CD practices, and cloud data platforms (AWS, Azure, or GCP); experience with containerization (Docker) is a plus.
- Demonstrated experience using AI-assisted development tools to write, debug, and refactor code; comfort evaluating AI-generated outputs for correctness, security, and suitability in a regulated laboratory data environment.
- Understanding of laboratory informatics concepts and data management in accredited or regulated settings.
Technologies
Key technologies include PostgreSQL, MySQL, NoSQL, time-series databases, Python, pandas, SQLAlchemy, FastAPI, Streamlit, React, Plotly Dash, Apache Airflow, Dagster, Prefect, GitHub Copilot, Cursor, Claude Code, Git, Docker, AWS, Azure, GCP, and LIMS.
Benefits
- 401(k)
- Dental insurance
- Flexible spending account
- Health insurance
- Health savings account
- Life insurance
- Paid time off
- Parental leave
- Professional development assistance
- Vision insurance