Founding Data Analytics Engineer
Job Description
Broccoli AI is hiring a Founding Data Analytics Engineer to build the unified data layer, own the architecture and tooling, and ensure every dashboard and metric is trustworthy.
Responsibilities
- Build and run ClickHouse ingestion pipelines to reliably move data from all sources into ClickHouse, including selecting the tooling and owning the end-to-end flow.
- Model and document data by transforming raw feeds into clean, well-defined tables, including entity resolution so a single customer is consistent across billing, support, and call data.
- Create a source-of-truth library with canonical views and metric definitions that dashboards and analyses can rely on.
- Make data Human & AI-ready by structuring models, definitions, and documentation so both people and AI agents can query and get correct answers, plus building internal tools that enable others to ask data questions with confidence.
- Ensure data trustworthiness through freshness checks, quality tests, and alerts to detect pipeline failures before customers are impacted.
- Support deep dives via ad-hoc analysis, segment investigations, and responses to partner questions.
- Partner with engineering to understand system storage and production details (schemas, events, architecture) and provide early input so warehouse data is usable, stable, and straightforward to model.
Requirements
- 4β8+ years in data or analytics engineering, including building and operating production pipelines end to end and being responsible when pipelines break.
- Strong SQL and solid Python, with hands-on experience using ETL/orchestration tooling such as Airbyte, Fivetran, Dagster, dbt, or custom ETL.
- Columnar/OLAP warehouse experience, with ClickHouse preferred (BigQuery, Snowflake, or Redshift also acceptable).
- Data modeling as a craft, with experience designing tables others query and a focus on meaning and metrics, not only pipeline reliability.
Technologies
- ClickHouse
- BigQuery
- Snowflake
- Redshift
- SQL
- Python
- Airbyte
- Fivetran
- Dagster
- dbt
Nice to have
- Self-directed ownership, such as being the first or only data person or building a data platform from scratch.
- ClickHouse-specific experience, including materialized views and performance tuning on event-scale data.
- Multi-source identity / entity resolution experience.
- Customer-facing or multi-tenant analytics experience, including strict customer-level data isolation.
- B2B SaaS operational data experience such as calls, bookings, jobs, billing, or CRM/field-service data (for example, ServiceTitan).