Data Analytics Engineer
Job Description
EXL Service is seeking an experienced Data Analytics Engineer to design, build, and improve scalable data pipelines and analytics infrastructure that support financial product decisions. In this onsite role in San Francisco, CA, you will combine ETL/ELT engineering with cloud-native orchestration and strong SQL and PySpark capabilities within a regulated environment.
This position focuses on reliable data movement, transformation, and governance, ensuring analytics-ready data with lineage and auditability. You will also help automate ingestion, optimize performance and cost, and collaborate across data science, analytics, product, and risk or compliance teams.
Responsibilities
- Design, build, and maintain scalable, reliable ETL/ELT pipelines across cloud and on-prem sources, with emphasis on data quality, lineage, and auditability.
- Develop and optimize Python/PySpark and SQL transformations for large, high-volume financial datasets.
- Architect and manage pipeline orchestration using tools such as Airflow, Databricks Workflows, or Step Functions to automate ingestion, transformation, and delivery.
- Create and maintain CI/CD pipelines with GitHub and GitHub Actions for automated testing, deployment, and version-controlled infrastructure changes.
- Build cloud-based solutions on AWS (including S3, Glue, EMR, Redshift, Lambda, IAM) to support analytics, reporting, and downstream ML use cases.
- Deploy and manage pipelines and infrastructure as code, supporting environment promotion, rollback, and monitoring best practices.
- Monitor and troubleshoot pipeline performance, improve query efficiency, and optimize cloud cost across the data stack.
- Work with data scientists, analysts, product, and risk/compliance teams to translate requirements into robust data solutions.
- Apply data governance, security, and regulatory compliance standards appropriate for financial data, including PII, SOX, and PCI.
- Document pipeline architecture, data models, and processes, and support engineering standards and code review practices.
Requirements
- Advanced proficiency in Python for scripting, automation, and data engineering workflows.
- Strong hands-on experience with PySpark for distributed processing at scale.
- Expert-level SQL, including Advanced SQL topics such as window functions, query optimization, complex joins, and performance tuning.
- Solid experience with AWS and cloud-based application/data development, including S3, Glue, EMR, Redshift, Lambda, IAM, and CloudWatch.
- Proven expertise building and orchestrating data pipelines using Airflow, Databricks Workflows, Step Functions, or equivalent solutions.
- Hands-on CI/CD experience using GitHub and GitHub Actions for automated build, test, and deployment.
- Deep understanding of ETL/ELT design patterns, data modeling, and data warehousing concepts.
- Experience deploying infrastructure and pipelines via code (for example, version-controlled deployments).
- Ability to optimize pipeline performance, query execution, and cloud resource or cost efficiency.
- Bachelor’s degree in computer science, Engineering, Data Science, or a related field, or equivalent practical experience.
- 5+ years of experience in data engineering, analytics engineering, or a related technical role.
- Prior experience in banking, fintech, or financial services, with awareness of regulatory and data-security requirements.
Technologies
- Python, PySpark, SQL
- Amazon Web Services (AWS): S3, Glue, EMR, Redshift, Lambda, IAM, CloudWatch
- Airflow, Databricks Workflows, Step Functions
- GitHub, GitHub Actions
- ETL, ELT
- Terraform, CloudFormation
- Kafka, Kinesis, Spark Structured Streaming
- Great Expectations, dbt
- Databricks, Delta Lake, Unity Catalog, notebooks
Preferred / Desirable Skills
- Hands-on experience with Databricks (Delta Lake, Unity Catalog, notebooks, cluster optimization).
- Familiarity with Terraform or CloudFormation for infrastructure as code.
- Experience with streaming data technologies (Kafka, Kinesis, Spark Structured Streaming).
- Exposure to data quality/testing frameworks (Great Expectations, dbt tests).
- Knowledge of dbt for transformation and analytics engineering workflows.
- Understanding of financial data domains such as payments, lending, risk, fraud, or accounting data.
- Relevant certifications including AWS Certified Data Analytics or Solutions Architect, and Databricks Certified Data Engineer.
Soft Skills
- Strong analytical and problem-solving skills with attention to detail and data accuracy.
- Excellent communication skills, including translating technical concepts for non-technical stakeholders.
- Collaborative mindset working cross-functionally with analysts, engineers, and business teams.
- Self-directed ownership of projects end-to-end in a fast-paced, regulated environment.
- Strong ownership mindset around data quality, reliability, and documentation.
Compensation: USD 140,000 - 155,000 per year.