Data Engineer (Databricks), Assistant Vice President
Manager
Big Data
Bigdata
CI/CD
Cloud
Cloud Infrastructure
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Governance
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Security
Data Warehouse
Database
Databases
Databricks
Databricks Workflows
Delta Lake
DevOps
Devops Tools
DevSecOps
Engineering
ETL
Informatica
Information Technology (IT)
Infrastructure As Code
Lakehouse
Platform Engineering
Programming Language
Programming Languages
Pyspark
Security Automation
Software Development
Spark
Spark Sql
SQL
Job Description
State Street is hiring an Assistant Vice President to design, build, and support a Legal Data Lakehouse platform using AWS and Databricks. In this role, you will focus on scalable data pipelines, governance, and trusted data capabilities that support legal operations, compliance analytics, reporting, and AI/ML use cases, all within enterprise security and audit expectations.
What you’ll do
- Design, build, and maintain scalable data pipelines using PySpark, Python, and Spark SQL
- Develop and optimize ETL/ELT workflows on Databricks using Delta Lake
- Implement lakehouse architecture with Bronze/Silver/Gold layers for enterprise data platforms
- Build and manage Databricks Jobs, Workflows, and Notebooks for batch and streaming workloads
- Create reusable frameworks for data ingestion, processing, and orchestration
- Containerize data workloads with Docker and automate processes using scripting
- Integrate Databricks pipelines with Power Platform solutions, including Power Apps and Power Automate
- Enable data exposure for business users through APIs, connectors, and curated datasets
- Integrate data from enterprise sources including SQL Server and Oracle
- Apply data modeling techniques such as dimensional modeling and implement partitioning and optimization strategies
- Work with structured and semi-structured data such as JSON and Parquet
- Tune performance using caching, indexing, and Spark optimization techniques
- Publish curated datasets for consumption in Power BI dashboards and Power Apps/Power Automate workflows
- Support data quality through unit testing, validation frameworks, and automated checks
- Monitor and troubleshoot distributed Spark workloads and production pipelines across Databricks and cloud environments
- Maintain data lineage, consistency, and audit readiness
- Collaborate with Legal, Security, Compliance, and Enterprise Data teams to deliver scalable solutions
- Translate business requirements into robust data engineering designs, and act as a Subject Matter Expert (SME) in Databricks and lakehouse architecture
- Lead initiatives with minimal supervision and take ownership of deliverables
- Implement and support governance using Databricks Unity Catalog and AWS controls such as IAM and KMS
- Ensure adherence to data privacy and regulatory requirements, including GDPR, as well as internal security and audit standards
- Design and maintain data access controls along with data classification and handling standards, working with IAM and security teams
- Design and maintain CI/CD pipelines using Harness, Azure DevOps, or GitHub
- Automate deployment of Databricks assets using Databricks Repos and the Databricks CLI
- Monitor, schedule, and optimize workflows using Databricks orchestration tools
- Maintain clear documentation including architecture, data flows, and runbooks
- Continuously improve performance, scalability, and cost efficiency
Requirements
- 8+ years of experience in Data Engineering or data platform development
- Strong hands-on experience with Databricks and Apache Spark
- Proficiency in PySpark, Python, and SQL
- Experience with AWS data platform services, including S3, Glue, Lambda, and IAM
- Experience working with Delta Lake and lakehouse architecture
- Solid understanding of distributed data processing
- Solid understanding of ETL/ELT frameworks
- Solid understanding of data modeling techniques
Technologies
- Databricks, AWS, PySpark, Python, Spark SQL
- ETL/ELT, Delta Lake, Lakehouse architecture (Bronze/Silver/Gold layers)
- Databricks Jobs, Databricks Workflows, Notebooks
- Docker, Power Platform (Power Apps, Power Automate), APIs, Connectors, Power BI
- SQL Server, Oracle, JSON, Parquet
- Unit testing, Validation frameworks, Databricks Unity Catalog
- IAM, KMS, GDPR
- CI/CD pipelines with Harness, Azure DevOps, or GitHub
- Databricks Repos, Databricks CLI, S3, Glue, Lambda
Benefits
- 401K with company match
- Insurance coverage including basic life, medical, dental, vision, long-term disability, and other optional additional coverages
- Paid-time off including vacation, sick leave, short term disability, and family care responsibilities
- Access to the Employee Assistance Program
- Incentive compensation including eligibility for annual performance-based awards (excluding certain sales roles subject to sales incentive plans)
- Eligibility for certain tax advantaged savings plans
Location: Quincy, MA (onsite)
Salary: USD 110,000 - 177,500 per yearly
Minimum experience: 8 years
Education: Bachelor’s or Master’s degree in computer science, Data Engineering, Information Systems, or a related technical discipline
Preferred qualifications
- Hands-on experience with Databricks platform components including Delta Lake, Workflows, and Unity Catalog
- Strong experience building end-to-end data pipelines (batch and streaming) using AWS and Databricks
- Familiarity with performance optimization techniques in Spark and Delta Lake
- Experience supporting analytics, reporting, or AI/ML use cases on a lakehouse platform
- Understanding of data governance, metadata management, and security controls
Nice to have
- Experience in Legal, Compliance, Financial Services, or regulated industries
- Understanding of legal data constructs such as contracts, clauses, obligations, and matters
- Exposure to unstructured data processing or document/NLP pipelines
- Experience with Power BI, Power Apps, or Power Platform
- Experience handling sensitive data in audit-driven environments