Data Engineer - AI (Spark, Databricks and GCP)
Job Description
Data Engineer - AI (Spark, Databricks and GCP) at Cotiviti, remote, with a salary range of USD 101,000 - 132,000 per year, requires a minimum of 2 years of experience and is responsible for building data pipelines, data management and validation, and collaborating with engineering, DBA and cloud teams to ensure data integrity and operations.
Responsibilities
- Develop, maintain and run intermediate to advanced Spark scripts for data management, validation and integration.
- Develop, maintain and run basic to intermediate SQL scripts for data management and validation.
- Optimize queries to improve the efficiency of daily tasks.
- Perform data analysis and identify issues.
- Collaborate with Engineering, DBA, Cloud Ops and other groups to troubleshoot environmental or network issues impacting work; extend support to after hours or weekends as needed.
- Build and maintain data pipelines as required.
- Validate task results to ensure all requirements are met.
- Adhere to industry and organizational compliance rules to maintain data integrity.
- Track personal productivity and complete task assignments within the department ticketing system by deadline.
- Align with organizational and individual goals as identified in performance reviews and goal setting.
- Complete special projects and other duties as assigned.
- Perform duties with or without reasonable accommodation.
Requirements
- Bachelor’s degree in Computer Science, Information Technology or equivalent work experience.
- 3+ years of working knowledge of big data technologies such as Spark, S3, Kafka, Ray, Hadoop, and related tools.
- 2+ years of working knowledge of big data or cloud technologies including Databricks, AWS, Azure, Hadoop, Spark, Snowflake, etc.
- 3+ years of working knowledge of cloud platforms (AWS, Azure, GCP, OCI, etc.).
- 3+ years of experience with RDBMS (Oracle, MS SQL, Vertica, etc.) and use of SQL, PL/SQL or other data integration/ETL tools.
- Databricks or AWS certifications are a plus.
- Familiarity with data pipeline orchestration tools such as Airflow or Databricks Workflows.
- 3+ years of data analysis experience, preferably in Healthcare enrollment, medical claims or pharmacy claims.
- Proficient with Microsoft Office Suite (PowerPoint, Word, Excel, Outlook).
- Flexible work schedule.
- Experience with project management tools like JIRA.
- Databricks and/or Snowflake environment familiarity is a plus.
Technologies
Spark, Amazon S3, Kafka, Ray, Hadoop, Databricks, AWS, Azure, Snowflake, Google Cloud Platform (GCP), Oracle, MS SQL, Vertica, SQL, PL/SQL, Airflow, Databricks Workflows, JIRA, Microsoft Office Suite
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- Disability insurance
- Life insurance
- 401(k) savings plan
- Paid family leave
- Holidays (9 per year)
- PTO 17-27 days per year
Mental requirements
- Strong analytical skills
- Excellent verbal, listening and written communication skills
- Ability to multitask and prioritize projects to meet deadlines and tight turnaround times
- Ability to work well independently or in a team environment
Working conditions and physical requirements
- Remaining in a stationary position, often sitting for prolonged periods
- Repeating motions that may include wrists, hands, and fingers
- Must be able to provide a dedicated, secure work area
- Must be able to provide high-speed internet access and maintain a suitable office setup