Data Engineer
Job Description
Scale Marketing is hiring a Data Engineer to expand and optimize its data pipeline architecture, data flow, and extraction/transform/load processes.
Responsibilities
- Create and maintain optimal data pipeline architecture
- Assemble large, complex data sets that satisfy functional and non-functional business requirements
- Identify, design, and implement internal process improvements, including:
- Automating manual processes
- Optimizing data delivery
- Redesigning infrastructure for greater scalability
- Build infrastructure for extraction, transformation, and loading (ETL) from a wide variety of data sources using SQL and big data technologies
- Develop analytics tools that leverage the data pipeline to deliver actionable insights into customer acquisition, operational efficiency, and other business performance metrics
- Collaborate with stakeholders across Partner, Client, and Data teams to support data-related technical issues and data infrastructure needs
- Keep data separated and secure across national boundaries through multiple data centers
- Create data tools that support analytics and data scientists in building and optimizing products
- Work with data and analytics experts to improve the functionality of data systems
Requirements
- Advanced working SQL and experience with relational databases, including:
- Query authoring (SQL)
- Working familiarity with a variety of databases
- Experience building and optimizing big data pipelines, architectures, and data sets
- Ability to perform root cause analysis on internal and external data and processes to address business questions and identify improvement opportunities
- Strong analytical skills with unstructured datasets
- Experience building processes supporting:
- Data transformation
- Data structures
- Metadata
- Dependency and workload management
- Proven history manipulating, processing, and extracting value from large disconnected datasets
- Working knowledge of message queuing, stream processing, and highly scalable big data data stores
- Project management and organizational skills
- Experience supporting and working with cross-functional teams in a dynamic environment
- 5+ years in a Data Engineer role
- Graduate degree in Computer Science, Statistics, Informatics, Information Systems, or another quantitative field
- Experience with big data tools: Apache Hadoop, Apache Hive, Apache Spark, Apache Pig, etc.
- Experience with at least three relational SQL and/or NoSQL databases: MongoDB, HBase, Oracle, SQL Server, Cassandra, Neo4j, Redis, and/or Riak
- Experience with data pipeline/workflow management tools, including: Apache Kafka and/or Apache Spark, Azure DevOps, AWS Data Pipeline
- Experience with at least one cloud service: AWS and/or Microsoft Azure
- Experience with at least one stream-processing system: AWS Kinesis, Azure Stream Analytics, and/or Apache Flink, etc.
- Experience with at least three object-oriented/object function/operating system scripting languages: Python, Java, Scala, C++, Linux/Unix, etc.
Location
- Chicago, IL (onsite)
Technologies
- SQL
- Apache Hadoop, Apache Hive, Apache Spark, Apache Pig
- MongoDB, HBase, Oracle, SQL Server, Cassandra, Neo4j, Redis, Riak
- Apache Kafka, Azure DevOps, AWS Data Pipeline
- AWS, Microsoft Azure
- AWS Kinesis, Azure Stream Analytics, Apache Flink
- Python, Java, Scala, C++, Linux/Unix