Senior Data Scientist
Job Description
Remote role for a Senior Data Scientist focused on data acquisition, data quality, and building validated data science and AI/ML models.
Responsibilities
- Lead data acquisition efforts, ensuring datasets are properly formatted, accurately described, and aligned with GitHub privacy policies
- Mentor others in data cleaning and data analysis best practices
- Identify gaps in existing datasets and drive onboarding of new datasets from production systems or third-party vendors
- Collaborate with relevant teams to resolve data integrity issues and support upstream change for long-term quality
- Use expertise in data modeling, AI/ML tools, programming languages, and query languages to:
- Create models
- Conduct experiments
- Analyze results
- Evaluate methodology and performance of team model outputs
- Recommend improvements
- Drive best practices for model validation, implementation, and application
- Partner across the organization to identify and explore opportunities for transformative solutions for stakeholders and customers
- Develop and communicate data-driven strategies aligned to business priorities
- Lead conversations with end customers and/or internal stakeholders to define and solve business problems
- Communicate complex statistics and machine learning topics to diverse audiences (multidisciplinary teams, customers, technical and non-technical)
- Independently write efficient, readable, extensible code across multiple features or solutions
- Contribute to code and model review by providing feedback and implementation improvement suggestions
- Drive operational excellence for model deployment, including performance, scalability, monitoring, maintenance, integration into engineering production systems, and stability
- Produce project plans to define steps needed for completion and deliver measurable improvement in business performance metrics over time
Requirements
- Bachelor’s Degree in Data Science, Mathematics, Physics, Statistics, Economics, Operations Research, Computer Science, or related field AND 5+ years experience in data science (e.g., managing structured and unstructured data, applying statistical techniques)
- OR Master’s Degree in the same fields AND 3+ years experience in data science (e.g., managing structured and unstructured data, applying statistical techniques)
- Technical understanding of data science techniques including regression, classification, time-series analysis, experimental design, and causal inference
- Proficiency in Python or R
- Experience with query languages such as SQL and KQL
- Experience with data manipulation tools such as Spark and Airflow
- Ability to communicate findings to non-technical stakeholders using storytelling and visualization with tools such as Jupyter notebooks or Azure Data Explorer/PowerBI dashboards
Technologies
- Python
- R
- SQL
- KQL
- Spark
- Airflow
- Jupyter notebooks
- Azure Data Explorer
- PowerBI
Benefits
- Annual bonus
- Stock
- Sales incentives based on revenue or utilization (depending on the plan terms and employee role)
- Benefits to support you, wherever you are
GitHub Leadership Principles
- Customer-obsessed
- Ship to learn
- Growth mindset
- Own the outcome
- Better together
- Diverse and inclusive
- Model
- Coach
- Care
- Create clarity
- Generate energy
- Deliver success
Compensation
Base salary range: USD 124,000.00 - USD 329,200.00 per year.