Principal Machine Learning Engineer
Job Description
The Principal Machine Learning Engineer role at Oracle focuses on implementing and productionizing machine learning models. This position owns deployment readiness and drives end-to-end operational excellence across monitoring, troubleshooting, workflow automation, and continuous improvement.
Key Responsibilities
- Implement machine learning (ML) models for production use and ensure deployment readiness.
- Automate ML workflows, including data extraction, transformation, and loading (ETL), as well as model deployment and monitoring, to support continuous integration and continuous delivery of ML solutions.
- Build infrastructure and frameworks to monitor model performance in deployment and alignment with design criteria.
- Proactively monitor deployed models and troubleshoot independently or in collaboration with Data Science.
- Evaluate potential data quality, security, and privacy issues and assess impacts on modeling and downstream data analysis.
- Support troubleshooting and debugging for ML infrastructure and workflows, including addressing root causes and implementing robust solutions to prevent recurrence.
- Transform ML prototypes into production-ready models, scaling models and cleaning model code to meet production quality standards.
- Develop novel metrics that provide analytical insights into model operating performance for non-technical stakeholders.
- Perform data cleaning, preprocessing, and feature identification to prepare data for model training.
- Collaborate with stakeholders to integrate ML models into new or existing systems, including Development Leads, Product Management, Operations, and Release Management.
- Maintain the partnership between model development and operations to enable smooth deployment and continuous model improvement.
- Develop, maintain, and refine tools, platforms, environments, and services for internal use, including professional documentation for technical processes (experimentation, data collection and analyses, and model building).
- Produce efficient, bug-free medium-complexity code from scratch, properly maintain and organize the existing codebase, and test and review code for defects.
- Apply best practices for version control, code review, and code delivery/deployment.
- Keep current with developments in the machine learning field and integrate relevant knowledge into model development, including evaluating third-party ML frameworks and libraries.
- Manage and coordinate moderately complex tasks, monitoring timelines and deliverables, and delegate, monitor, and prioritize work across multiple projects with technical oversight.
- Leverage understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet needs, while seeking diverse perspectives to support inclusivity.
- Identify and address moderately complex issues by analyzing a wide range of information in alignment with standard practices, and proactively escalate unresolved or critical issues with thorough assessment and solution suggestions.
- Review, contribute to, and document problem-solving strategies; pursue continuous learning and proactively seek feedback and training.
- Coach and mentor junior team members and support knowledge sharing across teams.
- Recommend and collaborate on process improvements, evaluate their impact on key stakeholders, and solicit feedback on alternative approaches.
- Contribute to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Required Technologies
- PyTorch
- TensorFlow
- Keras
Work Location
United States (onsite)