Associate Data Scientist
Job Description
National Computing Group seeks an entry-level Associate Data Scientist for training, evaluating, and managing machine learning models on an enterprise AI platform within a large-scale manufacturing setting; hybrid work is available near Pittsburgh, NE Minnesota, NW Indiana, or St. Louis.
Responsibilities
- Learn the enterprise AI platform architecture and how it connects to manufacturing data across the facility network.
- Train machine learning models on industrial datasets including process sensor data, quality measurements, and operational metrics.
- Evaluate model performance using appropriate statistical methods and determine readiness for production, selecting champion models.
- Manage the full model lifecycle from training and validation to deployment, monitoring, and retraining.
- Collaborate with data engineers to specify data transformations and feature engineering needs.
- Create visualizations and reports that communicate model outputs and insights to operations and leadership.
- Process, cleanse, and verify the integrity of data used for analysis.
- Perform ad hoc analyses on structured and unstructured datasets and present findings clearly.
- Stay current on developments in machine learning, industrial AI, and related platform capabilities.
- Design and execute model training experiments using robust train/validation/test methodologies.
- Compare candidate models across metrics such as accuracy, precision/recall, RMSE, and drift.
- Document evaluation rationale and maintain clear records of model versions and decisions.
- Recommend champion model promotions and flag underperforming models for retraining or replacement.
- Develop an understanding of how model outputs influence real manufacturing decisions and weigh that context in evaluations.
Requirements
- Bachelor's or Master’s degree in Data Science, Statistics, Computer Science, Engineering, Mathematics, or a related quantitative field
- Solid foundation in machine learning concepts: supervised/unsupervised learning, model evaluation, overfitting, cross-validation
- Proficiency in Python, including core data science libraries (pandas, NumPy, scikit-learn)
- Understanding of statistical fundamentals: hypothesis testing, distributions, regression, and model diagnostics
- Exposure to data visualization tools or libraries (Tableau, Power BI, Matplotlib, Seaborn, or similar)
- Strong analytical thinking and the ability to approach ambiguous problems methodically
- Excellent communication skills; able to explain what a model does and why it matters to non-technical stakeholders
- Intellectual curiosity and eagerness to learn in an industrial environment that may be new to you
- Ability to manage your own time and work independently while collaborating within a broader team
- At least 1 year of hands-on machine learning experience
- At least 1 year of programming experience
Technologies
- Python
- pandas
- NumPy
- scikit-learn
- Tableau
- Power BI
- Matplotlib
- Seaborn
- SQL
- Hadoop
- Spark
- data lakes
About the Role
This entry level opportunity supports a motivated, analytically strong new graduate to grow into the enterprise data science function for a major steel producer. You will receive training to become the dedicated expert on the enterprise AI platform, the system that drives data driven decision making across multiple facilities.
You will not be required to know everything at the outset; you will learn quickly, ask thoughtful questions, and focus on obtaining accurate answers. Over time, this role will own the core function of training, evaluating, and managing the machine learning models powering the enterprise AI platform, determining which models are production ready (champion models) and driving ongoing improvements.
Why This Role
Many early career data scientists work on isolated models with limited real world impact. This role, by contrast, connects your work directly to operations running 24/7 in a highly data rich industrial environment. You will develop deep expertise in enterprise AI platform operations at scale across multiple facilities, with meaningful ownership.