Senior AI Data Scientist
Job Description
Knowtion Health is seeking a remote Senior AI Data Scientist to lead end-to-end data science and ML engineering for revenue-cycle use cases, spanning model development, data pipelines, and MLOps in production.
Responsibilities
- Translate complex revenue-cycle challenges into tractable modeling problems and own the end-to-end lifecycle, including claim and invoice viability scoring, denial and underpayment prediction, revenue/ARR forecasting, and work prioritization, from hypothesis through production-ready, monitored results.
- Develop and validate models with solid statistical foundations, including thoughtful feature design, handling of missing data and class imbalance, calibration, benchmarking against simple baselines, and transparent treatment of uncertainty.
- Create and sustain data pipelines powering these models, integrating source systems, billing and EHR data, and internal data stores; manage data collection, cleansing, field mapping, and normalization of payer and plan data with a focus on reliability and data quality.
- Own the MLOps program, including model monitoring and drift detection (population stability), calibration and thresholding, benchmarking against simple baselines, explainability (SHAP and interpretable coefficients), experiment and version tracking, and documentation such as model cards.
- Collaborate with subject-matter experts to encode business rules and heuristic logic alongside statistical models when it yields greater accuracy or defensibility.
- Gather and document business requirements for revenue-cycle initiatives, covering workflows, decision points, stakeholder objectives, operational constraints, success criteria, data availability, and underlying assumptions; maintain traceability as requirements evolve.
Requirements
- Multiple years of building, validating, and deploying models in production, with clear ownership of at least one model relied upon by a product or business function.
- A quantitative degree in statistics, computer science, mathematics, or data science, or equivalent demonstrated ability.
- Strong statistical foundation in experimental design, inference, evaluation of imbalanced real-world data, and distinguishing real signals from artifacts.
- Fluency in core modeling fundamentals: train/test discipline, overfitting and regularization, data leakage, class imbalance, calibration, and honest model evaluation.
- Proficiency in production-grade Python (pandas, NumPy, scikit-learn; deep-learning frameworks are a plus) and strong SQL, including stored procedures and performance-aware queries against large tables, with Git and reproducible workflows.
- Hands-on MLOps experience: monitoring, drift detection, experiment tracking, versioning, and maintaining model health in production.
- Comfort owning the data pipeline end-to-end, including source-system integration, scheduling, and data-quality controls.
Technologies
- Python
- pandas
- NumPy
- scikit-learn
- deep-learning frameworks
- SQL
- Git
- SHAP
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Short-term disability
- Long-term disability
- Paid holidays
- 401k
- Generous PTO
PHOTO IDENTIFICATION AND IMAGE CAPTURE NOTICE
- As part of our hiring process, candidates are required to present a valid government-issued photo ID during screening and/or interviews.
- We may also capture a photograph during the screening or interview process for identity verification, security, and record-keeping purposes. All identification documents and images will be handled in accordance with applicable privacy laws and accessed only by authorized personnel.
- By proceeding with your application, you acknowledge and consent to this process.