Data Scientist I
Job Description
Memorial Sloan Kettering Cancer Center (MSK) is hiring a Data Scientist I for the Strategy & Growth team. In this hybrid role in New York, you will help design, build, and deploy AI agents and generative AI applications that automate operational workflows and support better decision-making across business initiatives.
The position is focused on early-career impact in areas including large language models (LLMs), retrieval-augmented generation (RAG), analytics, and production reliability. You will work closely with Strategy & Growth stakeholders to translate business needs into scalable AI-enabled workflows.
What you’ll do
- Design, build, and deploy AI agents and generative AI applications from early prototypes through production-ready tools to automate operational workflows and support commercialization and revenue-generation initiatives.
- Assess performance, accuracy, and real-world impact of models and AI solutions using metrics, testing, and monitoring to ensure reliability once deployed.
- Partner with Strategy & Growth stakeholders to convert business challenges into scalable AI-enabled workflows.
- Support broader data analytics and data engineering needs, including building and maintaining ETL pipelines, developing dashboards and reporting tools, and conducting ad hoc data analyses.
- Provide data science support for in-house digital products with potential commercial value.
- Evaluate emerging AI technologies, tools, and frameworks and recommend practical applications for the team.
- Use modern AI development platforms, coding assistance tools, and software engineering best practices to speed up innovation.
What you bring
- 1 to 3 years of experience in data science, machine learning, or applied AI.
- Hands-on experience with large language models including Claude, GPT-4, and open-source LLMs, with exposure to prompt engineering, retrieval-augmented generation (RAG), and/or API-based development.
- Understanding of machine learning fundamentals and data analysis techniques.
- Proficiency in Python and/or R for data science and ML workflows.
- Comfort with relational databases and SQL.
- Familiarity with coding assistant tools (e.g., Claude Code, GitHub Copilot), modern software development practices, and version control with Git.
- Exposure to or interest in agentic AI systems (multi-step agents, tool use, orchestration frameworks such as AWS Strands and LangChain). Direct experience is a strong plus.
- Familiarity with cloud platforms, especially AWS (e.g., Bedrock) is a plus.
- Experience building AI applications, AI agents, RAG solutions, or productivity-focused automation tools is strongly preferred.
Technologies you may work with
- Claude, GPT-4, open-source LLMs
- Prompt engineering, retrieval-augmented generation (RAG), API-based development
- Python, R, SQL
- Claude Code, GitHub Copilot, Git
- AWS Strands, LangChain, AWS, Bedrock
Example projects
- Build an AI agent to automate patent operations and compliance workflows.
- Develop an LLM pipeline to automate chart review by extracting outcomes (such as whether a patient had a change in diagnosis) from unstructured clinical notes.
- Build a RAG pipeline that extracts insights from a large portfolio of prior strategic research and interviews.
Schedule and location
- Schedule: 9:00 AM – 5:00 PM EST, Monday through Friday
- Location: Hybrid; 1x a week on site with additional flexibility as needed; 633 Third Avenue, New York, NY
- Reporting to: Director, Data Science
Compensation
- Annual pay range: USD 113,100 - 180,900
- FSLA status: Exempt