Data Scientist Cancer Neuroscience
Job Description
Support cancer neuroscience research by building reliable data infrastructure, analytical workflows, and predictive modeling capabilities for CNP projects.
Responsibilities
- Create and maintain a structured, up-to-date inventory of neuro-focused datasets, cohorts, registries, biospecimen-linked datasets, clinical data assets, imaging datasets, patient-reported outcomes, molecular datasets, and other relevant research resources.
- Document institutional data asset attributes, including dataset ownership, access requirements, data dictionaries, cohort definitions, refresh cadence, data limitations, analytical readiness, and potential research use cases.
- Act as a technical connector between CNP-supported investigators, neuro-focused research initiatives, institutional data teams, the Institute for Data Science in Oncology, the institutional Research Data Office, and other enterprise stakeholders to improve awareness, responsible use, and strategic alignment.
- Find opportunities to harmonize, link, reuse, and scale CNP-aligned datasets while reducing duplicative data collection and supporting governance, documentation, and interoperability standards.
- Help CNP leadership use inventory and analytical outputs for project prioritization, infrastructure planning, collaboration opportunities, grant development, philanthropy reporting, and long-range strategic planning.
- Design, build, maintain, and document data pipelines for prospective and retrospective oncology cohorts, including clinical, research, operational, and multimodal datasets.
- Collect, clean, harmonize, preprocess, and validate data from multiple sources to support accuracy, completeness, usability, and reproducibility.
- Develop and maintain transparent, auditable, and well-commented code bases for routine and advanced data science tasks.
- Create reusable workflows, templates, and documentation that support efficient data access, analysis, reporting, and collaboration across CNP-supported projects.
- Work with large datasets using Python, R, SQL, Spark, Foundry, Microsoft Fabric, AWS, Azure, GCP, or comparable environments.
- Conduct exploratory data analysis to identify trends, patterns, anomalies, missingness, bias, and potential signals relevant to cancer neuroscience research.
- Develop, test, validate, and refine predictive and prescriptive models using statistical methods, machine learning, artificial intelligence, and other advanced techniques.
- Design and implement analytical experiments to validate hypotheses, evaluate model performance, and support scientific decision-making.
- Provide onboarding, hands-on training, and technical support to research staff and project teams using Foundry and related institutional platforms, including best practices for access, workflow development, documentation, and collaboration.
- Assist users with analytical question definition, technical issue resolution, workflow improvements, and translating research needs into feasible data science approaches.
- Collaborate with clinicians, laboratory scientists, data scientists, data engineers, statisticians, bioinformaticians, research staff, and program leadership to optimize data insights and support CNP priorities.
- Communicate findings via visualizations, dashboards, reports, presentations, and clear written summaries for technical and non-technical audiences.
- Collaborate on scientific publications, abstracts, grant applications, technical reports, and other documentation for faculty, leadership, advisory boards, philanthropy, and project teams.
- Prepare technical reports and other documentation for upper management and team members for short- and long-range projects and planning.
- Contribute to institutional data integration by aligning CNP-specific data needs with broader IDSO-supported platforms, standards, and analytical capabilities.
- Ensure data governance, privacy, security, documentation, and compliance standards are followed in all data science activities.
- Stay current with emerging data science, machine learning, artificial intelligence, clinical informatics, and big data technologies.
- Other duties as assigned.
Requirements
- Bachelor's degree in Biomedical Engineering, Electrical Engineering, Computer Engineering, Physics, Applied Mathematics, Science, Engineering, Computer Science, Statistics, Computational Biology, or related field.
- Three years of scientific software or industry development/analysis experience.
- With a Master's degree: one year required experience.
- With a PhD: no experience required.
Technologies
- Python, R, SQL, Spark
- Foundry
- Microsoft Fabric
- AWS, Azure, GCP
Benefits
- Medical
- Dental
- Paid time off
- Retirement
- Tuition benefits
- Educational opportunities
- Individual and team recognition
Preferred
- Experience in academic healthcare, oncology, cancer research, clinical research, translational research, or biomedical research.
- Experience working with EPIC-derived clinical data, electronic health record data, prospective or retrospective cohorts, clinical registries, institutional research datasets, real-world clinical data, or multimodal biomedical datasets.
- Experience with Foundry or similar enterprise data platforms.
- Experience developing analytical workflows, dashboards, technical reports, reusable data products, or data pipelines for research or operational stakeholders.
- Experience contributing to peer-reviewed publications, abstracts, grant applications, scientific presentations, philanthropy reports, or institutional strategy documents.
- Experience with machine learning, natural language processing, large language model-enabled workflows, or AI applications in healthcare or research settings.
- Experience working collaboratively with clinicians, scientists, statisticians, bioinformaticians, data engineers, research staff, and institutional data teams.
Education
- Preferred: Master's degree or PhD in Science, Engineering, Data Science, Computer Science, Statistics, Biomedical Informatics, Computational Biology, Bioinformatics, Biostatistics, Public Health, or related field.
Additional Information
- Requisition ID: 183428
- Employment Status: Full-Time
- Employee Status: Regular
- Work Week: Days
- Work Location: Houston, TX (onsite)
- Work Location: Hybrid Onsite/Remote
- Pivotal Position: Yes
- Referral Bonus Available?: Yes
- Relocation Assistance Available?: Yes
- FLSA: Exempt and not eligible for overtime pay
- Fund Type: Soft
- Salary: USD 106,500 - 159,500 per year
- Minimum Salary: USD 106,500
- Midpoint Salary: USD 133,000
- Maximum Salary: USD 159,500
- #LI-Hybrid