Data Scientist II - RiskOS
Job Description
This role supports Socure’s RiskOS Workforce Verification vertical by delivering an end-to-end data science pipeline to identify hiring and workforce identity fraud. The position also includes GenAI and NLP components for resume verification and explanation over unstructured text.
Key Responsibilities
- Own the full data science lifecycle for Workforce Verification use cases on RiskOS, covering data exploration and hypothesis generation through model development, evaluation, deployment, and monitoring.
- Explore and analyze workforce-related data sources including applications, resumes, device and behavioral telemetry, background checks, and ATS/HRIS integrations to detect patterns of workforce fraud.
- Design, implement, and iterate on rules, conditions, and heuristic logic in RiskOS workflows to identify high-risk workforce events, such as repeated identities across multiple resumes, suspicious device patterns, and anomalous hiring flows.
- Develop and evaluate machine learning models for workforce risk and identity assessment, including fraud risk scoring, clustering related identities, and anomaly detection across hiring funnels, using Socure’s broader identity and device signals where applicable.
- Collaborate with RiskOS and Workforce product teams on GenAI-powered features like the Resume Verification Agent and explanation agents by defining input/output requirements, building evaluation datasets, and establishing quantitative and qualitative evaluation frameworks for LLM components.
- Work closely with engineering to productionize models, rulesets, and GenAI components within RiskOS by defining interfaces, supporting integration and testing, and contributing to monitoring, alerting, and feedback loops.
- Translate model and rule performance into clear customer-facing narratives with product, Workforce GTM, and solution consulting, including outcomes such as blocking fake applicants and reducing deepfake interviews or identity rental in hiring.
- Incorporate customer feedback and outcome data to continuously improve Workforce Verification logic and models, and support experimentation and offline “test harness” design for safe evaluation of new workflows and templates.
- Operate with a product mindset by documenting assumptions, decisions, and evaluation results, communicating trade-offs clearly, and proactively surfacing risks, limitations, and opportunities.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Engineering, or a related quantitative field, or equivalent practical experience.
- 3–6 years of hands-on experience in data science, machine learning, or applied analytics, with meaningful experience in fraud, risk, trust & safety, or workforce or hiring analytics preferred.
- Experience owning end-to-end analytics and/or model development projects, including problem framing, data wrangling, feature engineering, model training, evaluation, and deployment support.
- Strong proficiency in Python and SQL, including experience with common data science and ML libraries such as pandas, scikit-learn, XGBoost, PySpark, or similar.
- Comfort working with large, messy, heterogeneous datasets including JSON workflows, logs, event streams, and third-party enrichments, plus the ability to create reusable abstractions or utilities.
- Exposure to Natural Language Processing and/or unstructured text analytics such as resume or document parsing, entity extraction, similarity search, or basic embedding-based methods applied in real-world products.
- Some hands-on experience with Generative AI or LLM-based products (for example, commercial LLM APIs, prompt design, RAG-style retrieval, or evaluation of LLM outputs), with interest in further developing these skills.
- Strong analytical and problem-solving skills, including reasoning about ambiguous signals and adversarial behavior in fraud or workforce contexts.
- Ability and willingness to perform light data engineering or production-oriented tasks as needed in collaboration with engineering, such as building ETL transforms, contributing to Airflow/Spark jobs, or instrumenting basic monitoring.
- Clear, concise communication skills for explaining complex analyses, models, and GenAI behavior to non-technical stakeholders such as product, GTM, and customers.
- A bias toward ownership, learning, and collaboration in a fast-paced, evolving environment, with comfort receiving guidance from senior data scientists while expanding scope and autonomy.
Technologies
Python, SQL, pandas, scikit-learn, XGBoost, PySpark, Natural Language Processing, GenAI, LLMs, RAG-style retrieval, ETL, Airflow, Spark
Nice to Have
- Direct experience with workforce, HR tech, ATS/HRIS data, or hiring funnel analytics.
- Prior work on identity verification, device intelligence, or orchestration and rules engines (such as RiskOS or similar systems).
- Familiarity with evaluation and monitoring of GenAI systems, including offline benchmarks, human-in-the-loop review, and safety or hallucination checks.
Role Details
- Location: Miami, FL (onsite)
- Experience: 3+ years
- Compensation: USD 140,000 - 170,000 per year