Machine Learning Engineer
Job Description
WireScreen seeks a Machine Learning Engineer in New York, NY (hybrid) to advance entity resolution and knowledge graph development, scale data ingestion, and deploy ML models across millions of records, reporting to the VP of Engineering and collaborating with data, enrichment, and product teams.
Responsibilities
- Refine current entity resolution algorithms to reveal hidden links between individuals and organizations in China
- Enrich the knowledge graph by integrating alternative data sources to map power structures in China
- Train, validate, and deploy ML models handling tens of millions of records daily
- Collaborate with Product to define and implement evaluation frameworks for traditional ML and agent-based systems
- Integrate agent workflows into internal tools to boost Research team scale and speed
Requirements
- 4+ years of experience tackling clustering oriented ML problems, preferably in knowledge graphs or entity resolution; other relevant domains include recommendation systems, cohort analysis, and anomaly/outlier detection
- End-to-end ML model production experience, including building and operating a service from experimentation and training through testing, deployment, and ongoing maintenance; model families may include clustering, classification/regression, dimensionality reduction and embeddings, nearest-neighbor or similarity methods (e.g. KNN, SVM), ensembles, NLP, and deep learning
- Strong proficiency in Python and SQL
Technologies
- Python
- SQL
- PySpark
- Temporal
- FastAPI
- Scikit-learn
- NumPy
- Docker
- Terraform
- Kubernetes
Benefits
- Competitive compensation including salary, equity, and rapid growth potential
- 100% company-paid Medical, Dental, and Vision coverage for employees
- FSA, HSA, and 401(k) options to help you plan for healthcare expenses and retirement
- Generous paid time off plus company-wide holidays
- Pre-tax commuter benefits to help you save on transit and parking
- Hybrid office schedule designed to give you flexibility while staying connected with your team
NICE TO HAVE
- Experience with frontier or state-of-the-art models and fine-tuning LLMs for task-specific applications
- Experience handling large, heterogeneous unstructured datasets and or with semantic search, computer vision (especially OCR), or linear optimization problems
- Familiarity with PySpark, Temporal, FastAPI, Scikit-learn, NumPy, Docker, Terraform, Kubernetes
- Early-stage startup experience (Series B or earlier)
- B2B SaaS experience
- #LI-LG1