Senior Machine Learning Engineer
Job Description
Senior Machine Learning Engineer at Siemens in Raleigh, NC (onsite), leading LLM powered application development on AWS and owning the end-to-end model lifecycle with MLOps.
Responsibilities
- Develop LLM driven applications by architecting retrieval augmented generation pipelines, coordinating prompts, building tools and agents, implementing safety controls, and creating evaluation harnesses; measure latency, cost, and quality.
- Oversee the full ML lifecycle from data curation and feature engineering to training or fine tuning (LoRA or QLoRA), A/B testing, deployment, monitoring, and ongoing enhancement of models and prompts.
- Deploy scalable AWS services using EKS, ECS and Lambda; leverage SageMaker, Bedrock, EMR, MSK, and Step Functions; implement observability with CloudWatch and OpenTelemetry and apply cost-control practices.
- Establish MLOps and governance practices including CI/CD for models, model/version registries, data and prompt lineage, evaluation gates, and responsible AI controls.
- Collaborate with Brightly across asset-management use cases to design ML and LLM solutions, working with product managers and UX to ship customer‑facing features that improve reliability, safety, and sustainability.
- Conduct exploratory data analysis across structured, semi-structured, and unstructured data to reveal patterns, correlations, feature importance, and data quality issues.
- Perform in-depth research on asset-related, operational, and domain datasets to uncover root causes, trends, and predictive signals.
- Maintain a pragmatic, product‑oriented approach with a bias toward measurable outcomes and rapid iteration with stakeholders.
- Practice engineering excellence by delivering production‑quality Python, designing reliable APIs and services, and upholding testing and observability standards.
- Provide collaborative leadership by mentoring peers and influencing architectural decisions across teams.
Requirements
- 8 to 10 years of overall software or ML engineering experience, including at least 2 years building and operating production ML systems.
- More than 1 year of hands-on LLM application development (RAG, fine‑tuning, prompt engineering, evaluators/guardrails, agentic workflows) using Langchain and Langgraph.
- 3+ years of AWS proficiency, with core services (EKS/ECS, Lambda, S3, DynamoDB or RDS, Step Functions, IAM) and ML stack familiarity (SageMaker, Bedrock or HF on AWS).
- Modeling and frameworks expertise with Python, PyTorch, and Hugging Face ecosystem; experience with vector stores (OpenSearch, PGVector, Pinecone), embeddings, retrieval, and NLP/LLM evaluation metrics.
- MLOps experience including CI/CD for ML, model registries, experiment tracking, telemetry/monitoring, and automated retraining; Docker/Kubernetes and CI pipelines (GitHub Actions or GitLab CI).
- Data engineering fluency covering ETL/ELT, streaming and batch processing (Spark/Flink), and data quality and governance controls for ML.
Technologies
- Python
- PyTorch
- Hugging Face ecosystem
- Langchain
- Langgraph
- SageMaker
- Bedrock
- OpenSearch
- PGVector
- Pinecone
- Spark
- Flink
- AWS
- EKS
- ECS
- Lambda
- S3
- DynamoDB
- RDS
- Step Functions
- IAM
- EMR
- MSK
- CloudWatch
- OpenTelemetry
- Docker
- Kubernetes
- GitHub Actions
- GitLab CI
- MLflow
- Kedro
- SageMaker Pipelines
Nice to Have
- Experience with distributed training (FSDP, DeepSpeed), RLHF, or optimization on Inferentia/Trainium.
- Exposure to sustainability, asset management, or intelligent operations domains.
- Security and compliance knowledge for enterprise ML systems.