Persona AI Inc is building robotics intelligence from multimodal data. In this Staff, Robotics ML/Data Engineer role (onsite in Houston, TX), you will architect and scale end-to-end robotics data pipelines that convert raw in-the-wild egocentric video and dense sensor streams into high-fidelity training assets for downstream foundation model training. The work is hands-on, pipeline-focused, and designed to improve data quality, validation rigor, and researcher visibility.
What youβll work on
- Cross-modal validation systems that confirm agreement across video, proprioception, force/haptic signals, and language annotations, including consistency checks such as reprojection of robot state into the image plane and VLM-assisted verification that instructions match observed behavior.
- Robust orchestration of multimodal modules for hand tracking, segmentation, depth estimation, 3D reconstruction, and pose tracking. You will also retarget human demonstrations into robot trajectories and run simulation-in-the-loop validation (kinematic feasibility, physics replay, motion-consistency filtering) to keep synthesized data physically grounded.
- Data augmentation for reliability using spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection to expand expert trajectories and improve model robustness.
- Unified state-action representations across differing embodiments, coordinate frames, rotation conventions, gripper/hand parameterizations, and sampling rates, with per-dimension validity masking and per-source normalization. The goal is adding new robots or sensors via configuration rather than rewrites.
- Dataset tooling and audit workflows that enable researchers to query, visualize, and audit datasets (clip browsers, trajectory viewers, annotation review UIs). You will translate model-failure analyses into new curation rules and targeted re-collection requests.
- End-to-end ingestion pipelines that take raw, unstructured recordings (egocentric video, teleoperation sessions, third-party open datasets) and produce indexed, queryable, training-ready datasets. This includes temporal segmentation into action clips, metadata and scene-graph extraction, embedding-based retrieval, and language annotation workflows.
What you bring
- M.S. or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field.
- Deep expertise in Python and extensive experience with PyTorch, especially implementing custom dataloaders for multimodal datasets.
- Experience processing complex time-series data from force-torque (F/T) sensors, load cells, or tactile arrays, with pristine alignment to visual frames.
- Strong video processing experience with OpenCV, FFmpeg, and Decord, including managing I/O bottlenecks for terabyte-scale video datasets.
- Solid working knowledge of 3D geometry and robotics data: coordinate frames and transforms, rotation representations, camera intrinsics/extrinsics, forward/inverse kinematics, and URDF.
- Proven ability to implement programmatic and generative data augmentation for computer vision and time-series data.
Tools and technologies you may use
Python, PyTorch, OpenCV, FFmpeg, Decord, URDF, Ray, Apache Spark, Open X-Embodiment, DROID, AgiBot World, EgoDex, SAM-family, MANO, SMPL, Omniverse, MuJoCo, NVIDIA robotic software stack, VLM.
Benefits
- Competitive compensation
- Performance-based bonus
- 99% employer covered medical benefits
- Early-stage equity
- Competitive PTO
- Company-wide paid winter break between December 24th and January 2nd
- Full access to advanced tools
Bonus skills
- Experience with NVIDIAβs robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).
- Comfort composing and evaluating perception modules as part of a pipeline, including segmentation (SAM-family), monocular depth, hand/body pose estimation (MANO/SMPL), and 6-DoF object pose tracking or point tracking.
- Familiarity with distributed data processing systems such as Ray and Apache Spark for cluster computing.
- Background using or generating synthetic robotic data via simulation (Omniverse, MuJoCo).
- Experience integrating spatial awareness or tactile data representations (for example, Fourier encoding) into visual pipelines.
Department: Software. Reports to: Teleoperations Lead. Employment type: Full-Time. Location: Houston, TX or Pensacola FL.