DataJobs.io
← Back to all jobs

Job Description

In this role, you will define and own Engine Infra metrics and telemetry, build analytics pipelines and dashboards, and use experiments and causal analysis to improve game-server efficiency, reliability, and player experience.

Responsibilities

  • Own Engine Infra metrics, telemetry logging, and validation for the RCC fleet and edge compute.
  • Define and maintain KPIs for game-server efficiency, reliability, and player experience, including:
    • CPU and memory
    • Heartbeat
    • Occupancy and multiplexing efficiency
    • Crash and OOM rates
    • Runtime error codes
    • Latency quality (for example, DataPing vs RakPing)
  • Access and process raw RCC and engine telemetry.
  • Build analytics pipelines and deliver operating views by turning signals across per-instance, per-game, datacenter, and channel into dashboards and decision systems.
  • Support improvements that help Engine Infra pack more players onto fewer cores, including correct CPU and memory allocation and keeping multiplexed and large-server processes healthy.
  • Lead channel tests, hardware A/Bs, and causal analyses for:
    • Multiplexing
    • ARM64 bursting
    • Regional capacity
    • Edge latency
  • Quantify trade-offs among efficiency, stability, and quality.
  • Separate impacts from host, container, engine, content, and release effects.
  • Estimate player-experience impact, including latency-sensitive cohorts (not only global averages), to ensure launches ship on evidence.
  • Collaborate with Engine Infra, Game Engine, Networking, Matchmaking, Infrastructure, Product, and TPMs, plus teams such as Creator Analytics and Growth, to debug ambiguous production issues and identify capacity and quality opportunities.
  • Influence how Roblox runs live experiences on owned edge and cloud hardware:
    • A more efficient RCC fleet
    • Better regional capacity
    • Lower edge latency
    • Interactive performance as servers scale toward larger, denser configurations

Requirements

  • Advanced degree in Computer Science, Statistics, Engineering, or a related field, with a focus on large-scale system analysis.
  • 4+ years as a Data Scientist, with a proven record in high-volume telemetry and product quality optimization.
  • Expertise processing massive, noisy datasets (for example, crash logs or telemetry pings) and building scalable pipelines for reporting.
  • Experience using first-principles reasoning to debug complex, ambiguous system failures where root cause is not always clear.
  • Ability to present deep technical results to Engineering leaders and use data storytelling to prioritize work and future investment.

Nice to have

  • Direct experience with game engines or systems performance.
  • Experience in game development or working on real-time interactive systems.
  • Background or interest in memory management, OS-level interactions, or network stack optimization.

Compensation and location

  • Location: San Mateo, CA, United States (onsite).
  • Annual salary range: USD 221,380 - 263,670.
  • Actual base pay depends on factors including professional background, training, work experience, location, business needs, and market demand.
  • In some circumstances, actual salary could fall outside the expected range.
  • Full-time employees are also eligible for equity compensation and benefits as described on the company’s page.

Work schedule

  • Onsite Tuesday, Wednesday, and Thursday.
  • Optional presence on Monday and Friday (unless otherwise noted).

US based roles only

  • The company may not be able to employ candidates for this role who have US work authorization related to certain U.S. visa categories, or support future H-1B sponsorship at this time.

Similar Jobs