DataJobs.io
← Back to all jobs

Job Description

The Data & AI Engineer at The Carlyle Group is a hands-on individual contributor who translates data and AI architecture into production systems. The role focuses on building AI-ready data pipelines, retrieval and semantic components, and data products that support analytics and generative AI applications across the organization.

Key Responsibilities

  • Build and operate AI-ready data pipelines for embedding generation, chunking, indexing, and refresh workflows to support reliable enterprise data retrieval by LLMs, agents, and generative AI applications.
  • Implement retrieval-augmented generation (RAG) capabilities, including vector store integrations, hybrid search, re-ranking, and grounding logic, aligned to architectural patterns defined by the Senior AI & Data Architect.
  • Develop tool and function interfaces enabling agents and copilots to query and act on enterprise data with guardrails, logging, and evaluation hooks.
  • Partner with Data Science and AI Engineering teams to operationalize feature stores, evaluation datasets, and reusable AI data products.
  • Contribute to semantic and context engineering for natural-language analytics, conversational reporting, and business-focused AI-driven insights.
  • Design, build, and maintain production-grade ELT, streaming, and transformation pipelines using dbt, Fivetran, and Snowflake.
  • Implement ingestion, modeling, and consumption practices that meet enterprise expectations for scalability, performance, security, resiliency, and cost efficiency.
  • Write clean, well-tested Python and SQL and apply engineering best practices, including version control, code review, CI/CD, modular design, and automated testing.
  • Support production onboarding of new sources and domains under a federated operating model, working with domain data engineers to apply shared platform capabilities.
  • Implement semantic models, data contracts, and analytical or dimensional models that enable trusted self-service analytics and dependable AI grounding.
  • Build and maintain reusable data products with clear ownership, documented contracts, and contextual metadata for human and AI consumers.
  • Collaborate with the Senior AI & Data Architect to refine and extend enterprise semantic standards based on production learnings.
  • Support discovery and consumption tooling so analysts, applications, and agents can access data products with minimal friction.
  • Implement data quality checks, lineage capture, and pipeline observability across both data and AI workloads.
  • Build logging, evaluation, and monitoring for AI systems, including prompt and response capture, retrieval metrics, and model performance signals in line with governance standards.
  • Partner with Data Governance to operationalize metadata, stewardship, and access controls so AI systems consume enterprise data with the same rigor as human users.
  • Surface issues early, propose remediation steps, and feed lessons learned back into architectural patterns.
  • Participate in architectural design reviews and contribute hands-on engineering perspective to evolving patterns and standards.
  • Mentor junior data engineers and analysts on modern data and AI engineering practices.
  • Document patterns, write runbooks, and share knowledge to accelerate adoption of reusable platform capabilities.

Required Qualifications

  • Bachelor’s degree.
  • 6+ years of overall relevant technical experience.
  • Experience in data engineering, analytics engineering, or platform engineering, including 1-2 years of direct, hands-on experience building generative AI or AI/ML systems in production.
  • Proven experience implementing retrieval, grounding, and semantic components for LLM- or agent-based applications, including RAG pipelines, vector stores, embedding workflows, and structured tool use.
  • Hands-on experience with modern AI platforms and tooling categories, such as AWS Bedrock, Databricks ML, Snowflake Cortex, OpenAI/Anthropic APIs, LangChain/LlamaIndex (or equivalents), MLflow, and vector databases like Databricks Vector Search, pgvector, or Pinecone.
  • Strong, demonstrable expertise in Python and SQL, including working knowledge of distributed processing frameworks such as Spark.
  • Deep, hands-on experience with modern data stacks (dbt, Fivetran, Snowflake) in AWS-based environments.
  • Track record of building data pipelines and products consumed by AI systems, beyond BI tools and human analysts.
  • Experience operating within federated data operating models and complex, regulated enterprise environments, with financial services experience preferred.

Technologies

Python, SQL, dbt, Fivetran, Snowflake, ELT, streaming, Snowflake Cortex, AWS Bedrock, Databricks ML, OpenAI/Anthropic APIs, LangChain/LlamaIndex, MLflow, Databricks Vector Search, pgvector, Pinecone, Spark, RAG, vector stores, hybrid search, re-ranking, embedding workflows, structured tool use, feature stores, vector databases, prompt and response capture, lineage capture, CI/CD.

Location and Work Schedule

Location: Washington, D.C. or New York, NY (onsite)

In-office requirement: 4 days per week

Compensation and Benefits

  • Anticipated base salary range: $160,000 to $180,000 (USD, per year).
  • Compensation range is specific to the applicable office location and considers factors including required and preferred skill sets, prior experience and training, and licenses and/or certifications.
  • Comprehensive benefits package including retirement benefits, health insurance, life and disability insurance, paid time off, paid holidays, family planning benefits, and wellness programs.
  • Eligible for an annual discretionary incentive program based on individual and organizational performance.

What Success Looks Like

Within the first 12 months, the role will deliver foundational AI-ready data pipelines and retrieval components based on the target-state architecture, productionize one or more priority RAG or agent-grounding use cases, and establish reusable engineering patterns that other domain teams can adopt across the federated data platform.

Education

Bachelor’s degree.

Similar Jobs