Applied AI Engineer
Backend Developer
Agentic Ai
Ai Agent
Ai Agent Platform
Analytics
Artificial Intelligence
Big Data
Business Intelligence
Clickhouse
Data Analysis
Data Analytics
Data Engineer
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Visualization
Data Warehouse
Database
Databases
Digital Marketing
Generative AI
Programming
Programming Language
Programming Languages
Rag Architectures
Reporting and Analytics
SQL
Job Description
Block Labs is building governed, production-grade AI agents and the data and agent platform behind decisioning that directly affects business outcomes. As an Applied AI Engineer, you will connect LLM-powered intelligence to audited data, risk-aware tool execution, and operator-controlled surfaces, spanning agent workflows, decision models, and analytics dashboards. This onsite role is based in Malta, MT.
What you’ll do
- Deliver a production SQL BI analyst agent first: a Slack-native agent that answers business questions using governed SQL over the analytical warehouse, with validated queries, sanity-checked results, and cited evidence behind every number.
- Build and own analyst agents end to end, including executive P&L Q&A, daily health briefings, and root-cause analysis for metric movements, guarded by SQL and schema validation, result sanity checks, and cross-checks against canonical reporting views.
- Extend capabilities into customer-facing agents with intent triage and routing, retrieval-grounded (RAG) responses over versioned knowledge bases, and strict abort-and-escalate fallbacks; include multi-turn conversational state machines, localized brand voice, and escalation logic integrated with helpdesk and CRM platforms (webhook ingestion, session lifecycle management, intent metadata tagging, and automated escalation tickets with pre-packaged tool context).
- Create a risk-stratified tool layer between agents and back-office APIs using read-only context gathering first, information-first validation before actions, multi-turn confirmation workflows, and per-tool switching that escalates low-risk mutations toward autonomous execution as evidence accumulates.
- Harden agents against adversarial input using prompt injection screening, confidence-threshold freezes on sensitive intents, silent security escalation paths, and defenses against tool misuse and data exfiltration.
- Build an autonomy ladder using LangGraph, the Anthropic Agent SDK and Model Context Protocol (MCP), or equivalent orchestration frameworks.
- Engineer the closed feedback loop with corrections capture, proven-query and semantic memory, decision audit logging, and evaluation harnesses, including regression suites to ensure new tools or intents introduce zero degradation to existing paths.
- Build and productionise models behind decision signals, including churn, lifetime value, and bonus-sensitivity models; composite player risk scores across identity, payment, gameplay, bonus, and network signals; collusion, bot-play, and multi-accounting detection; and anomaly detection for treasury and payments.
- Ship models as governed signals with versioned, SLA’d contracts to the decision engine and agents, including freshness, drift, and calibration monitoring and automated retraining paths; deliver through real-time (Kafka/MSK), near-real-time, and batch (ClickHouse) tiers.
- Own multi-vector withdrawal risk scoring with cited rationale and confidence, evidence-aware aggregation, and automatic re-scoring when late evidence lands.
- Codify business rules with domain owners so policies remain auditable: translate policy into deterministic configurable rules, simulate and backtest threshold changes on historical data before activation, design holdouts and control groups to measure uplift, and run deep-dive analyses feeding both agents and the executive team.
- Build supervisor and approval surfaces including review queues with one-click action proposal cards for high-risk mutations, searchable session replay (prompts, model outputs, reasoning chains, and tool calls), and a structured grading module feeding evaluation and fine-tuning datasets.
- Design and ship dashboards the business uses: decision audit views, agent performance dashboards (correction rate, failure rate, and decision volume by rule and vector), risk review queues, and KPI views built with the BI team, including movement toward AI-assisted anomaly detection and explanation.
- Ensure humans set objectives, budgets, and approval gates while AI and ML models author, score, simulate, and optimise within them; the engine executes deterministically inside approved boundaries and logs everything. The goal is grounded, auditable actions for real customers and real money, gated by human approval where stakes demand it.
Requirements
- 4+ years of experience in software, data science, or machine learning engineering, including 1+ years building LLM-powered agents in production (tool use and function calling, structured outputs, retrieval and memory, multi-step orchestration with LangGraph or the Anthropic Agent SDK).
- Shipped a production RAG system and can discuss grounding, chunking and retrieval quality, hallucination control, and when to refuse to answer.
- Can treat customer-facing agents as an attack surface, including defenses against prompt injection, tool-call abuse, and data leakage through model outputs.
- Production ML lifecycle ownership across feature engineering, training, serving, monitoring, and retraining. Fraud, risk, or abuse detection experience is a strong signal (imbalanced classes, adversarial users, cost-asymmetric decisions).
- Statistical rigor in experiment design, holdouts and control groups, uplift measurement, and score calibration, with the ability to defend threshold choices.
- Evaluation discipline for non-deterministic systems, including building evaluation harnesses and regression suites to catch quality drift early.
- Strong Python for production services, comfort in TypeScript for review and approval surfaces, and strong SQL skills on columnar analytical databases (with ClickHouse preferred).
- Can take work to stakeholder-ready surfaces (review queues, approval interfaces, dashboards, lightweight internal apps) using tools such as Streamlit.
- Experience designing systems where model outputs feed deterministic execution inside a clean governed boundary, including LLM observability and tracing (such as Langfuse or LangSmith), and ownership of what ships and how it fails.
Technologies
- SQL, Python, TypeScript, ClickHouse, Kafka, MSK, Slack, LangGraph, Anthropic Agent SDK, Model Context Protocol (MCP), Langfuse, LangSmith, Streamlit, RAG
Nice to have
- Experience in iGaming or other high-trust, transaction-intensive environments with security, fraud prevention, auditability, traceability, data integrity, and robust operational controls.
- Helpdesk or CS-platform integration experience (Intercom, Zendesk, or similar) including webhooks, conversation APIs, agent-assist, or full automation.
- Exposure to blockchain or crypto-native transaction flows, including on-chain data, wallet clustering, or stablecoin settlement.
- Experience with constrained optimisation, bandits, or reinforcement learning under hard business constraints (budgets, caps, exclusion lists).
- Experience with rule engines or decision-management systems, and Slack app development.
- Event-driven and streaming experience, including Kafka or MSK consumers, idempotent processing, and failure handling.
How the team works
- Fully remote with asynchronous-first communication, with EU timezone overlap preferred.
- Small, high-autonomy Intelligence team within the Data function, reporting to the Head of Data, coordinating with AI, BI, and Infrastructure Teams, and for customer-facing agents with the Head of CS and product squads whose platforms the agents serve.
- Architecture decisions are documented and debated; you participate in design reviews and own domain decisions.
- Autonomy is earned by evidence: models and agents start propose-only and human-gated, then graduate step by step backed by simulation and outcome data.
- Multi-tenant scale from day one, so what you ship for one operator must absorb the next without additional engineering effort.