Staff AI Engineer
Job Description
Pulley is hiring a Staff-level AI Engineer to build AI intelligence for permitting workflows. In this hybrid role based in the San Francisco Bay Area, you will set the technical direction for deploying LLM capabilities across product surfaces, owning work from early ambiguity through architecture, production quality, and evaluation.
Role Responsibilities
- Define the technical direction for how Pulley applies LLMs across multiple product surfaces, owning the AI problem space from ambiguity through architecture to shipped product and ongoing iteration.
- Transform unstructured permitting documents, city regulations, and jurisdiction workflows into structured and reliable outputs, including extraction, classification, retrieval, and agentic workflows over document sources that were not built for machine readability.
- Set company-wide standards for evaluation and observability by defining ground truth, measuring quality and regressions, and establishing when a model change represents real improvement, supported by systems that make this the default across teams shipping LLM features.
- Use AI agents as a daily practice, directing, reviewing, and shipping agent-driven work at high velocity while maintaining accountability for quality.
- Make technical bets that shape what Pulley can build next year, including choices around model selection, architectures, and build versus buy decisions, with ownership of production consequences.
- Multiply engineering impact by setting repeatable patterns for LLM feature development, mentoring senior engineers toward broader scope, and improving team speed through systems, standards, and abstractions.
Required Qualifications
- 8+ years of software engineering experience, with a substantial portion focused on production LLM or ML systems.
- Demonstrated end-to-end ownership of a meaningful AI domain, covering problem identification, architecture, delivery, and ongoing production ownership, including data quality, evaluation design, cost and latency, and failure handling.
- Hands-on production experience with large language models, including prompting, retrieval-augmented generation, structured extraction, tool use, and agentic workflows, with the judgment to use the right approach for the task.
- Experience designing evaluations and implementing practices that make LLM-powered features reliable in production.
- Experience building with AI coding agents in real deployments, where agents performed substantive implementation under your direction.
- Ability to architect durable systems while making pragmatic tradeoffs.
- Experience mentoring engineers and providing technical direction that others operationalize.
- Based in the San Francisco Bay Area and able to work in person 4 days a week.
Technology Stack
- LLMs, ML
- Retrieval-augmented generation
- TypeScript
- React
- Google Cloud
Location and Employment Details
- Location: San Francisco, CA (hybrid)
- Employment type: Full time
- Department: Engineering
- Department: Engineering
Compensation and Benefits
- Salary: USD 300,000 - 350,000 per year
- Equity: Offers Equity
What You Bring
- Comfort operating in ambiguity and preference for defining the right problem over executing a prewritten spec.
- Product mindset with interest in validating outcomes with users and ensuring solutions address customer needs.
- Rigorous approach to defining and measuring success, relying on evaluations rather than demos and building measurement before delivery.
- Commitment to both quality and velocity, using tools, abstractions, and process improvements to achieve both.
- Strong ownership instincts at organizational scale, including identifying and fixing systems or process gaps without needing permission.
Nice to Have
- Experience with document understanding at scale, including OCR, layout-aware parsing, or vision-language models over scanned PDFs, drawings, or forms.
- Experience fine-tuning models or building data pipelines to produce training and evaluation datasets from real-world usage.
- Experience in construction tech, govtech, proptech, or similar domains where real-world documents and processes are inherently messy.
- Experience with modern full-stack development, including TypeScript and React, plus willingness to work in application code that delivers AI features to users.
- Experience operating as the most senior AI engineer in a domain, including being the escalation point when answers are unclear.
- Startup experience, particularly at a stage where you helped build the team in addition to shipping the product.