Job Description
The Mission: The enterprise software landscape is fracturing. We are transitioning from "Systems of Record" to "Systems of Action." Zendesk is leading this shift with the Resolution Platform, aiming to be the first CX company to secure $1B in AI-driven revenue. To achieve this, we must solve the "Orchestration Trap." We are not just building features; we are building the high-performance runtime that allows fleets of AI Agents to Perceive, Reason, and Act.
The Opportunity: You will operate at the cutting edge of large language models, planning algorithms, and multi-agent coordination. By designing advanced memory structures and autonomous learning mechanisms, you will bridge the gap between offline model capabilities and dynamic, real-world task execution.
What You Will Architect & Research
Building the World's Best Task Agents: You are tasked with achieving state-of-the-art performance against the industry's most rigorous benchmarks. You will optimize our agentic workflows to push past the 2026 frontiers, targeting elite-level autonomous problem solving on complex multi-step tool-use benchmarks.
Advanced Memory & Cognitive Architectures: You will design sophisticated memory systems inspired by human cognition, allowing agents to dynamically filter interference, maintain context, and leverage long-term historical knowledge effectively across extended interactions.
Trajectory Analysis & Reasoning Refinement: You will analyze complex agentic AI trajectories and trace patterns to understand how models navigate non-deterministic, multi-step problems. By leveraging these insights, you will build systems that capture the agent's entire chain-of-thought, analyzing conditional branches to actively detect anomalies and halt hallucination loops prior to failure.
Self-Improving Loops & Skill Discovery: You will architect scalable, autonomous self-improving loops that allow agents to operate and learn continuously without human intervention. You will design frameworks where tool search patterns, error handling, and sophisticated retry logics are actively fed back into the system to dynamically improve the structure of the agentic AI planner. This includes enabling agents to autonomously discover, synthesize, and incrementally acquire new reusable skills based on environmental feedback and task completion.
Enterprise Guardrails & Content Safety: You will engineer multi-layered defenses to secure agentic workflows against unique risks such as tool misuse, cascading action chains, and unintended control amplification. This includes designing strict input validation to block malicious prompt injections or jailbreak attempts, as well as output filtering to ensure responses remain within the application's domain boundary.
Supervisor Patterns & Governance: You will implement governance-centric architectures, such as supervisor or manager agent patterns, to explicitly regulate actions and decision sequences during runtime. You will enforce capabilities-based access, ensuring that an agent's available tools are strictly determined by the user's role, and align system evaluations with compliance standards like the NIST AI Risk Management Framework.
Rigorous Agentic Evaluation: You will move beyond static leaderboards to build continuous, multi-turn evaluation frameworks. You will design preference data pipelines to rigorously test emergent multi-agent coordination, ensuring our agents act safely and align perfectly with enterprise intent.