We are looking for an exceptional Production Troubleshooting & Incident Lead to take end-to-end ownership of complex production incidents in a mission-critical banking environment.
This is not a traditional support role.
The person in this position will be expected to quickly understand complex problems, build a clear picture of the incident across multiple systems, identify the relevant components and dependencies, determine the root cause, engage the appropriate engineering teams, and drive the incident through to complete resolution.
Our production environment consists of multiple applications, services, infrastructure components, business processes, and integrations with the Bank's systems. Successful troubleshooting therefore requires both deep technical capabilities and strong system-level thinking.
The ideal candidate is a natural problem solver — someone who is comfortable entering an unclear situation, asking the right questions, connecting information from multiple sources, and rapidly turning uncertainty into a structured understanding of the problem.
עיקרי התפקיד
Troubleshooting & Root Cause Analysis
Build a clear technical picture of an incident and identify the root cause, even when the initial information is incomplete or ambiguous.
Follow the complete transaction and process flow across multiple systems.
Form hypotheses, validate them systematically, and eliminate possible causes.
The role requires the ability to understand not only what each component does, but also how the entire system behaves as one ecosystem.
Debugging & Investigation
Perform advanced troubleshooting using logs, monitoring tools, traces, metrics, databases, APIs, and other diagnostic information.
Navigate between different systems and data sources to reconstruct what happened.
Identify patterns, anomalies, timing issues, dependencies, and unexpected system behaviour.
Incident & Crisis Management
Identify the right technical stakeholders.
Separate facts from assumptions.
Communicate status, findings and risks to C-level management.
Assign clear actions to relevant stakeholders.
Verify that the production environment has fully recovered before closing the incident.
דרישות
Exceptional hands-on troubleshooting and problem-solving skills with at least +5 years of experience in this field.
Strong holistic thinking and ability to understand complex multidisciplinary systems.
Ability to quickly learn unfamiliar systems and connect information from multiple technical and business sources.
Ability to ask the right questions, challenge assumptions and distinguish facts from hypotheses.
Strong ownership and a drive-to-resolution mindset.
Ability to act well under intensive and high-pressure environment.
Ability to move comfortably between business processes and deep technical investigation.
Technical Background
Production troubleshooting in complex, distributed environments.
Microservices and distributed architectures.
Cloud and On-Prem environments.
Kubernetes.
APIs and system integrations.
SQL and database investigation.
Application and infrastructure logs.
Monitoring, observability, metrics and tracing.
CI/CD and modern DevOps environments.
The candidate does not need to be the developer or infrastructure expert for every component, but must have the technical depth and curiosity to understand how the pieces work together and where to investigate when something goes wrong.
Interpersonal & Leadership Skills
Exceptional interpersonal skills are a critical requirement for this role.
The candidate must be able to:
Communicate clearly and accurately.
Build excellent working relationships with engineers, managers, business stakeholders and Bank teams.
Know when to push, when to listen, and how to keep everyone focused on solving the problem.
The goal is - Bring clarity to complex problems and drive them to resolution.
יתרון
Experience in banking, financial systems or other mission-critical environments is a strong