Senior Production Reliability Engineer
We are seeking a highly technical Senior Production Reliability Engineer to lead our local Excellence in Engineering team. This is not a software development role: its focus is the reliability, diagnosability and operational effectiveness of complex, largely non-cloud production systems . The successful candidate will combine deep knowledge of system architecture, data flows and databases with advanced debugging, monitoring and troubleshooting skills. They will lead difficult investigations through to root cause, drive AI-enabled automation and operational improvement, and provide technical direction and mentoring to the local team. Infrastructure, capacity and performance are important supporting areas, but cloud-platform engineering is not the primary remit.
Advanced Incident Investigation & Problem Management
Diagnose issues across applications, services, operating systems, databases, messaging and networks using logs, traces, metrics, configuration, queries and controlled reproduction.
AI, Automation & Operational Improvement
Automate diagnostics, health checks, evidence collection, recovery activities and recurring operational workflows using scripts, tools and governed AI-assisted solutions.
Build searchable operational knowledge from incidents, runbooks, architecture information and known errors, and measure reductions in manual effort and time to recovery.
System Architecture, Data Flows & Databases
Trace transactions and data across APIs, services, message brokers, batch processes and databases to isolate defects and performance bottlenecks.
Diagnose database query, locking, contention, indexing, connectivity, data-quality and throughput issues, and identify resilience gaps and operational risks.
Local Team Leadership & Operational Readiness
Coach and mentor engineers in architecture, debugging, observability, automation and effective incident practice, while fostering collaboration with global teams.
Observability , Monitoring & Reliability
Help direct improvements to dashboards, alerts and diagnostic views covering service health, dependencies, latency, throughput, saturation and errors.
Skills, Experience & Qualifications Required:
4–7 years of progressively responsible experience in site reliability, production engineering, application support, systems engineering or a similarly deep technical operations role, ideally in financial services.
Strong understanding of system architecture, APIs, microservices, messaging, data flows and multi-tier failure modes.
Strong relational database and SQL skills, including query performance, locking, indexing, connectivity and data integrity.
Hands-on experience with observability platforms such as Dynatrace and with logs, metrics, traces, dashboards and alerts.
Scripting capability in PowerShell, Python, Bash or SQL; working knowledge of Java, C#, Kubernetes, Kafka and API gateways is advantageous, but regular feature development is not expected.
Understanding of performance, capacity, resilience, networking and operating-system fundamentals; on-premises or hybrid experience is strongly valued, while Azure familiarity is useful but secondary.
Experience providing technical leadership, prioritising team activity, mentoring engineers and coordinating across engineering, product, infrastructure and customer-facing teams.
Across the globe, institutional investors rely on us to help them manage risk, respond to challenges, and drive performance and profitability. We keep our clients at the heart of everything we do, and smart, engaged employees are essential to our continued success.
We are committed to fostering an environment where every employee feels valued and empowered to reach their full potential. As an essential partner in our shared success, you’ll benefit from inclusive development opportunities, flexible work-life support, paid volunteer days, and vibrant employee networks that keep you connected to what matters most. Join us in shaping the future.
As an Equal Opportunity Employer, we consider all qualified applicants for all positions without regard to race, creed, color, religion, national origin, ancestry, ethnicity, age, disability, genetic information, sex, sexual orientation, gender identity or expression, citizenship, marital status, domestic partnership or civil union status, familial status, military and veteran status, and other characteristics protected by applicable law.
Discover more information on jobs at StateStreet.com/careers
Read our CEO Statement