Description:
We're looking for a Site Reliability Engineer to keep the trading platforms we depend on healthy, performant and continuously improving. You'll operate within a set of platforms that span Windows and Linux: a distributed platform built on Service Fabric, real-time news infrastructure, and a large Python production environment, running across development, integration, and production in multiple regions. When something breaks you'll dive in, diagnose quickly, and fix it - then make the change that stops it happening again.
This is a hands-on engineering role that sits at the boundary between development and infrastructure. The team is evolving into reliability engineering: automation-first, monitoring-as-code, self-healing systems, and shared tooling built to production standard. You'll work in a tight-knit, business-facing team alongside developers, production and infrastructure engineers in Dublin, and partner closely with global platform-engineering teams - consuming their products and contributing to them upstream rather than building regional one-offs.
In This Role You Will
OPERATE Help run the production estate day to day: deployments, platform operations, capacity, certificates, and health checks across environments and tenants. Triage and resolve incidents quickly, perform root-cause analysis, and drive the permanent fix so the same problem doesn't come back.
AUTOMATE Take an automation-first approach to everything the team touches: build auto-remediation for known failure modes, replace manual runbooks with code, and treat recurring toil as a defect to be engineered away.
OBSERVE Operate the monitoring and alerting stack (Checkmk, Elasticsearch/Kibana, Grafana, Prometheus): move its configuration into code, tune signal from noise, and make sure every alert that fires is actionable.
RELEASE Own safe delivery to production: design and maintain CI/CD pipelines in GitLab and deployment orchestration in Octopus Deploy.
ENGINEER Build and productionise the team's tooling with CI, tests, and documentation as standard, and contribute improvements upstream to the global platform products the estate runs on rather than maintaining local one-offs.
COLLABORATE Act as the bridge between development teams and the infrastructure teams (compute, network, storage, data-centre operations), applying your knowledge of both worlds: translate what developers need into infrastructure terms and infrastructure constraints back into design decisions, and build tooling that can be handed off to infrastructure teams to run.
What We're Looking For
| Organization | Susquehanna International Group |
| Industry | Engineering |
| Occupational Category | Site Reliability Engineer |
| Job Location | Dublin,Ireland |
| Shift Type | Morning |
| Job Type | Full Time |
| Gender | No Preference |
| Career Level | Intermediate |
| Experience | 2 Years |
| Posted at | 2026-08-03 6:38 pm |
| Expires on | 2026-09-17 |