Site Reliability Engineer
New
C
CXMFintech / Trading
Remote (Americas, LatAm preferred), Americas time zones (UTC-3 to UTC-8)Full-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3–5 years
- Required Skills
- AWSPostgreSQLC#Grafana.NETPrometheusTerraform
Requirements
- 3–5 years of professional experience in Site Reliability Engineering or production operations.
- Strong experience debugging and supporting .NET/C# applications in production.
- Hands-on experience with Windows Server environments.
- Strong PowerShell scripting skills.
- Proficiency with Python or Bash.
- Experience with Grafana, Prometheus, and Loki for monitoring and observability.
- Solid understanding of metrics, logging, tracing, and alerting best practices.
- Experience with modern CI/CD pipelines, deployment strategies, and rollback mechanisms.
- Experience working with AWS.
- Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
- Experience troubleshooting Aurora PostgreSQL or other relational database platforms.
- Practical experience with SLIs, SLOs, Error Budgets, and Incident Response.
Responsibilities
- Participate in on-call rotation for production trading systems and lead incident response.
- Investigate production incidents, perform root cause analysis, and implement preventive actions.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views.
- Instrument .NET services to improve telemetry, metrics, logging, and visibility.
- Define, implement, and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL, and AWS infrastructure.
- Improve deployment safety, release automation, and rollback strategies.
- Automate operational tasks through scripting and infrastructure automation.
- Maintain runbooks and operational documentation.
View Full Description & ApplyYou'll be redirected to the employer's site