Senior Site Reliability Engineer
New
F
FleetioFleet management software
Remote - USA; open to candidates in the United States.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of Ruby/Rails Experience; 3+ years of AWS Experience
- Required Skills
- AWSKubernetesRuby on RailsTerraformDatadog
Requirements
- Have 5+ years of Ruby/Rails experience.
- Have 3+ years of AWS experience.
- Have Kubernetes experience.
- Have experience profiling and benchmarking source code.
- Be effective at code review and identifying potential performance problems before production.
- Have experience with Datadog or other APM tools.
- Bring a strong background in site reliability and infrastructure engineering for Rails applications.
- Have experience building AI agents, LLM-powered automations, or integrations for engineering or operations workflows (plus).
- Have experience with Infrastructure as Code tools such as Terraform (plus).
- Have a deep understanding of cloud network fundamentals, or experience with distributed event and data stores such as Kafka, Redis, Elasticsearch, Memcached, and TimescaleDB (plus).
Responsibilities
- Identify, triage, and resolve performance issues.
- Monitor performance metrics across Ruby, Rails, and database systems, including SLOs and SLIs.
- Build and maintain AI agents, skills, and automations to reduce operational toil.
- Use AI-assisted tooling for performance analysis, root-cause investigation, and code review.
- Collaborate with SREs to identify and address performance bottlenecks.
- Help product engineers adopt AI-driven performance and reliability workflows.
- Lead database capacity planning and upgrade initiatives.
- Manage database disaster recovery planning and execution, and oversee backup systems and pre-production databases.
- Maintain infrastructure and operations documentation, including runbooks, and participate in the on-call rotation.
View Full Description & ApplyYou'll be redirected to the employer's site