Cloud Systems Engineer - Site Reliability
New
T
TherapyNotes LLCHealthcare Software
Philadelphia, Pennsylvania, United StatesFull-TimeSenior
Salary$110,000-$150,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonBashCloud ComputingKubernetesAzureLinuxTerraformAnsibleDatadog
Requirements
- BS degree in Information Systems, Engineering, or equivalent experience.
- 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE.
- Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies.
- Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems.
- Expertise with an observability platform (Datadog, Prometheus, Grafana, or New Relic).
- Experience with scripting and operational automation using Bash, PowerShell, or Python.
- Proficiency with infrastructure-as-code and configuration-management practices.
- Experience participating in production on-call rotations, incident response, and root cause analysis.
- Experience working in Agile/DevOps environments and operating production services using ITSM practices.
- Prior software development experience or experience investigating application behavior via code/logs is a plus.
Responsibilities
- Own and continuously improve reliability metrics, logs, traces, and service-level views using Datadog.
- Design and maintain high-availability, high-throughput, data- and compute-intensive systems for a 24/7 SaaS platform.
- Partner with service owners to define SLIs, SLOs, and error budgets to drive operational readiness.
- Participate in and drive incident management as a technical responder or incident commander.
- Investigate issues across infrastructure and application layers using metrics, logs, and distributed traces.
- Improve deployment safety and service resilience through automated validation and recovery capabilities.
- Eliminate operational toil using automation tools like Bash, PowerShell, Python, or Ansible.
- Manage infrastructure using Terraform/OpenTofu and configuration automation with Ansible.
View Full Description & ApplyYou'll be redirected to the employer's site