Senior DevOps Engineer
New
K
KaratSaaS infrastructure
Remote (India - Bangalore ONLY); only available to candidates residing in Bengaluru, Schedule must overlap with U.S. business hours.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Clear written and verbal English communication skills
- Experience
- 5+ years of experience in DevOps, Site Reliability Engineering, infrastructure engineering, platform engineering, or a closely related discipline
- Required Skills
- AWSDockerCI/CDLinuxDatadogNetworking
Requirements
- Have 5+ years of experience in DevOps, Site Reliability Engineering, infrastructure engineering, platform engineering, or a closely related discipline.
- Have significant hands-on production experience with AWS.
- Have experience designing, operating, and improving CI/CD systems using CircleCI, GitHub Actions, Jenkins, GitLab CI, or another major platform; CircleCI is strongly preferred.
- Have strong experience with Docker and containerized application environments.
- Have strong Linux, networking, security, and cloud-infrastructure fundamentals.
- Have practical experience applying SRE principles to production systems, including observability, alerting, incident response, root-cause analysis, and capacity planning.
- Have hands-on experience with a leading telemetry and observability platform; Datadog is strongly preferred.
- Have experience designing monitoring and alerting systems that reduce noise and support incident response and service ownership.
- Have experience with cloud cost management and optimization.
- Have experience partnering with application-engineering teams on shared infrastructure and operational practices.
- Have clear written and verbal English communication skills.
Responsibilities
- Own and evolve Karat’s AWS SaaS infrastructure to keep services secure, scalable, reliable, observable, and cost-efficient.
- Design, improve, and operate CI/CD pipelines using CircleCI and related tooling.
- Build and maintain observability capabilities, including metrics, logs, traces, dashboards, and actionable alerting.
- Apply SRE principles to improve reliability, availability, performance, capacity planning, incident response, root-cause analysis, and operational learning.
- Partner with software engineering teams to establish shared infrastructure and operational practices and influence delivery decisions.
View Full Description & ApplyYou'll be redirected to the employer's site