Site Reliability Engineer
New
Y
YunoPayments infrastructure
You can work from everywhereFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English — advanced proficiency, written and spoken; Spanish proficiency is nice to have.
- Experience
- +7 Years of Experience
- Required Skills
- AWSDockerPostgreSQLKafkaKubernetesTerraformDatadog
Requirements
- Have more than 7 years of experience.
- Design and own event-driven systems using message queues such as Kafka, NATS, or RabbitMQ; understand at-least-once delivery, consumer groups, dead letters, and backpressure.
- Have experience migrating systems from synchronous to asynchronous communication.
- Bring deep AWS experience with EC2, VPC, IAM, S3, and RDS, plus strong networking fundamentals.
- Use Terraform or Pulumi for infrastructure as code.
- Have production experience with Kubernetes and Docker, including container lifecycle, resource limits, health checks, and orchestration at scale.
- Have observability experience with Datadog or an equivalent, including dashboards, monitors, APM, and distributed tracing; have defined and operated SLOs, SLIs, and error budgets.
- Have hands-on experience with chaos engineering and resilience testing, such as fault injection, game days, or chaos experiments.
- Be able to debug distributed systems and production failures, and code automation and tooling in Go, Python, or a similar language.
- Have solid SQL experience with PostgreSQL and NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning.
- Have proven technical leadership setting reliability standards, influencing architecture across teams, and mentoring engineers.
- Have advanced written and spoken English proficiency.
- Preferred: production AI/MLOps infrastructure experience, multi-tenant container platforms, data pipelines and orchestration, incident-management tooling, or payments-industry experience.
Responsibilities
- Define reliability standards, SLO culture, error-budget policy, and incident practices across engineering teams.
- Drive platform architecture decisions and evolve the infrastructure as the platform matures.
- Design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable async messaging.
- Own AWS cloud infrastructure and automate provisioning with infrastructure as code.
- Build monitoring, tracing, and alerting for platform health.
- Act as the senior escalation point for difficult production problems.
- Run blameless postmortems and root-cause analyses, turning findings into permanent fixes.
- Mentor senior and mid-level engineers and raise reliability standards.
- Run fault-injection and resilience experiments to identify and prevent failures.
View Full Description & ApplyYou'll be redirected to the employer's site