Site Reliability Engineer

New
Y
YunoPayments infrastructure
You can work from everywhereFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English — advanced proficiency, written and spoken; Spanish proficiency is nice to have.
Experience
+7 Years of Experience
Required Skills
AWSDockerPostgreSQLKafkaKubernetesTerraformDatadog

Requirements

  • Have more than 7 years of experience.
  • Design and own event-driven systems using message queues such as Kafka, NATS, or RabbitMQ; understand at-least-once delivery, consumer groups, dead letters, and backpressure.
  • Have experience migrating systems from synchronous to asynchronous communication.
  • Bring deep AWS experience with EC2, VPC, IAM, S3, and RDS, plus strong networking fundamentals.
  • Use Terraform or Pulumi for infrastructure as code.
  • Have production experience with Kubernetes and Docker, including container lifecycle, resource limits, health checks, and orchestration at scale.
  • Have observability experience with Datadog or an equivalent, including dashboards, monitors, APM, and distributed tracing; have defined and operated SLOs, SLIs, and error budgets.
  • Have hands-on experience with chaos engineering and resilience testing, such as fault injection, game days, or chaos experiments.
  • Be able to debug distributed systems and production failures, and code automation and tooling in Go, Python, or a similar language.
  • Have solid SQL experience with PostgreSQL and NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning.
  • Have proven technical leadership setting reliability standards, influencing architecture across teams, and mentoring engineers.
  • Have advanced written and spoken English proficiency.
  • Preferred: production AI/MLOps infrastructure experience, multi-tenant container platforms, data pipelines and orchestration, incident-management tooling, or payments-industry experience.

Responsibilities

  • Define reliability standards, SLO culture, error-budget policy, and incident practices across engineering teams.
  • Drive platform architecture decisions and evolve the infrastructure as the platform matures.
  • Design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable async messaging.
  • Own AWS cloud infrastructure and automate provisioning with infrastructure as code.
  • Build monitoring, tracing, and alerting for platform health.
  • Act as the senior escalation point for difficult production problems.
  • Run blameless postmortems and root-cause analyses, turning findings into permanent fixes.
  • Mentor senior and mid-level engineers and raise reliability standards.
  • Run fault-injection and resilience experiments to identify and prevent failures.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now