Site Reliability Engineer (SRE) / DevOps Engineer

New
G
GradionDigital Innovation, Tech
Ho Chi Minh City / Hanoi / Can Tho / Da Nang City, Global follow-the-sun modelFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Good English - daily communication with European stakeholders is a core requirement
Experience
4+ years
Required Skills
AWSPythonBashGCPKubernetesPrometheusCI/CDNetworking

Requirements

  • 4+ years in a DevOps / SRE / Platform Engineering role within an international team
  • Solid Kubernetes knowledge - cluster operations, troubleshooting, and configuration
  • Hands-on cloud experience with AWS and/or GCP
  • Good understanding of networking fundamentals - DNS, load balancing, firewalls, VPC
  • Scripting and automation skills (Python, Bash, or similar)
  • Experience with CI/CD tools and GitOps-based delivery
  • Working knowledge of monitoring and observability systems (Prometheus, ELK, or equivalent)
  • Good English - daily communication with European stakeholders
  • Self-directed and proactive mindset

Responsibilities

  • Own platform availability: monitor, triage, and resolve incidents within defined SLA windows
  • Manage cloud infrastructure on AWS and/or GCP - provisioning, scaling, and day-to-day operations
  • Maintain and improve CI/CD pipelines and GitOps workflows
  • Operate observability systems: monitoring, logging, and alerting at production scale
  • Participate in on-call rotation as part of the global follow-the-sun coverage model
  • Configure, deploy, and manage AI tooling and MCP servers in production environments
  • Contribute to infrastructure automation, scripting, and internal tooling
  • Write clear post-incident reviews and contribute to the monthly operational report
  • Collaborate closely with engineering teams across multiple time zones
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now