Site Reliability Engineer (SRE) / DevOps Engineer
New
G
GradionDigital Innovation, Tech
Ho Chi Minh City / Hanoi / Can Tho / Da Nang City, Global follow-the-sun modelFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Good English - daily communication with European stakeholders is a core requirement
- Experience
- 4+ years
- Required Skills
- AWSPythonBashGCPKubernetesPrometheusCI/CDNetworking
Requirements
- 4+ years in a DevOps / SRE / Platform Engineering role within an international team
- Solid Kubernetes knowledge - cluster operations, troubleshooting, and configuration
- Hands-on cloud experience with AWS and/or GCP
- Good understanding of networking fundamentals - DNS, load balancing, firewalls, VPC
- Scripting and automation skills (Python, Bash, or similar)
- Experience with CI/CD tools and GitOps-based delivery
- Working knowledge of monitoring and observability systems (Prometheus, ELK, or equivalent)
- Good English - daily communication with European stakeholders
- Self-directed and proactive mindset
Responsibilities
- Own platform availability: monitor, triage, and resolve incidents within defined SLA windows
- Manage cloud infrastructure on AWS and/or GCP - provisioning, scaling, and day-to-day operations
- Maintain and improve CI/CD pipelines and GitOps workflows
- Operate observability systems: monitoring, logging, and alerting at production scale
- Participate in on-call rotation as part of the global follow-the-sun coverage model
- Configure, deploy, and manage AI tooling and MCP servers in production environments
- Contribute to infrastructure automation, scripting, and internal tooling
- Write clear post-incident reviews and contribute to the monthly operational report
- Collaborate closely with engineering teams across multiple time zones
View Full Description & ApplyYou'll be redirected to the employer's site