Senior Support Engineer
New
J
JobgetherFinancial Technology
Based in BrazilFull-TimeSenior
SalaryTotal compensation range of approximately USD 72K–90K, plus equity opportunities.
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 3–5 years
- Required Skills
- AWSDockerKubernetesCI/CDDatadogDistributed SystemsTechnical support
Requirements
- 3–5 years of experience in technical support, Site Reliability Engineering (SRE), Forward Deployed Engineering, or full-stack engineering roles.
- Strong experience troubleshooting distributed systems and production environments.
- Expertise with observability and monitoring platforms such as DataDog, ELK Stack (Elasticsearch, Logstash, Kibana), Prometheus, Grafana, or similar tools.
- Solid knowledge of infrastructure and DevOps practices, including Infrastructure as Code (IaC), containerization (Docker, Kubernetes), cloud-native architectures, and CI/CD pipelines.
- Strong experience with AWS infrastructure, database management, and system monitoring.
- Experience with incident management processes and high-intensity support environments.
- Ability to create automation solutions and improve operational workflows.
- Strong English communication skills, with the ability to collaborate effectively with global teams and external stakeholders.
- Proactive mindset with strong ownership, problem-solving ability, and experience working with SLA/SLO environments.
Responsibilities
- Lead complex technical troubleshooting and incident resolution from investigation through the implementation of permanent solutions.
- Own production issues across distributed systems, infrastructure, and application layers.
- Design and improve monitoring processes, including playbooks, runbooks, incident documentation, and knowledge base resources.
- Leverage AI-powered tools and automation workflows to improve issue detection, diagnosis, and resolution speed.
- Collaborate with engineering, customer experience, and client-facing teams to resolve technical challenges and improve service delivery.
- Develop and maintain automation scripts and tools to reduce manual work and increase system reliability.
- Establish best practices for incident response, monitoring, logging, and proactive issue prevention.
- Support the creation of scalable monitoring systems and operational processes.
- Analyze recurring technical problems and implement strategic improvements to prevent future incidents.
- Act as a communication bridge between technical teams and customers during critical incidents and implementation challenges.
View Full Description & ApplyYou'll be redirected to the employer's site