Senior Support Engineer
New
A
Alternative PaymentsFinTech
Brazil, RemoteFull-TimeSenior
Salary72k-90k USD, plus equity
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 3-5 years
- Required Skills
- AWSDockerKubernetesGrafanaPrometheusCI/CDDatadogDistributed SystemsTroubleshooting
Requirements
- 3-5 years of experience in technical support, SRE, Forward Deployed Engineer, or full-stack engineering.
- Strong expertise with observability and monitoring tools such as DataDog, ELK Stack (Elasticsearch, Logstash, Kibana), Prometheus, Grafana or similar platforms.
- Solid infrastructure and DevOps knowledge including infrastructure as a code (IaC), containerization (Docker/Kubernetes) and cloud-native architecture.
- Strong skills in AWS infrastructure, database management, and CI/CD pipelines.
- Solid experience in troubleshooting distributed systems and production environments.
- Strong communication skills in English to collaborate effectively across engineering teams, customer success, and external clients.
- Proactive mindset with the ability to solve complex technical problems, work with SLA/SLO, and drive automation initiatives.
- Experience with incident management processes and high-intensity support environments.
Responsibilities
- Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
- Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios.
- Draft detailed documentation including post-incident reviews and knowledge base articles.
- Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
- Collaborate with cross-functional teams to deliver solutions, optimize support processes, and implement scalable monitoring systems.
- Take ownership of complex production issues and provide technical expertise across distributed systems, infrastructure, and application layers.
- Develop and maintain automation scripts and tools to reduce manual intervention and improve system reliability.
- Bridge communication between engineering teams, customer experience, and clients during critical incidents and implementation challenges.
View Full Description & ApplyYou'll be redirected to the employer's site