Senior Backend Engineer - Monitoring and Anomaly Detection
New
J
JobgetherMonetization Systems
IndiaFull-TimeSenior
SalaryCompetitive compensation package aligned with experience and market standards.
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Required Skills
- PythonRuby on RailsSalesforceClickhouseGrafanaPrometheus
Requirements
- Professional experience developing applications with Ruby on Rails.
- Strong background in backend engineering, site reliability engineering, or observability-focused roles.
- Experience designing and operating monitoring, alerting, and reliability solutions using tools such as Prometheus, Grafana, and OpenTelemetry.
- Experience building anomaly detection, monitoring, risk management, or data quality solutions.
- Experience working with business-critical systems such as billing, financial, or transactional platforms.
- Familiarity with data platforms and reporting systems, including technologies such as ClickHouse and event-driven pipelines.
- Experience with change data capture, event streaming, or messaging systems is preferred.
- Ability to independently own projects from concept through production delivery and ongoing optimization.
- Strong understanding of software architecture, scalability, reliability, and operational best practices.
- Experience using Python for data analysis, automation, or anomaly detection is an advantage.
- Familiarity with platforms such as Salesforce or Zuora is preferred.
- Strong written and verbal communication skills in English.
- Ability to collaborate effectively in a remote, distributed, and asynchronous work environment.
Responsibilities
- Design, build, and maintain backend solutions for monitoring, telemetry, alerting, and anomaly detection across critical systems.
- Develop metrics, logs, and tracing capabilities using observability tools such as Prometheus, Grafana, and OpenTelemetry.
- Build automated detection mechanisms for billing anomalies, data inconsistencies, event failures, and transactional issues.
- Create reconciliation processes and data integrity checks across usage, billing, and customer-related pipelines.
- Define and maintain service reliability practices, including SLOs, SLIs, operational dashboards, runbooks, and incident response processes.
- Investigate and apply artificial intelligence and machine learning techniques to predict anomalies and improve issue resolution.
- Own engineering initiatives from design and technical planning through implementation, deployment, and production monitoring.
- Review code contributions and provide constructive feedback to improve engineering quality and maintainability.
- Collaborate with product, finance, support, and engineering stakeholders to transform operational needs into scalable technical solutions.
View Full Description & ApplyYou'll be redirected to the employer's site