Senior Backend Engineer - Monitoring and Anomaly Detection

New
J
JobgetherMonetization Systems
IndiaFull-TimeSenior
SalaryCompetitive compensation package aligned with experience and market standards.
Apply NowOpens the employer's application page

Job Details

Languages
English
Required Skills
PythonRuby on RailsSalesforceClickhouseGrafanaPrometheus

Requirements

  • Professional experience developing applications with Ruby on Rails.
  • Strong background in backend engineering, site reliability engineering, or observability-focused roles.
  • Experience designing and operating monitoring, alerting, and reliability solutions using tools such as Prometheus, Grafana, and OpenTelemetry.
  • Experience building anomaly detection, monitoring, risk management, or data quality solutions.
  • Experience working with business-critical systems such as billing, financial, or transactional platforms.
  • Familiarity with data platforms and reporting systems, including technologies such as ClickHouse and event-driven pipelines.
  • Experience with change data capture, event streaming, or messaging systems is preferred.
  • Ability to independently own projects from concept through production delivery and ongoing optimization.
  • Strong understanding of software architecture, scalability, reliability, and operational best practices.
  • Experience using Python for data analysis, automation, or anomaly detection is an advantage.
  • Familiarity with platforms such as Salesforce or Zuora is preferred.
  • Strong written and verbal communication skills in English.
  • Ability to collaborate effectively in a remote, distributed, and asynchronous work environment.

Responsibilities

  • Design, build, and maintain backend solutions for monitoring, telemetry, alerting, and anomaly detection across critical systems.
  • Develop metrics, logs, and tracing capabilities using observability tools such as Prometheus, Grafana, and OpenTelemetry.
  • Build automated detection mechanisms for billing anomalies, data inconsistencies, event failures, and transactional issues.
  • Create reconciliation processes and data integrity checks across usage, billing, and customer-related pipelines.
  • Define and maintain service reliability practices, including SLOs, SLIs, operational dashboards, runbooks, and incident response processes.
  • Investigate and apply artificial intelligence and machine learning techniques to predict anomalies and improve issue resolution.
  • Own engineering initiatives from design and technical planning through implementation, deployment, and production monitoring.
  • Review code contributions and provide constructive feedback to improve engineering quality and maintainability.
  • Collaborate with product, finance, support, and engineering stakeholders to transform operational needs into scalable technical solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive compensation package aligned with experience and market standards.
Apply Now