Staff Software Engineer, Alerting Platform
New
J
JobgetherSecurity & IT
IndiaFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonCloud ComputingMachine LearningApache KafkaGoLLMDistributed Systems
Requirements
- Demonstrated experience designing, building, or operating high-reliability alerting, detection, or notification systems in production at significant scale.
- Strong experience implementing and operating both rule-based and ML-driven detection or event-processing logic against real-world production traffic.
- Advanced experience with Apache Kafka, particularly producing and consuming event streams for reliable downstream processing and delivery.
- Strong understanding of telemetry and protocol data, including security tooling output, monitoring telemetry, syslog, OpenTelemetry, NetFlow or sFlow, SNMP, ICMP, and firewall logs.
- Ability to transform diverse technical data sources into meaningful network and system topologies, dependency maps, and operational context.
- Proven experience building systems where webhook reliability, idempotency, delivery guarantees, and performance under sustained load are critical.
- Expert-level proficiency in Go and/or Python, with strong experience across container orchestration and multi-cloud environments.
- Demonstrated ability to make organization-level architectural and technology decisions rather than focusing solely on individual implementation tasks.
- Strong systems-thinking, problem-solving, and technical communication skills, with the ability to influence teams and stakeholders across an organization.
- Experience mentoring engineers and providing technical direction across multiple teams or workstreams.
Responsibilities
- Define the architecture for high-fidelity alerting across rule-based thresholds, machine-learning anomaly scores, correlation logic, and customer-defined alert rules.
- Own the reliability of the alerting pipeline end to end, from alert evaluation through Kafka-based delivery and downstream notification systems.
- Establish robust guarantees around webhook delivery, idempotency, reliability, and sustained-load performance.
- Drive correlation strategies that transform security tooling, monitoring telemetry, syslog, OpenTelemetry, and network protocol data into meaningful network and system topologies and dependency maps.
- Make topology and dependency context usable for AI and LLM-driven reasoning across investigations, troubleshooting, correctness analysis, and remediation workflows.
- Establish the technical foundations for generating alert definitions from normalized data models using LLM-powered capabilities.
- Provide architectural leadership across multiple teams and workstreams, stepping into complex initiatives where deep technical expertise is required.
- Balance strategic investments in detection and correlation capabilities with the long-term scalability, reliability, and maintainability of the alerting platform.
- Make organization-level architecture and technology decisions and communicate technical direction effectively to engineering and business stakeholders.
- Mentor engineers across the teams you support, promote strong engineering practices, and raise the technical quality of the overall platform.
View Full Description & ApplyYou'll be redirected to the employer's site