Staff Machine Learning Engineer
New
D
DragosCybersecurity
United StatesFull-TimeStaff
Salary$190,000.00
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years of engineering experience with at least 4 years focused on machine learning
- Required Skills
- DockerPythonSQLJavaKubernetesPyTorchGoRustTensorflowMLOps
Requirements
- 6+ years of engineering experience with at least 4 years focused on machine learning implementations in production environments.
- Strong software engineering foundation with expertise in Python and SQL.
- Experience with at least one additional language (Go, Rust, Java, or JVM-family languages).
- Demonstrated experience building and deploying ML systems using modern frameworks and libraries (scikit-learn, PyTorch, TensorFlow, HuggingFace, or similar).
- Proven track record implementing ML solutions such as classification systems, time series analysis, anomaly detection, or NLP applications.
- Experience with MLOps practices, including model versioning, monitoring, pipeline orchestration, and deployment in high-reliability environments.
- Familiarity with data engineering concepts, including data pipelines, stream processing, message queuing, and working with medium-to-large scale datasets.
- Knowledge of containerized deployment solutions and cloud-native architectures (Kubernetes, Docker).
- Strong communication skills with the ability to explain technical concepts to diverse stakeholders.
- Cybersecurity domain knowledge, particularly in threat detection, threat intelligence, or ICS/OT operations.
- Experience with LLMs, RAG, or advanced NLP techniques is beneficial.
Responsibilities
- Design and implement production-grade machine learning systems that expand Dragos product capabilities, with consideration for both cloud and resource-constrained on-premises environments.
- Build and optimize ML model architectures for ICS/xOT cybersecurity use cases, including threat detection, asset classification, behavioral analysis, anomaly detection, and natural language processing systems.
- Develop robust data pipelines and ML workflows that integrate with existing data infrastructure, supporting both real-time and batch processing requirements.
- Collaborate with OT detection experts to translate research concepts and prototypes into scalable, production-ready ML systems.
- Partner with Data Engineers to establish data contracts and implement observability frameworks for ML pipelines, including monitoring, versioning, and deployment best practices.
- Contribute to ML infrastructure improvements, including automated testing frameworks, CI/CD pipelines, and deployment strategies for containerized environments (Kubernetes, Docker).
- Evaluate and adapt state-of-the-art ML research and open-source models to domain-specific cybersecurity applications.
- Troubleshoot and optimize ML model performance in production environments, addressing issues related to latency, accuracy, and resource utilization.
View Full Description & ApplyYou'll be redirected to the employer's site