Software Engineer, Infrastructure & Reliability

New
C
CrewAIAI Infrastructure
United StatesFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSDockerPostgreSQLPythonKubernetesGoRedisCI/CD

Requirements

  • Strong infrastructure/platform engineering experience in production SaaS environments.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers.
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.

Responsibilities

  • Own and improve the infrastructure that runs CrewAI's platform including AWS, ECS, Docker, Kubernetes, and networking.
  • Build and maintain CI/CD pipelines for build, test, deployment, and environment promotion.
  • Improve reliability across cloud and enterprise deployments through health checks, alerting, incident response, and capacity planning.
  • Partner with runtime and product engineers on optimizing Celery, FastAPI, Redis, Rails, and Postgres workloads.
  • Manage production observability and telemetry infrastructure, including logs, metrics, traces, and dashboards.
  • Harden security and compliance posture across IAM, secrets management, and vulnerability scanning.
  • Build tooling and automation, such as Helm charts and release artifacts, to support self-hosted customer installs.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now