Software Engineer, Infrastructure & Reliability
New
C
CrewAIAI Infrastructure
United StatesFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSDockerPostgreSQLPythonKubernetesGoRedisCI/CD
Requirements
- Strong infrastructure/platform engineering experience in production SaaS environments.
- Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
- Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
- Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
- Strong debugging instincts across app, infra, network, deploy, and dependency layers.
- Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
- Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
- Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Responsibilities
- Own and improve the infrastructure that runs CrewAI's platform including AWS, ECS, Docker, Kubernetes, and networking.
- Build and maintain CI/CD pipelines for build, test, deployment, and environment promotion.
- Improve reliability across cloud and enterprise deployments through health checks, alerting, incident response, and capacity planning.
- Partner with runtime and product engineers on optimizing Celery, FastAPI, Redis, Rails, and Postgres workloads.
- Manage production observability and telemetry infrastructure, including logs, metrics, traces, and dashboards.
- Harden security and compliance posture across IAM, secrets management, and vulnerability scanning.
- Build tooling and automation, such as Helm charts and release artifacts, to support self-hosted customer installs.
View Full Description & ApplyYou'll be redirected to the employer's site