Software Engineer, Infrastructure & Reliability
New
C
CrewAIAI Infrastructure
United StatesFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSDockerPostgreSQLPythonKubernetesGoRedisCI/CDGitHub Actions
Requirements
- Strong infrastructure/platform engineering experience in production SaaS environments.
- Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
- Experience with ECS and/or Kubernetes.
- Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
- Strong debugging instincts across application, infrastructure, network, deployment, and dependency layers.
- Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
- Ability to write reliable automation in Python, Ruby, Go, Bash, or similar languages.
- Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Responsibilities
- Own and improve the infrastructure that runs CrewAI's platform including AWS, ECS, ECR, Docker, Kubernetes, and networking.
- Build and maintain CI/CD pipelines for build, test, image publishing, migrations, and deployment safety.
- Improve reliability across cloud and enterprise deployments through health checks, alerting, capacity planning, and operational runbooks.
- Partner with runtime and product engineers on optimizing Celery, FastAPI, Redis, Rails, and Postgres production workloads.
- Manage production observability and telemetry infrastructure including logs, metrics, traces, and dashboards.
- Harden security and compliance posture across IAM, workload identity, secrets management, and vulnerability scanning.
- Build tooling and automation such as Helm charts and release artifacts to enable customer self-hosted installations.
- Reduce operational toil by automating recurring workflows.
View Full Description & ApplyYou'll be redirected to the employer's site