Staff Software Engineer - Backend
New
J
JobgetherSoftware Engineering
Remote position within the United StatesFull-TimeStaff
SalaryEstimated base salary ranging from $140,000–$210,000 USD for most U.S. locations; $157,000–$235,000 USD for Austin, D.C. Metro, non-Bay Area California, Hawaii, Illinois, Massachusetts, New Hampshire, Oregon, Virginia, and Washington; $166,800–$250,200 USD for the New York City Metro and Kirkland/Seattle areas; $182,000–$273,000 USD for the Bay Area and Los Angeles.
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- AWSDockerPythonKubernetesTypeScriptGoMicroservicesDistributed Systems
Requirements
- 7+ years of professional software engineering experience building and operating large-scale, production-grade distributed systems and microservices.
- Strong proficiency in Python, Go, and/or TypeScript, with deep knowledge of API design, system architecture, scalability, resiliency, and performance optimization.
- Hands-on experience developing CI/CD pipelines, cloud-native applications, and production infrastructure on AWS.
- Strong experience with containerization and orchestration technologies such as Docker, Kubernetes, ECS, or comparable platforms.
- Experience designing and operating highly available production workloads in distributed environments.
- Experience with machine learning or generative AI infrastructure and model-serving technologies.
- Familiarity with observability frameworks and tools such as OpenTelemetry, Prometheus, Grafana, and CloudWatch.
- Experience with vector databases, feature stores, caching technologies, and infrastructure-as-code solutions.
- Understanding of GPU infrastructure, workload scheduling, performance tuning, and cloud cost optimization.
- Proven ability to lead complex technical initiatives, influence architecture, and collaborate effectively with diverse technical and business stakeholders.
- Demonstrated experience mentoring engineers and providing technical leadership within distributed or globally collaborative teams.
Responsibilities
- Lead the design, architecture, and evolution of infrastructure, platforms, and backend services supporting machine learning and generative AI workloads.
- Build and operate highly scalable, reliable, and observable distributed systems and microservices for production environments.
- Deploy and manage ML and generative AI workloads using serving technologies such as vLLM, Triton, TorchServe, SageMaker Endpoints, or comparable frameworks.
- Develop cloud-native infrastructure and services on AWS, supporting high-availability workloads across services such as ECS, EKS, Lambda, DynamoDB, S3, IAM, and CloudWatch.
- Establish and improve observability practices using technologies such as OpenTelemetry, Prometheus, Grafana, and CloudWatch.
- Work with vector databases, feature stores, caching systems such as Valkey/Redis, and infrastructure-as-code technologies including CDK, CloudFormation, or Terraform.
- Contribute to GPU infrastructure management, workload scheduling, performance optimization, and cloud cost efficiency.
- Drive complex technical initiatives, influence architectural decisions, and establish engineering standards across distributed teams.
- Partner closely with ML Scientists, Data Engineers, Product teams, and other stakeholders to translate machine learning capabilities into reliable production solutions.
- Mentor engineers and serve as a technical leader, helping elevate engineering quality and accelerate software delivery through AI-assisted development tools.
View Full Description & ApplyYou'll be redirected to the employer's site