Staff AI Cloud Engineer
New
J
JobgetherAI Cloud Infrastructure
Based in the United StatesFull-TimeStaff
SalaryAnnual base salary range of $180,000–$225,000 USD, plus variable compensation.
Apply NowOpens the employer's application page
Job Details
- Experience
- 15+ years
- Required Skills
- PythonKubernetesSoftware ArchitectureTypeScriptGoCI/CDDistributed Systems
Requirements
- 15+ years of experience in software architecture, engineering, or closely related technical leadership roles.
- Experience as a founding engineer or in a comparable environment where you helped scale technology and engineering practices.
- Proven ability to influence multiple teams and drive technical decisions without formal management authority.
- Strong experience designing, building, deploying, and operating production-grade software systems at scale.
- Deep expertise across AI systems, backend or distributed systems, cloud architecture, and infrastructure.
- Practical understanding of modern AI technologies, including foundation models, inference, retrieval, orchestration, evaluation, latency, monitoring, and cost optimization.
- Strong working knowledge of full-stack software development and modern engineering practices.
- Hands-on experience with Python, Go, TypeScript, Kubernetes, and major cloud platforms.
- Demonstrated ability to operate independently and take ownership of complex technical initiatives.
- Strong architectural judgment and ability to balance long-term technical strategy with practical business needs.
- Excellent communication and collaboration skills with the ability to explain complex technical concepts.
Responsibilities
- Define and continuously evolve architecture, engineering standards, and technical practices across multiple engineering teams.
- Partner with Engineering Managers, technical leads, and senior engineers on complex system-design decisions and high-impact technical initiatives.
- Lead technical designs, RFCs, architecture reviews, and critical code reviews while ensuring solutions remain practical, scalable, and maintainable.
- Build and ship production systems spanning AI, backend services, distributed systems, cloud infrastructure, and full-stack applications.
- Design and operate production AI capabilities, including foundation model integration, inference, retrieval, orchestration, evaluation, monitoring, latency optimization, and cost management.
- Improve CI/CD pipelines, testing strategies, observability, reliability, security, infrastructure, and overall developer experience.
- Create reusable platforms, libraries, reference implementations, and standardized engineering patterns.
- Provide hands-on technical leadership during complex production incidents and contribute directly to resolving challenging infrastructure issues.
- Mentor engineers and technical leaders through hands-on guidance, knowledge sharing, and technical coaching.
View Full Description & ApplyYou'll be redirected to the employer's site