Senior Engineer, Infrastructure
New
J
JobgetherAI Infrastructure
Germany; Fully remote position with the flexibility to work from Europe.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSPostgreSQLGCPKubernetesCI/CDNetworkingDistributed Systems
Requirements
- Strong software engineering background with experience writing and maintaining production-quality code.
- Proven experience designing, building, and operating infrastructure on major cloud platforms such as GCP, AWS, or similar environments.
- Hands-on production experience with Kubernetes and containerized infrastructure.
- Strong understanding of cloud networking and security concepts, including VPCs, IAM, load balancing, DNS, firewalls, WAFs, CDNs, and service identity.
- Experience with infrastructure-as-code tools and automated deployment systems.
- Ability to debug and solve complex issues across software applications, infrastructure components, networking layers, and distributed systems.
- Experience operating large-scale distributed technologies such as OpenSearch, Elasticsearch, PostgreSQL, or similar systems.
- Strong focus on reliability, security, developer experience, and infrastructure cost efficiency.
- Proactive mindset with the ability to identify improvements, eliminate operational complexity, and take ownership of critical systems.
- Strong engineering fundamentals, problem-solving skills, and the ability to quickly learn unfamiliar technologies.
- Experience deploying or operating large language models with serving frameworks such as vLLM or SGLang is considered an advantage.
Responsibilities
- Design, build, and operate reliable infrastructure platforms supporting AI-powered products and large-scale production workloads.
- Own and improve Kubernetes environments, including cluster architecture, networking, ingress, service communication, workload isolation, autoscaling, and deployment practices.
- Develop and maintain cloud infrastructure across areas such as networking, security, identity management, secrets, firewalls, content delivery, and access controls.
- Improve system reliability through enhanced observability, monitoring, alerting, incident response processes, disaster recovery planning, and resilience improvements.
- Build and optimize infrastructure automation using infrastructure-as-code, CI/CD pipelines, GitOps workflows, and reusable platform tooling.
- Investigate and resolve complex production issues across application, infrastructure, networking, and data system layers.
- Improve infrastructure efficiency through cost optimization, capacity planning, resource management, and operational improvements.
- Support scalable data systems and distributed technologies, ensuring strong performance and reliability.
- Identify architectural bottlenecks and proactively develop solutions that prepare the platform for future growth and increasing AI workloads.
- Reduce operational complexity by replacing manual processes with automated, scalable engineering solutions.
View Full Description & ApplyYou'll be redirected to the employer's site