Senior Engineer, Infrastructure

New
J
JobgetherAI Infrastructure
Germany; Fully remote position with the flexibility to work from Europe.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPostgreSQLGCPKubernetesCI/CDNetworkingDistributed Systems

Requirements

  • Strong software engineering background with experience writing and maintaining production-quality code.
  • Proven experience designing, building, and operating infrastructure on major cloud platforms such as GCP, AWS, or similar environments.
  • Hands-on production experience with Kubernetes and containerized infrastructure.
  • Strong understanding of cloud networking and security concepts, including VPCs, IAM, load balancing, DNS, firewalls, WAFs, CDNs, and service identity.
  • Experience with infrastructure-as-code tools and automated deployment systems.
  • Ability to debug and solve complex issues across software applications, infrastructure components, networking layers, and distributed systems.
  • Experience operating large-scale distributed technologies such as OpenSearch, Elasticsearch, PostgreSQL, or similar systems.
  • Strong focus on reliability, security, developer experience, and infrastructure cost efficiency.
  • Proactive mindset with the ability to identify improvements, eliminate operational complexity, and take ownership of critical systems.
  • Strong engineering fundamentals, problem-solving skills, and the ability to quickly learn unfamiliar technologies.
  • Experience deploying or operating large language models with serving frameworks such as vLLM or SGLang is considered an advantage.

Responsibilities

  • Design, build, and operate reliable infrastructure platforms supporting AI-powered products and large-scale production workloads.
  • Own and improve Kubernetes environments, including cluster architecture, networking, ingress, service communication, workload isolation, autoscaling, and deployment practices.
  • Develop and maintain cloud infrastructure across areas such as networking, security, identity management, secrets, firewalls, content delivery, and access controls.
  • Improve system reliability through enhanced observability, monitoring, alerting, incident response processes, disaster recovery planning, and resilience improvements.
  • Build and optimize infrastructure automation using infrastructure-as-code, CI/CD pipelines, GitOps workflows, and reusable platform tooling.
  • Investigate and resolve complex production issues across application, infrastructure, networking, and data system layers.
  • Improve infrastructure efficiency through cost optimization, capacity planning, resource management, and operational improvements.
  • Support scalable data systems and distributed technologies, ensuring strong performance and reliability.
  • Identify architectural bottlenecks and proactively develop solutions that prepare the platform for future growth and increasing AI workloads.
  • Reduce operational complexity by replacing manual processes with automated, scalable engineering solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now