Senior Software Engineer (Serverless)
New
J
JobgetherCloud Infrastructure
UKFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- KubernetesGoServerlessDistributed Systems
Requirements
- 7+ years of professional software engineering experience with a proven track record of building and operating large-scale distributed systems.
- Strong programming experience with Golang, or the ability and willingness to quickly become proficient in the language.
- Deep experience with Kubernetes and container orchestration systems, including real-world operation of production environments.
- Strong understanding of distributed systems concepts such as consistency, availability trade-offs, queueing, backpressure, retries, idempotency, and multi-tenancy.
- Experience designing and operating high-throughput, low-latency services with a strong focus on performance optimization and reliability.
- Proven ability to take ownership of complex technical challenges, lead design discussions, unblock teams, and deliver impactful solutions.
- Strong software engineering fundamentals with the ability to write reliable, maintainable code and investigate complex technical problems.
- Experience with serverless platforms or function-as-a-service technologies such as Knative, AWS Lambda, GCP Cloud Run, Cloudflare Workers, or similar solutions.
- Familiarity with GPU scheduling technologies, including Kubernetes device plugins, MIG, MPS, time-slicing, or NVIDIA GPU Operator.
- Experience with ML inference technologies such as vLLM, TensorRT-LLM, Triton Inference Server, SGLang, or similar platforms.
- Knowledge of runtime optimization techniques including cold-start reduction, image streaming, checkpoint/restore, or sandboxing technologies.
- Experience developing Kubernetes operators using Go and frameworks such as controller-runtime or kubebuilder.
Responsibilities
- Design, develop, and maintain core components of a GPU-native serverless AI platform, including control planes, schedulers, runtimes, autoscaling systems, and customer-facing APIs.
- Solve complex engineering challenges related to cold-start optimization, GPU scheduling, multi-tenant isolation, fair resource allocation, request routing, and platform scalability.
- Own technical architecture decisions for key platform areas by creating design documents, evaluating solutions, and aligning engineering teams around effective approaches.
- Establish and maintain high engineering standards through code reviews, design reviews, technical guidance, and active collaboration with team members.
- Operate services with an SRE mindset by defining reliability objectives, improving observability, supporting incident response, and driving continuous platform improvements.
- Collaborate directly with customers on architecture discussions, performance optimization, and complex production issues requiring deep technical expertise.
- Partner with product, infrastructure, and go-to-market teams to translate customer needs into scalable technical roadmaps.
- Support knowledge sharing and technical mentorship to help raise the overall engineering capability of the team.
View Full Description & ApplyYou'll be redirected to the employer's site