Staff Software Engineer, IDSM
New
J
JobgetherAI/ML infrastructure
Ability to work remotely from Canada.Full-TimeStaff
SalaryAnnual US base salary range of USD $200,000–$235,000.
Apply NowOpens the employer's application page
Job Details
- Experience
- At least 8 years of experience in software engineering, primarily in an individual contributor capacity.
- Required Skills
- AWSDockerArtificial IntelligenceKubernetesMachine LearningMicrosoft AzuregRPCCI/CDRESTful APIsDistributed Systems
Requirements
- At least 8 years of software engineering experience, primarily in an individual contributor capacity.
- Strong knowledge of AI and machine learning systems, including model development, training, optimization, deployment, and lifecycle management.
- Experience designing, building, and maintaining scalable backend systems in distributed computing environments.
- Proficiency designing and implementing secure, high-performance RESTful APIs and/or gRPC services.
- Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
- Understanding of backend performance profiling, troubleshooting, and optimization in cloud-based or distributed environments.
- Familiarity with containerization and orchestration technologies such as Docker and Kubernetes.
- Experience implementing unit, integration, and end-to-end automated tests and maintaining CI/CD pipelines.
- Ability to collaborate across engineering teams and integrate backend services with frontend applications and external systems.
- Experience with Apache Spark, AWS SageMaker, or Azure Machine Learning is an advantage.
- Familiarity with model registries, model monitoring, enterprise MLOps, and LLM hosting infrastructure is beneficial.
Responsibilities
- Design, develop, and maintain high-performance backend services for enterprise-scale data science and machine learning workflows.
- Collaborate with customers and internal teams to implement model deployment solutions with cloud machine learning platforms.
- Improve capabilities for model development, training, registration, deployment, and ongoing management.
- Help design and launch a centralized data science catalog for discovering, exploring, and summarizing platform resources.
- Integrate monitoring capabilities for deployed model health, performance, and operational status.
- Expand tagging and metadata functionality across platform entities.
- Extend large language model hosting capabilities for scalability, performance, and operational logging.
- Build secure APIs, including RESTful APIs and gRPC services, and integrate backend systems with frontend interfaces and third-party services.
- Profile, troubleshoot, and optimize backend applications and distributed workloads across cloud and containerized environments.
- Implement testing practices and contribute to CI/CD pipelines and architectural decisions.
View Full Description & ApplyYou'll be redirected to the employer's site