Research Engineer Intern - AI Systems

New
Y
Yotta LabsAI Infrastructure
United States, Hong Kong, Canada, SingaporeInternshipEntry
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonFGPA ArchitecturePyTorchC++

Requirements

  • Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field.
  • Solid programming skills in Python and familiarity with C++.
  • Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy).
  • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels.
  • Understanding of AI frameworks such as PyTorch, Dynamo, or LMCache.
  • Familiarity with model architectures and profiling tools like Nsight, ROCm Profiler, or Neuron Profiler.
  • Ability to work independently in a collaborative, remote environment.

Responsibilities

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium.
  • Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes including kernel fusion, scheduling, KV-cache, and memory optimizations.
  • Build benchmarks, identify performance regressions, and convert profiler traces into performance improvements.
  • Ship code upstream to open-source AI infrastructure projects, including tests and documentation.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now