Research Engineer Intern - AI Systems
New
Y
Yotta LabsAI Infrastructure
United States, Hong Kong, Canada, SingaporeInternshipEntry
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonFGPA ArchitecturePyTorchC++
Requirements
- Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field.
- Solid programming skills in Python and familiarity with C++.
- Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy).
- Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels.
- Understanding of AI frameworks such as PyTorch, Dynamo, or LMCache.
- Familiarity with model architectures and profiling tools like Nsight, ROCm Profiler, or Neuron Profiler.
- Ability to work independently in a collaborative, remote environment.
Responsibilities
- Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium.
- Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
- Profile and improve inference performance in vLLM, SGLang, and custom runtimes including kernel fusion, scheduling, KV-cache, and memory optimizations.
- Build benchmarks, identify performance regressions, and convert profiler traces into performance improvements.
- Ship code upstream to open-source AI infrastructure projects, including tests and documentation.
View Full Description & ApplyYou'll be redirected to the employer's site