Software Engineer Project Intern (Model Infrastructure) - 2026 Start (BS/MS)
About this role
Public source summary from intern-list.com / Jobright.
TikTok is a short-form mobile video platform seeking a Software Engineer Project Intern for its Model Infrastructure team. The intern will optimize large-scale recommendation model training and inference, develop LLM-integrated recommendation infrastructure, and work on distributed streaming, hardware-aware co-design, and model state management.
Responsibilities
Engineering Efficiency at Scale: Drive the optimization of training and inference pipelines to maximize hardware utilization (MFU/HFU) for models featuring hundreds of billions of dense parameters LLM2Rec Infrastructure: Architect specialized systems to support the integration of LLMs into the recommendation stack, focusing on memory-efficient attention mechanisms and advanced KV cache management for long-sequence user modeling Massive Sparse & Dense Streaming: Build and optimize high-concurrency engines for Petabyte-scale streaming training, handling continuous parameter updates and high-frequency data ingestion without compromising stability Hardware-Aware Co-Design: Work closely with researchers to design next-generation recommendation architectures optimized for modern GPU/NPU interconnects, ensuring high-bandwidth utilization across the cluster Distributed State Management: Innovate on how we store and synchronize massive model states across heterogeneous memory hierarchies (HBM, DDR, and NVMe)
Qualifications: Currently pursuing an Undergraduate/Master in Software Development, Computer Science, Computer Engineering, or a related technical discipline Strong programming skills in C++ and Python Solid understanding of Computer Architecture and the GPU software stack (CUDA, Triton, or NCCL) Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and a desire to "look under the hood" of model execution runtimes A strong interest in solving system-level bottlenecks in large-scale distributed environments Experience with Transformer-based architectures, 3D parallelism (TP/PP/DP) Deep understanding of the torch.compile stack, including TorchDynamo (graph acquisition) and TorchInductor (lowering) Hands-on experience writing high-performance kernels or optimizing collective communication (e.g., customizing NCCL/UCX) Familiarity with RDMA networking, high-performance storage, or specialized Parameter Server architectures Success in programming competitions (ACM-ICPC) or contributions to prominent open-source AI infrastructure or high-performance computing projects
Benefits: Interns have day one access to health insurance, life insurance, wellbeing benefits and more. Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year). Interns who are not working 100% remote may also be eligible for housing allowance.