Position Overview & Description
SynthMind AI is building next-generation LLM serving infrastructure. We need a Senior ML Infra Engineer to optimize inference latencies and build distributed training pipelines.
Key Responsibilities
• Build high-throughput GPU serving infrastructure for LLMs.
• Reduce model inference latency and GPU memory overhead.
• Collaborate with AI research scientists to deploy experimental models.
Role Requirements & Qualifications
• Python, C++, PyTorch, CUDA, and Triton inference server experience.
• Hands-on experience scaling GPU cluster workloads.
• M.S. or B.S. in Computer Science or related field.
Benefits & Perks
• Top-tier competitive salary + Generous AI equity grant.
• Unlimited PTO policy.
• Premium hardware selection (MacBook Pro M3 Max / RTX Workstations).