Job Description
Welcome to the vanguard of technology. At 2026 Systems, we are building the neural backbone of the next decade. We are looking for a highly skilled Senior AI Infrastructure Engineer to architect the high-performance computing environments that power our next-generation generative AI models. If you are passionate about scaling AI, optimizing deep learning clusters, and solving complex engineering challenges, we want to meet you.
Our mission is to democratize access to advanced intelligence. As a key member of our infrastructure team, you will be responsible for ensuring our AI models run at lightning speed across distributed systems.
Responsibilities
- Design, build, and maintain scalable, high-performance AI inference and training infrastructure on cloud platforms (AWS/GCP).
- Implement and optimize containerization strategies (Docker, Kubernetes) for machine learning workloads.
- Collaborate with research scientists to integrate new models into production pipelines with minimal latency.
- Monitor system health, optimize resource allocation, and implement cost-reduction strategies for GPU clusters.
- Establish and enforce security best practices and compliance standards for sensitive AI data.
Qualifications
- 5+ years of experience in DevOps, SRE, or Systems Engineering, with a focus on AI/ML.
- Expert knowledge of Python, Bash scripting, and container orchestration tools.
- Strong proficiency with Deep Learning frameworks (TensorFlow, PyTorch) and model serving tools (Triton, TorchServe).
- Experience managing large-scale GPU clusters and optimizing for throughput.
- Excellent problem-solving skills and the ability to thrive in a fast-paced, agile environment.