Max Shad

Senior Director of AI/ML Research Engineering

Contact Information

About

Max Shad is the Senior Director of AI/ML Research Engineering at Harvard University’s Kempner Institute, where he leads the institute’s computational program. He combines hands-on research engineering with technical and strategic leadership, overseeing foundation model development, AI platforms and tools, and large-scale computing infrastructure. His team builds models and the systems researchers use to train, serve, and study them at scale. Max is the architect of the Kempner AI cluster and founded Harvard’s first central Research Software Engineering team in 2019.

AI/ML Research Engineering and Foundation Models

Max leads the team’s AI/ML research engineering projects and remains hands-on in the code. The team pre-trains billion-parameter vision-language models from scratch to advance model performance and develop new architectures. Its broader work spans mixture-of-experts models, reinforcement learning systems, flow mapping optimization, NeuroAI foundation models, and agentic shared research infrastructure on AI clusters.

AI Platforms and Agentic Tools

Max also directs the development of platforms, workflows, agentic AI solutions, and tools that let researchers train, serve, and monitor models at scale. These include KempnerForge, a PyTorch-native framework for fault-tolerant, large-scale distributed training of foundation models; KempnerInsight, a web application that monitors the AI cluster in real time; and KempnerPulse, a GPU monitoring tool. 

The team also publishes open-source recipes for serving open-weight models and builds MLOps pipelines for the AI cluster and the cloud. Max leads the development of the Kempner Computing Handbook which has been accessed from more than 110 countries.

Large-Scale AI Infrastructure

As the architect of the Kempner AI cluster, Max leads its design and long-term planning in partnership with FAS Research Computing and the Massachusetts Green High Performance Computing Center (MGHPCC). His work spans compute, storage, and network architecture; vendor partnerships; procurement and service negotiations; and performance benchmarking. He also leads monitoring and automation efforts that have improved cluster reliability.

The cluster includes more than 1,100 NVIDIA GPUs across A100, H100, H200, and RTX PRO 6000 Blackwell models, with InfiniBand networking and high-performance storage. The cluster ranked third among U.S. academic supercomputers on the November 2024 and June 2026 TOP500 lists. It also ranked 32nd worldwide on the November 2024 Green500 list.

Research Engineering at Harvard

In 2019, Max founded Harvard’s first central Research Software Engineering (RSE) team and helped define the university’s RSE job family. The team partnered with researchers across Harvard’s schools on projects spanning scalable data platforms, generative AI, computational neuroscience analysis platforms, and cloud-native solutions. Its collaborations extended across neuroscience, life sciences, medical sciences, engineering, computer science, applied mathematics, astrophysics, business, public health, and design.