
Systems ML Engineer (Member of the Technical Staff)
TransfyrMEMBER OF THE TECHNICAL STAFF - SYSTEMS ML ENGINEER
About Transfyr
Transfyr is building physical AI for science.
Why is it that a professional athlete has dramatically more information about every play they make than a scientist has about the cause of any experimental failure? At Transfyr, we are building the infrastructure to make real-world scientific work legible, transferable, and reproducible.
Modern science is capable of extraordinary outcomes, but much of the most important insights never become explicit: how experiments are actually executed, protocols drift, how experts make gametime decisions on the fly, why experiments fail on Tuesdays. This tacit knowledge is rarely captured, making it difficult to reliably reproduce results, much less hand off protocols to new team members or collaborators. We believe our systematic failure to capture tacit knowledge is holding back the entire industry.
We're building systems that operate directly in real laboratory environments to elucidate, capture, and interpret this missing information. Our platform records and analyzes multimodal data about how scientific work is performed and turns it into durable, operational knowledge. In doing so, we are also building the world's largest commercial dataset on real-world scientific execution.
This foundation is critical not only for driving elite human performance today, but for enabling meaningful automation tomorrow. Physical AI systems cannot learn from outcomes alone; they require rich, grounded records of how work is actually done in the real world.
The Role
Systems ML Engineers at Transfyr ensure our models train and run efficiently across their full lifecycle – from large-scale training through production inference. You will own performance optimization across our ML stack, working embedded with the research team to make training faster and more efficient, and with our perception and production systems to get models running well at inference, across both cloud and edge infrastructure.
We are looking for engineers who combine deep understanding of ML systems with hands-on performance engineering skill. The ideal Systems ML Engineer has experience profiling and optimizing large-scale models for both training and inference, is comfortable writing custom GPU kernels when off-the-shelf ops aren't fast enough, and can manage cloud and edge infrastructure that holds up in real lab environments.
This role spans deep ML and performance engineering and cloud/DevOps responsibilities. You will profile and optimize training and inference workloads, manage cloud and edge infrastructure, optimize cloud spend, ensure security and compliance, and work closely with the perception and research teams to squeeze more performance out of every model we ship.
We're building a team, and we have needs across levels, from hands-on builders early in their careers to senior engineers who enjoy shaping training infrastructure and technical direction.
This role is in-person in Cambridge, MA (other locations may open in the future, feel free to reach out even if Boston is not currently an option for you).
What You'll Accomplish with Us
- Profile and Optimize Performance: Use profiling tools (e.g., Nsight, PyTorch Profiler) to identify bottlenecks in data loading, gradient computation, and communication, and implement optimizations like kernel fusion, sharding, and tiling to improve step time.
- Optimize Distributed Training: Improve the efficiency of distributed training pipelines using frameworks like PyTorch Distributed, working closely with the research team on training performance.
- Develop Custom Kernels: Design and maintain high-performance GPU kernels in Triton or CUDA for performance-critical ML workloads.
- Build Data Pipelines: Design and optimize data loading pipelines that maximize training throughput, and inference pipelines that reliably serve models on real-world, multimodal lab data.
- Handle Cloud and Edge: Manage deployment across both cloud infrastructure and edge devices running in active lab environments, where compute and connectivity are more limited.
- Keep Production Running: Debug and resolve performance bottlenecks, resource issues, and failures across the training and deployment stack.
- Work Closely with Research and Perception: Partner with the research team on training efficiency and with perception engineers to get multimodal data (vision, audio, sensor, metadata) flowing reliably through your pipelines.
- Make Updates Safe: Build monitoring, versioning, and rollback into deployments so model updates don't break production.
Who You Are
- High agency. You don't wait for perfect datasets or well-posed problems. You identify what needs to be learned, build the right scaffolding, and push work forward.
- Biased toward action. You prototype quickly, test assumptions against real data, and iterate based on failure rather than waiting for theoretical certainty.
- Successful in ambiguity. You can make progress when labels are incomplete, feedback is delayed, and success criteria evolve over time.
- Thoughtful. You understand when sophistication helps and when it obscures, and you make deliberate tradeoffs between model complexity, robustness, and operational cost.
- Clear, direct communicator. You can explain model performance and limitations to collaborators across engineering, science, and operations.
- Intense. You care deeply about the mission, work hard when it matters, and help keep the team oriented toward what actually moves the needle.
What You Know
- Systems ML Engineering Expertise: Demonstrated expertise in ML systems engineering, including optimizing and deploying large-scale models in production, debugging and fixing performance and stability issues in deployed systems, building infrastructure for reproducible, monitored ML deployments, and optimizing inference throughput and resource utilization across cloud and edge.
- Distributed Training & Serving: Deep knowledge of distributed training and serving frameworks, including PyTorch/JAX distributed strategies, gradient accumulation, mixed precision training, and checkpoint/recovery systems, and how to apply that knowledge to efficient, reliable model serving.
- Cloud & Edge Infrastructure: Strong cloud administration skills, including AWS services, infrastructure as code (Terraform), Kubernetes orchestration, cost optimization, security best practices, and compliance requirements — plus experience deploying and optimizing ML systems on edge or on-prem infrastructure.
- Full ML Stack Understanding: Understanding of the ML stack from hardware (GPUs, interconnects, storage) through frameworks (PyTorch, JAX) to deployment and serving, and what it takes to move a model from research prototype to reliable production.
- Cross-Stack Debugging: Skilled at debugging complex failures across the stack – GPU/NCCL issues, data loading bottlenecks, memory leaks, and performance or convergence problems in both training and deployment.
- Algorithm Optimization: Deep experience optimizing algorithms for cloud and edge environments, including computer vision and other ML algorithms, with GPU-level work like CUDA and kernel tuning.
Other Things We Like to See
- A passion for and experience in science
- A passion for and experience with AI
- Demonstrated experience working in fast-moving/ambiguous environments (like startups!)
The Basics
- Competitive compensation (cash + equity)
- Full benefits (low/no-cost health insurance options, HSA, 401K with matching, lunch subsidy, etc.)
About Transfyr
<cite index="27-1">Transfyr develops AI-powered systems and infrastructure to capture, interpret, and distribute scientific knowledge, making scientific work legible, transferable, and reproducible to accelerate innovation.</cite>
Interested in this role?
Apply now to join Transfyr.
