Verified from career page · Posted 4w ago
Research Engineer, Foundation Model (Multimodal Fusion & Distillation)
Laelaps AI · Research & Development
Zurich
- Posted
- 4w ago
- Workplace
- On-site
- Salary
- Not disclosed
- Visa sponsorship
- Not specified
Posted on 5 September 2026
Work model: On-site
Salary range not shared by the company
Visa sponsorship details unknown
OUR MISSION
At Laelaps AI, we believe robotics is entering a transformative decade, much like the arrival of the internet. Advances in AI, cloud computing, and hardware are reshaping what autonomous systems can do. Our mission is to build the intelligent software that powers physical security in the real world - enabling robots and sensors to handle dangerous and critical tasks that humans shouldn't have to. By engineering the orchestration layer for intelligent security, we aim to create a world that is safer, more secure, and more resilient.
We're a strong founding team based in Zurich, backed by visionary investors and advisors. We are engineering the future of security today!
THE ROLE
You will build the parts of LEA, our foundation model of the physical world, where field data is the input. LEA takes in cameras, infrared, thermal, lidar and robot state natively, around the clock, on real security sites. We stand on the best open models for the encoders and the reasoning head and do not pretrain from scratch. Your work goes where nobody else can do it: fusing every sensor into one stream, teaching an open vision-language model to reason over that stream, and distilling a fast model that sees every frame.
This is the first engineer dedicated to the model track, so you will also set how we train, evaluate and ship models. You will work with our Command Centre and Robot Autonomy teams to get what you train running on site appliances and robots, and you will see it tested at 03:00 in the rain.
WHAT YOU'LL WORK ON
- Native encoders: adapt the best open vision and point-cloud foundation models to RGB, IR, thermal and lidar, one encoder per sensor, raw stream in.
- Fusion: build the layer that aligns every sensor in time and space into one token stream, trained on time-aligned multi-sensor field data that exists nowhere else.
- Slow head: fine-tune an open frontier vision-language model to reason over fused sensor tokens instead of pictures.
- Fast head: distil a small model from the slow head's calls on our own events, so every frame is seen and only hard cases escalate.
- Evaluation: held-out sites, night and thermal slices, regression gates. A model that is worse anywhere does not ship.
- Training machinery: multi-GPU training, data loaders for multi-sensor clips, experiment tracking and reproducible runs.
WHO WE'RE LOOKING FOR
A research engineer who has shipped multimodal models, not only published them. You know when an open model is good enough and when to build, you measure on the slice that matters before claiming a number, and you care whether the model works on a real site, not only on a benchmark.
YOUR BACKGROUND
- 4+ years training and fine-tuning large vision or multimodal models, or a PhD plus industry experience.
- Strong PyTorch, comfortable with multi-GPU and multi-node training.
- Hands-on fine-tuning of vision-language models (instruction tuning, parameter-efficient methods, full fine-tunes).
- Experience with knowledge distillation or model compression for deployment.
- Work with video or non-RGB sensors: thermal, infrared, lidar or radar.
- Evaluation discipline and solid software engineering hygiene.
NICE TO HAVE
- Point-cloud or 3D foundation models.
- Self-supervised and latent-prediction methods (JEPA-style).
- Edge deployment with TensorRT or similar.
- Robotics or field data collection.
- Publications at top ML, vision or robotics venues.
WHAT WE OFFER
- Ownership: you are able to ship products and deliver project end-to-end.
- Mission: autonomous security that keeps people and critical sites safe, including in defence.
- Career path: a ground-floor seat with real runway. Prove your value and you will not have barriers to grow.
- Team: work directly with PhD-level co-founders in AI, Robotics, and Physics, alongside a strong (and fun) founding team.
- Compensation: Competitive equity/salary package
- Culture: International founding team that is serious about building but does not take itself too seriously.
About Laelaps AI
Laelaps AI builds perception and vision software for robotics and autonomous systems. A Swiss startup.