About Immunai:
Immunai is an AI-driven platform company focused on improving drug discovery and development by decoding the human immune system.
We combine large-scale single cell immune data, advanced machine learning, and strong engineering to help pharmaceutical and research partners make better, more informed decisions throughout the drug development process.
Our long-term goal is to reduce drug development failure rates and help more effective medicines reach patients. We’re building this platform thoughtfully and collaboratively, bringing together expertise across biology, AI, engineering, and business.
Immunai is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
We are looking for an ML Infra Engineer to play a central role in evaluating, and deploying Immunai’s AI models. You will own and evolve Immunai’s benchmarking and evaluation capabilities for foundation models and multimodal systems, while also working closely with modeling teams to support model development, iteration, and validation. This role sits at the intersection of software engineering, model understanding, and applied AI, with broad influence on how models are built, compared, and improved across the organization.
This is NOT a core algorithmic research role - but rather, a role focused on building out the ML engineering infrastructure for the evaluation and deployment of models.
Location: Ramat Gan, Israel (hybrid model)
What will you do?
Own & Evolve Benchmarking – Design, build, and maintain Immunai’s benchmarking suite for foundation models and multimodal AI systems.
Define Core Abstractions – Create clean, extensible abstractions and APIs for datasets, tasks, models, metrics, and evaluation workflows.
Develop Metrics & Evaluations – Implement metrics that capture predictive performance, biological relevance, and multimodal alignment.
Support Model Development – Work closely with AI scientists and data scientists to integrate new models, tweak architectures, and enable rapid, fair iteration.
Bring in New Models & Baselines – Add external and internal models to benchmarks and ensure meaningful comparisons.
Explore Data When Needed – Dive into data and results to debug evaluations, understand model behavior, and unblock modeling work.
Enable Rigor & Reproducibility – Ensure evaluations are consistent, well-versioned, and trustworthy over time.