Title: Group Lead - DriveNets Infrastructure Services (DIS)
Hybrid: Ra'anana
#LI-Hybrid
DriveNets is a leader in high-scale networking software for AI infrastructure and service providers. The company pioneered a disaggregated networking architecture that transforms the economics of large-scale networks while maximizing performance, utilization, and operational efficiency. DriveNets-powered networks are deployed by global leaders, including AT&T and Comcast, supporting more than 30% of total U.S. internet traffic. DriveNets AI Fabric delivers full-stack networking for AI infrastructures, providing the highest-performance, Ethernet-based alternative to InfiniBand. The solution is deployed by hyperscalers, NeoClouds, and enterprises worldwide. With over $1B raised, DriveNets continues to push the boundaries of modern networking infrastructure.
עיקרי התפקיד
DriveNets is seeking a Group Lead for its Infrastructure Services (DIS) team to be a key member of our customer-facing technical organization. Join a dynamic and forward-thinking company at the forefront of AI infrastructure. We leverage advanced technologies to develop innovative solutions that drive efficiency, scalability, and exceptional compute performance. Collaborate with the industry's best as we partner with hyperscalers, emerging NeoClouds, and enterprises building large-scale AI/HPC clusters, shaping the future of disaggregated AI networking and compute infrastructure. Our environment fosters creativity, teamwork, and growth, and offers you the opportunity to make a meaningful impact while leading a high-performing team on groundbreaking deployments.
As Group Lead for DIS, you will manage and develop a team of Solution Engineers and Solutions Architects responsible for designing, deploying, and optimizing DriveNets' AI/HPC infrastructure solutions at customer sites. You will provide technical leadership across the full customer lifecycle - from pre-sales architecture and POC execution through deployment, performance benchmarking, and ongoing operations. You will work cross-functionally with Sales, Product Management, and Engineering to ensure customer success, drive product feedback, and continuously raise the bar for technical delivery quality across the team.
Lead and develop the DIS team - a group of Solution Engineers and Solutions Architects - setting technical direction, managing execution, and fostering a culture of ownership, learning, and customer focus.
Oversee end-to-end customer engagement for DIS - from pre-sales technical support and solution architecture through POC planning, deployment execution, and post-deployment operations.
Serve as the senior technical escalation point for customer infrastructure challenges, including AI cluster performance issues, networking design trade-offs, and operational reliability concerns.
Partner with Sales Account Managers to support business opportunities, lead technical responses to RFP/RFQs, and influence technical decision-makers at the VP and CxO level.
Guide the team in conducting performance benchmarking activities - including NCCL/RCCL, RDMA, and LLM benchmarks - and ensure results are translated into actionable product and deployment insights.
Work with Product Management and Engineering to funnel customer requirements, field observations, and performance data into the product roadmap and development backlog.
Define and drive internal processes for deployment planning, operational readiness, monitoring standards, and technical documentation across the DIS team.
Build and maintain relationships with compute, NIC, and storage partners to support joint POCs, reference deployments, and solution validation.
Represent DriveNets at industry events and conferences, and contribute to external technical content including white papers, blogs, and design guides.
Recruit, mentor, and grow team members, and establish clear performance goals aligned with business objectives.
דרישות
What we need to see:
10+ years of experience in AI/HPC infrastructure, data center networking, or solutions architecture, with at least 2-3 years in a technical leadership or team lead capacity.
Hands-on technical depth across both compute infrastructure (GPU clusters, Linux systems, AI workloads) and data center networking (routing, switching, fabric design), with the ability to engage credibly across both disciplines.
Proven experience leading customer-facing technical teams through complex deployment and POC cycles in AI/HPC or data center environments.
Strong understanding of AI cluster architecture - including GPU platforms (NVIDIA, AMD), RDMA networking, storage connectivity, and the interaction between compute, network, and storage layers.
Experience with performance benchmarking methodologies (NCCL/RCCL, RDMA, LLM workloads) and the ability to interpret and act on results at a system level.
Demonstrated ability to work cross-functionally with Sales, Product Management, and Engineering teams, translating customer feedback into product improvements and go-to-market strategy.
Excellent communication and presentation skills, with proven ability to influence technical and executive stakeholders at customer organizations.
Ability to write extensive technical content (white papers, technical briefs, design guides, etc.) for external audiences with a balance of technical accuracy and clear messaging.
Ability to travel domestic and international.
Ways to stand out from the crowd:
Deep familiarity with AI-relevant infrastructure technologies - InfiniBand, RoCEv2, lossless Ethernet (PFC, ECN), GPU, NIC, DPU, and accelerated computing platforms.
Hands-on experience deploying and operating large-scale AI/HPC clusters, including GPU resource scheduling (Slurm, Kubernetes), monitoring (Prometheus, Grafana, DCGM), and operational tooling.
Understanding of scale-up (NVLink, UALink) and scale-out (Enhanced Ethernet, UEC, InfiniBand) interconnect technologies and their design trade-offs.
Experience with CCL tuning (NCCL/RCCL), GPU environment setup, and performance optimization across large multi-node GPU clusters.
Familiarity with AI/ML frameworks (PyTorch, TensorFlow) and how workload characteristics interact with infrastructure design decisions.
Proven experience with one or more Tier-1 Clouds (AWS, Azure, GCP, or OCI) or emerging NeoClouds, and cloud-native architectures and software.
Background in data center operations fundamentals - networking, cooling, power, and rack-level design.
Experience engaging compute, NIC, or storage vendors on joint solution definition, reference architecture development, or benchmarking programs.
EDUCATION
BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or other Engineering fields, or equivalent experience.
More About DriveNets
Based in Israel with locations in Romania, US, India and Japan as well as extended teams, DriveNets operations cover more than twelve countries. With recognition by industry analysts and through partnerships with market leaders such as AMD, Broadcom, Dell and others, DriveNets is pushing market momentum, delivering the scale and efficiency that modern AI workloads demand. Visit our website: https://drivenets.com/company
If your experience is close but doesn’t fulfil all requirements, please submit your application. DriveNets is on a mission to build a special company comprised of individuals with different backgrounds, perspectives, and experiences.
DriveNets is an equal opportunity employer. We do not discriminate based on upon race, religion, national origin, sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with disability, or other applicable legally protected characteristics.