VAST is building out a Presales Platform team responsible for the tooling, automation, and infrastructure that enables our field engineering organization to demonstrate value at scale. Our lab is a core asset for evaluations, demos, and internal enablement, and we're building a team whose job is to keep it running at production-quality reliability.
This role holds the senior technical ownership on that team. You'll own the reliability of the VAST clusters in our lab environment, influence the direction for the automation, tooling, and infrastructure-as-code practices the rest of the team builds on and executes within, and serve as the escalation point for the hardest technical problems in the lab. You'll partner with a growing team of lab and platform engineers, as well as the internal teams who rely on these tools day to day, to understand their needs and raise the operational bar of everything we run.
עיקרי התפקיד
Own the operational reliability of VAST clusters in the lab environment, including proactive health monitoring, upgrade planning, and issue resolution
Serve as the primary technical escalation point for complex cluster issues, working hands-on-keyboard to resolve them and partnering with VAST engineering when deeper investigation is needed
Shape the automation and tooling strategy for the lab environment, including provisioning scripts, CLI utilities, monitoring dashboards, and internal tooling that the rest of the team builds on
Experience with virtualization platforms (VMware vSphere, ESXi, Proxmox, or equivalent)
Methodical approach to troubleshooting complex issues across storage, networking, and compute layers
Comfort working across time zones with distributed team members
Excellent written and verbal communication, including ability to produce clear technical documentation and reports
יתרון
Existing hands-on experience with VAST Data clusters
Prior experience as a Customer Support Engineer at a storage or infrastructure company OR Reliability Engineer / SRE supporting a wide variety of infrastructure services
Familiarity with observability platforms (Grafana, Prometheus, Elasticsearch)