Operations and Support Lead
Operations, Customer Service
Miami, FL, USA · Remote
Hydra Host operates mission-critical AI infrastructure where customer success depends on operational excellence.
This role combines Customer Support Leadership with Infrastructure Operations Management, serving as the operational hub between customers, engineering, deployment teams, hardware vendors, and AI Factory partners.
You'll own customer-facing operational support while building the internal processes that keep our NeoCloud platform running efficiently. Whether responding to a GPU outage, coordinating infrastructure deployments, managing vendor escalations, improving SLAs, or building scalable operational workflows, you'll ensure both our customers and our infrastructure perform at the highest level.
This is not a traditional support management position.
This is an operations leadership role responsible for the daily execution, reliability, and continuous improvement of Hydra Host's AI Factory platform.
What You'll Do
Lead Customer Support Operations
Develop and lead Hydra Host's customer support organization supporting:
- Design and build Hydra Host’s customer support organization from the ground up
- Enterprise AI customers
- Interface with the Machine Learning engineering teams
- GPU infrastructure customers
- AI Factory operators
- Data center partners
- Build a high-performing support organization that delivers exceptional customer experiences while maintaining enterprise-grade service levels.
Own Operational Excellence
Drive the day-to-day operational health of Hydra Host's NeoCloud platform by coordinating activities across engineering, infrastructure, vendors, and customer-facing teams.
Ensure infrastructure operates reliably while continuously improving operational efficiency.
Lead Major Incident Management
- Serve as the Incident Commander during production-impacting events.
- Coordinate engineering, networking, infrastructure, vendors, and customers to resolve:
- GPU cluster failures
- Network outages
- Hardware failures
- Firmware issues
- Storage performance degradation
- Infrastructure capacity constraints
- Customer-impacting production incidents
- Own customer communications throughout incident response while driving rapid resolution and post-incident improvements.
Build Scalable Support Operations
- Design and implement:
- Ticketing systems
- Escalation procedures
- Knowledge management
- Support automation
- AI-powered support tools
- On-call rotations
- Operational playbooks
- Customer communication standards
- Create a support organization capable of scaling alongside Hydra Host's rapid growth.
Drive AI Factory Operations
Partner with deployment, engineering, and data center teams to coordinate:
- Infrastructure deployments
- Rack turn-up
- GPU cluster readiness
- Network activation
- Capacity planning
- Maintenance scheduling
- Production acceptance
- Operational readiness reviews
- Help ensure AI Factory infrastructure is deployed efficiently and operates reliably.
Vendor & Partner Management
Own operational relationships with:
- Hardware manufacturers
- GPU vendors
- Data center operators
- Construction partners
- Logistics providers
- Network carriers
- Service providers
- Track vendor performance, manage escalations, enforce SLAs, and ensure timely issue resolution.
SLA & Service Delivery
Establish, monitor, and improve operational KPIs including:
- SLA compliance
- MTTR
- Incident response times
- Customer satisfaction
- Infrastructure uptime
- Vendor performance
- Capacity utilization
- Operational readiness
- Provide executive reporting on operational performance and identify opportunities for continuous improvement.
Cross-Functional Leadership
Collaborate daily with:
- Infrastructure Engineering
- Network Engineering
- Platform Engineering
- Customer Success
- Deployment Program Managers
- AI Factory Partners
- Executive Leadership
- Enterprise Customers
- Act as the operational bridge between technical teams and customer-facing organizations.
Build the Organization
As Hydra Host grows, you'll help recruit, mentor, and develop the Operations & Support organization while establishing the culture, processes, and operational standards that define world-class infrastructure operations.
Required Qualifications
- 5+ years leading customer support, technical operations, infrastructure operations, or service delivery organizations
- Experience supporting enterprise infrastructure, cloud platforms, AI infrastructure, NeoCloud providers, or large-scale data center environments
- Experience managing production incidents in mission-critical environments
- Strong understanding of servers, networking, storage, and enterprise infrastructure
- Experience building operational processes that scale rapidly
- Experience working with third-party vendors and infrastructure providers
- Excellent communication skills with both technical and executive stakeholders
- Proven ability to lead cross-functional initiatives across engineering, operations, and customer organizations
- Strong organizational and project management skills
- Bias toward ownership, accountability, and continuous improvement
Preferred Qualifications
- Experience operating NeoCloud or AI Factory infrastructure
- Experience supporting NVIDIA GPU environments (HGX, DGX, H100, H200, Blackwell)
- Experience with bare-metal cloud infrastructure
- Familiarity with InfiniBand, RoCE, high-performance networking, and AI storage platforms
- Experience with ITIL Service Management, Incident Management, Change Management, and Problem Management
- Experience implementing AI-driven support automation and operational tooling
- Experience with CRM and service management platforms including Zendesk, Jira Service Management, Linear, HubSpot, Intercom, or similar platforms
- Previous experience managing technical support or operations teams