Skip to content
โ† All Web3 jobs

VP of Engineering

Hyperbolic

San FranciscoHybrid$300k to $350kToken compensation
About Hyperbolic Hyperbolic Labs is pioneering AI infrastructure with its open-access GPU cloud, aggregating computing resources across the globe to offer an innovative GPU marketplace and AI inference service at up to 75% cost savings compared to traditional cloud providers. The mission is to democratize AI by breaking down barriers to computing power, making it affordable and accessible to developers and researchers everywhere. Founded by co-founders with PhDs in AI, Math, and Computer Science, Hyperbolic has raised a Series A and is preparing for significant growth. The team is approximately 30 people, sitting at the intersection of AI and open-source technology, with an engineering-driven culture emphasizing ownership, execution, and building at the cutting edge of GPU infrastructure. About the Role This is Hyperbolic's first dedicated engineering executive hire, building the Infrastructure, Platform, and SRE functions from the ground up. The VP of Engineering will own infrastructure strategy, organizational growth, and executive-level decision making. This is NOT a step-back-and-manage role. Hyperbolic expects this person to be more than 40% hands-on: personally contributing to architecture reviews, debugging critical production issues, and partnering with engineers on implementation. This hire needs to come from the GPU cloud / AI infrastructure world, not traditional enterprise or big tech. Key Responsibilities - Lead the design and evolution of the AI cloud platform architecture: GPU orchestration, compute scheduling, networking, storage, and distributed systems - Build and scale large GPU clusters supporting customer workloads, including GPU provisioning, scheduling, utilization optimization, and capacity management - Personally participate in architecture reviews, system design, and key technical initiatives (expect 40%+ of time on technical contribution) - Act as the technical escalation point for complex infrastructure challenges: debug production issues, review proposals, and drive decisions - Establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence - Build SRE and Platform Engineering functions from scratch: define SLOs, SLIs, incident response, and capacity planning - Recruit and develop world-class Infrastructure, Platform, and SRE teams - Partner with executive leadership on company strategy and infrastructure investments - Manage infrastructure budgets, vendor relationships, and capacity planning Requirements - 12+ years building and operating large-scale infrastructure systems, with experience leading infrastructure organizations while remaining deeply hands-on technically - Previous experience building or operating a cloud platform at scale, ideally GPU-native cloud infrastructure supporting AI training and inference workloads - Expert-level Kubernetes knowledge and experience designing multi-region cloud infrastructure - Deep expertise in Linux, networking, distributed systems, and storage architecture - Proven track record scaling infrastructure in high-growth startup environments, not just maintaining systems at large companies - Strong understanding of Infrastructure-as-Code, automation frameworks, observability, monitoring, and reliability engineering - Experience building highly available production systems with clear SLOs and incident response processes Bonus Skills - Experience with GPU scheduling, Slurm, Kubernetes GPU operators, Ray, or distributed training systems - Experience managing thousands of GPUs in production environments - Background supporting AI training and inference platforms at scale - Experience with bare-metal provisioning and lifecycle management (IPMI/Redfish, BMC, PXE boot)
Apply on Deciml

Deciml matches Web3 professionals to roles with AI, free for candidates.