Clockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-Hours
Key Highlights
- ➤Clockwork.io raises $31 million, lifting total funding to $73 million
- ➤LinkedIn prevents tens of thousands of GPU-hours of downtime monthly
- ➤Together AI brings TorchPass to market on its GPU Clusters
- ➤WhiteFiber expands Clockwork.io across its global GPU-as-a-service footprint
- ➤TorchPass snapshots preserve distributed job progress without code changes
Expert Statements
Suresh Vasudevan, CEO of Clockwork.io
“Failures are inevitable at AI scale. Losing hours of useful work to them should not be”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“At AI infrastructure scale, a single network issue should never sideline healthy GPUs or interrupt running workloads. Before Clockwork.io, one InfiniBand NIC flap could remove an eight-GPU server from service, while a switch port flap could drain a second server, doubling the impact to 16 GPUs.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“Clockwork.io helped transform that operating model. Its network fault-tolerance technology automatically reroutes traffic onto healthy paths, allowing jobs to continue uninterrupted while link, optic, cable, or NIC faults are repaired. In aggregate, Clockwork.ioprevents tens of thousands of GPU-hours of downtime per month across our fleet. By turning what were once disruptive operational incidents into manageable maintenance events, Clockwork.io has helped improve infrastructure utilization and operational efficiency.”
Pavneet Ahluwalia, Product Lead, Together AI
“Our customers grade us on goodput, the share of their GPU-hours that actually move the model forward. Node repair already detects faults and provisions replacement capacity automatically. Clockwork.io's TorchPass and LinkPass build on that foundation and are designed to keep jobs moving through GPU faults and link failures, preserving progress. We are bringing them to market as the next layer of resilience in the platform.”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“Pressure-testing a cluster's reliability before it reaches production is critical, because a customer who inherits a hidden fabric fault pays for it later in failed jobs and lost GPU-hours.”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“Marginal optics, misconfigured NICs, and links that pass a basic test but degrade under load can slip through. Clockwork.io's automated fleet audit validates every link and node at once, localizes faults in minutes, and lets us correct them before acceptance. We bring clusters up faster, and a customer's first training run lands on a fabric validated end-to-end, not just powered on. With market demand growing as rapidly as it is, getting validated capacity to customers quickly is critical to our business, and it is why we are expanding Clockwork.io across our clusters.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“Cluster fault tolerance used to be a training problem. It is now an inference problem too.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“In our ClusterMAX, TorchPass cuts training goodput loss from 14% to under 3% for a gold-rated neocloud. Reinforcement Learning (RL) ties the two together: inference replicas generate rollouts, the trainer learns from them, and the updated weights go back to the replicas. Clockwork.io keeps replicas serving through link flaps and network failures. Its extremely fast checkpoints accelerate weight transfer back into the rollout fleet, so neither direction stalls the run. One fault-tolerance layer under training, inference, and RL is where this has to be solved.”
Greg Papadopoulos, Venture Partner, NEA
“The one thing that scales perfectly is unreliability: put enough GPUs in one machine and something is always failing”
Greg Papadopoulos, Venture Partner, NEA
“The old playbook: stop the job, reload a checkpoint, makes no sense at today's scale. Clockwork treats failure as the normal state: TorchPass migrates training off a failing GPU live, and now snapshots an entire running job with no code changes. It's already saving tens of thousands of GPU-hours a month. We first backed Clockwork.io in 2021 and are thrilled to keep supporting them as they define the performance layer of the AI cluster.”
Clockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-Hours PR Newswire
LinkedIn prevents tens of thousands of GPU-hours of downtime monthly; new TorchPass innovations preserve AI workload progress without code changes and speed reinforcement learning.
Get started
Create a free account to read the full story.
