Clockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-Hours
Key Highlights
- ➤Clockwork.io raises $31 million, bringing total funding to $73 million
- ➤LinkedIn prevents tens of thousands of GPU-hours of downtime monthly
- ➤Together AI brings TorchPass and LinkPass resilience software to market
- ➤WhiteFiber expands Clockwork.io adoption across its global GPU-as-a-service footprint
- ➤TorchPass adds code-free multi-node snapshots and faster asynchronous checkpoints
Expert Statements
Suresh Vasudevan, CEO of Clockwork.io
“Failures are inevitable at AI scale. Losing hours of useful work to them should not be”
Suresh Vasudevan, CEO of Clockwork.io
“Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done.”
Suresh Vasudevan, CEO of Clockwork.io
“We built our software alongside enterprises and cloud providers operating some of the largest GPU fleets, so it handles the failures they actually see.”
Suresh Vasudevan, CEO of Clockwork.io
“That protection belongs in the infrastructure enterprises and cloud providers rely on every day.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“At AI infrastructure scale, a single network issue should never sideline healthy GPUs or interrupt running workloads.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“Before Clockwork.io, one InfiniBand NIC flap could remove an eight-GPU server from service, while a switch port flap could drain a second server, doubling the impact to 16 GPUs.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“Clockwork.io helped transform that operating model.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“Its network fault-tolerance technology automatically reroutes traffic onto healthy paths, allowing jobs to continue uninterrupted while link, optic, cable, or NIC faults are repaired.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“In aggregate, Clockwork.io prevents tens of thousands of GPU-hours of downtime per month across our fleet.”
Raghu Hiremagalur, SVP, CTO Infrastructure, LinkedIn
“By turning what were once disruptive operational incidents into manageable maintenance events, Clockwork.io has helped improve infrastructure utilization and operational efficiency.”
Pavneet Ahluwalia, Product Lead, Together AI
“Our customers grade us on goodput, the share of their GPU-hours that actually move the model forward,”
Pavneet Ahluwalia, Product Lead, Together AI
“Node repair already detects faults and provisions replacement capacity automatically.”
Pavneet Ahluwalia, Product Lead, Together AI
“Clockwork.io's TorchPass and LinkPass build on that foundation and are designed to keep jobs moving through GPU faults and link failures, preserving progress.”
Pavneet Ahluwalia, Product Lead, Together AI
“We are bringing them to market as the next layer of resilience in the platform.”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“Pressure-testing a cluster's reliability before it reaches production is critical, because a customer who inherits a hidden fabric fault pays for it later in failed jobs and lost GPU-hours,”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“Marginal optics, misconfigured NICs, and links that pass a basic test but degrade under load can slip through.”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“Clockwork.io's automated fleet audit validates every link and node at once, localizes faults in minutes, and lets us correct them before acceptance.”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“We bring clusters up faster, and a customer's first training run lands on a fabric validated end-to-end, not just powered on.”
Tom Sanfilippo, Chief Technology Officer, WhiteFiber
“With market demand growing as rapidly as it is, getting validated capacity to customers quickly is critical to our business, and it is why we are expanding Clockwork.io across our clusters.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“Cluster fault tolerance used to be a training problem. It is now an inference problem too,”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“In our ClusterMAX, TorchPass cuts training goodput loss from 14% to under 3% for a gold-rated neocloud.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“Reinforcement Learning (RL) ties the two together: inference replicas generate rollouts, the trainer learns from them, and the updated weights go back to the replicas.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“Clockwork.io keeps replicas serving through link flaps and network failures.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“Its extremely fast checkpoints accelerate weight transfer back into the rollout fleet, so neither direction stalls the run.”
Dylan Patel, Founder, CEO, and Chief Analyst at SemiAnalysis
“One fault-tolerance layer under training, inference, and RL is where this has to be solved.”
Greg Papadopoulos, Venture Partner, NEA
“The one thing that scales perfectly is unreliability: put enough GPUs in one machine and something is always failing,”
Greg Papadopoulos, Venture Partner, NEA
“The old playbook: stop the job, reload a checkpoint, makes no sense at today's scale.”
Greg Papadopoulos, Venture Partner, NEA
“Clockwork treats failure as the normal state: TorchPass migrates training off a failing GPU live, and now snapshots an entire running job with no code changes.”
Greg Papadopoulos, Venture Partner, NEA
“It's already saving tens of thousands of GPU-hours a month.”
Greg Papadopoulos, Venture Partner, NEA
“We first backed Clockwork.io in 2021 and are thrilled to keep supporting them as they define the performance layer of the AI cluster.”
Clockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-Hours PR Newswire
LinkedIn prevents tens of thousands of GPU-hours of downtime monthly; new TorchPass innovations preserve AI workload progress without code changes and speed reinforcement learning.
Get started
Create a free account to read the full story.
