• Resources
  • Blogs
  • Synchronization in data centers: why timing is now an AI problem

Synchronization in data centers: why timing is now an AI problem

Stefano Ruffini
17 Jun 2026
Artificial Intelligence
Data Center
Synchronization in data centers: why timing is now an AI problem

Synchronization has quietly become one of the most consequential – and least visible – challenges in modern data center infrastructure.

As AI workloads scale and distributed systems grow more complex, getting timing right is no longer a nice-to-have. It’s a prerequisite for correct operation.

At WSTS 2026, Stefano Ruffini, Strategic Technology Manager at Calnex Solutions, presented the latest developments in data center synchronization – covering new standards activity, real-world test findings, and what the growing relationship between timing and AI actually looks like in practice.

Why synchronization is now business-critical

The case for precise timing in data centers now spans several distinct needs: consistent operation between distributed nodes, meaningful event logging and diagnostics, financial regulatory compliance, and – increasingly – the efficient operation of AI and machine learning workloads. Distributing synchronization across servers also carries potential benefits for power consumption, a commercially significant consideration as AI infrastructure scales.

AI training and inference involve coordinating work across large numbers of servers, sometimes spanning multiple sites. When clocks drift relative to one another, the consequences range from degraded observability to outright logical failures. Until recently, operators largely developed their own approaches. The industry is now moving toward standardized solutions.

The standards landscape

The standout development is the publication of ITU-T G Suppl.92 (October 2025) – a major step toward industry-wide agreement on timing requirements, solutions, and clock classifications for data centers. The Supplement defines three accuracy classes for Time Sync Clocks:

  • Class 1 (5 µs relative time error): suited to distributed databases.
  • Class 2 (1 µs): high-frequency telemetry and multi-node performance analysis.
  • Class 3 (200 ns): congestion control based on one-way delay and time-synchronized collective communication.

Alongside ITU-T, the Open Compute Project’s Time Appliances Project (TAP) has developed a PTP profile for data centers and a practical test guide covering time receivers, boundary clocks, and transparent clocks. IEEE is also active with P3335 (Timecard standard) and P1588.1 (a simplified timing protocol).

The fact that ITU-T, OCP, and IEEE are all working in coordination reflects how quickly synchronization has become a priority across data center infrastructure.

What the data actually shows

One of the most valuable aspects of the WSTS presentation was real measurement data – from lab testing and a purpose-built AI inference testbed developed in the context of the OCP/TAP Unified Intelligent Architecture project.

The PTM difference. Tests comparing clock behavior with and without Precision Time Measurement (PTM) revealed a significant gap. Without PTM, CPU load changes cause time errors of around 1 µs that the system cannot recover from even after load stabilizes. With PTM, errors remain in the region of 250 ns, and the system recovers to near its baseline. The implication: for AI workloads regularly varying CPU utilization, PTM becomes essential if you need to stay within Class 2 or Class 3 accuracy requirements.

Timing and AI causality. A more striking set of results came from an AI inference pipeline experiment measuring the impact of clock skew on system observability. The pipeline was tested with injected increasing skew values. The findings were consistent: zero violations at zero skew, then immediate and sustained causality violations once skew exceeded a small threshold – while token throughput remained completely unaffected.

The experiment tracked two key metrics: negative timing spans – when timestamps imply impossible causal relationships, such as events being received before they were sent – and causality health, scored as 1 when no violations occur within a rolling 30-second window and 0 otherwise. Negative timing spans increased exponentially, and the average causality health score dropped from 1.0 to 0.03.

The conclusion is clear: bad timing may not break an AI system’s output – but it breaks your ability to know whether the system is working correctly. Timing is no longer just a network concern. It is an application-layer requirement.

Timing as a network diagnostic tool

A separate thread explored how PTP itself can serve as a network observability tool. By using PTP between synchronized endpoints to measure mean path delay, it’s possible to detect latency changes that indicate emerging network issues – such as route changes, disruptions, or anomalies – before they cause visible failures. Synchronization infrastructure becomes not just a service to be managed, but a source of diagnostic signal in its own right.

Where the industry stands

The data center sector is at an early but rapidly moving point in formalizing its synchronization practice. Where operators once largely developed their own approaches, there is now coordinated industry effort – across ITU-T, IEEE, and OCP – defining what accuracy is required for which applications, how to meet the accuracy, how to verify equipment meets those requirements, and how to validate performance at the application layer where AI workloads actually run. The pace of that progress, across three major bodies simultaneously, reflects how quickly synchronization has moved from a background consideration to a core infrastructure requirement.

stefano

About the Author

Stefano Ruffini is Strategic Technology Manager at Calnex Solutions. He presented “What is Happening with Synchronization in Data Centers?” at WSTS 2026.