Transporting the native packetized traffic of the NoC directly across die boundaries using a stable, invariant interface.
The semiconductor industry is rapidly transitioning from monolithic SoCs to chiplet-based architectures. At the same time, AI workloads are evolving beyond cloud inference into physical AI: systems that perceive, reason, and act in the real world. Autonomous vehicles, industrial robots, humanoid robots, drones, and intelligent manufacturing systems all require continuous, high-volume data flow between heterogeneous compute engines.
However, these two trends expose a weakness in today’s die-to-die communication architectures. While coherent interconnect protocols have been highly successful for processor-centric systems, they are fundamentally mismatched to the traffic patterns generated by modern AI accelerators.
As a result, a new approach is needed that involves transporting the native packetized traffic of the NoC directly across die boundaries using a stable, invariant interface.
Traditional processors are dominated by memory accesses, synchronization, and cache coherence. Physical AI systems are different. A modern automotive compute platform may simultaneously execute:
Rather than exchanging individual cache lines, these engines often exchange continuous streams of feature maps, tensors, point clouds, video frames, and intermediate inference results. These data streams often consist of hundreds of kilobytes, or even megabytes, moving between specialized accelerators at deterministic rates. The communication pattern resembles a high-performance dataflow network more than a coherent symmetric multi-processor system. As physical AI systems become more capable, the amount of accelerator-to-accelerator traffic continues to grow faster than CPU-to-memory traffic.
Coherent die-to-die protocols remain the right solution for many applications and perform exceptionally well for:
These protocols were designed to maintain memory consistency while moving relatively small cache-line-sized transactions. For processor-centric workloads sharing memory, they are extremely effective.
Accelerator traffic follows very different patterns. Most coherent protocols divide communication into cache-line-sized transactions, typically 64 bytes each. Every cache line carries protocol overhead, including:
When transferring large tensors or feature maps, this overhead must be repeated thousands of times. For AI workloads, the actual payload accounts for only a portion of the transmitted bandwidth, with protocol overhead consuming an increasingly significant fraction of the available die-to-die bandwidth.
Some coherent protocols have added optimizations to help improve the efficiency of non-coherent, bursty traffic. While useful, these mechanisms are limited in scope. For example, they span only a small number of beats and cannot fully ameliorate the overhead associated with very long streaming transfers of arbitrary data. As accelerator bandwidth increases into the terabytes-per-second range inside future automotive AI systems, repeatedly transmitting cache-line metadata becomes increasingly difficult to justify.
Many non-coherent die-to-die approaches have been proposed, including:
These approaches extend AXI interfaces across die boundaries and work well for connecting peripherals and control-oriented IP. However, they still operate at the transaction layer. Each NoC packet of each virtual channel must be converted into AXI transactions before crossing the die boundary, then reconstructed as NoC packets on the receiving side. As systems scale to dozens of AI engines across multiple chiplets, this repeated conversion adds unnecessary latency, buffering, and implementation complexity.
Fig. 1: Multi-die interconnect in physical AI applications. Source: Arteris, Inc.
A more natural architecture is to allow the NoC itself to span multiple dies. Instead of terminating the NoC at each chiplet boundary, native NoC packets are transported directly across the die-to-die interface. The remote die simply becomes another portion of the same logical network. This provides several important advantages:
Automotive AI places unique demands on interconnects. Unlike cloud AI systems, automotive platforms must satisfy strict latency, determinism, power, and functional safety requirements while operating within tight thermal limits. Centralized vehicle compute platforms will integrate multiple AI chiplets connected through advanced packaging technologies. Individual chiplets may specialize in:
The largest traffic flows are often not between CPUs and memory but between accelerators themselves. For example:
These are fundamentally streaming dataflow applications. Transporting these streams using a native packetized NoC enables far more efficient bandwidth utilization, while reducing latency and buffering throughout the system. The result is higher sustained throughput, lower energy consumption per transferred bit, and greater scalability as additional accelerator chiplets are added.
Despite these advantages, directly exposing an implementation’s internal NoC protocol is problematic. Internal packet formats frequently evolve as NoC implementations gain new capabilities, such as:
Even relatively small NoC changes can significantly alter the internal protocol, making long-term die-to-die compatibility challenging. Routing introduces another challenge. Packet routing information may depend on the topology of the destination NoC. If the topology changes, routing vectors or addressing schemes may need to be updated. Without architectural stability, every NoC revision risks requiring die-to-die interface updates. While this may be manageable when a single team designs all dies, it becomes impractical when teams work independently.
The solution is to define an invariant architectural interface, rather than exposing the evolving internal NoC implementation. An invariant, die-to-die NoC interface provides a stable packet protocol specifically designed for die-to-die transport while allowing each chiplet’s internal NoC implementation to continue evolving independently.
Such an interface can:
This enables chiplets from different product generations to communicate through a stable architectural interface, while allowing incremental change on either side of the die-to-die boundary.
We propose a multi-die NoC packet transport interface based on an invariant die-to-die NoC interface. Instead of transporting cache-line transactions or extending AXI buses, the interface carries native, protocol-independent NoC packets directly between chiplets using a stable, implementation-independent packet format. The interface can be sized to match the bandwidth requirements of future AI systems, supports virtual channels for efficient sharing among diverse traffic types, and extends NoC routing seamlessly across die boundaries.
For physical AI, particularly in next-generation automotive computing platforms, this approach aligns the communication architecture with the system’s dominant traffic pattern: high-bandwidth, low-latency accelerator-to-accelerator data movement. By eliminating unnecessary protocol translation and minimizing transaction overhead, it creates a scalable, multi-protocol, multi-die interconnect capable of supporting increasingly complex AI pipelines, while maintaining the determinism, efficiency, and long-term architectural stability required for safety-critical deployments.
Learn more about Arteris’ multi-die solution.
Leave a Reply
Facts Only
* Transporting native packetized traffic of the NoC across die boundaries requires a stable, invariant interface.
* The semiconductor industry is transitioning from monolithic SoCs to chiplet-based architectures.
* Physical AI systems require continuous, high-volume data flow between heterogeneous compute engines for applications like autonomous vehicles and robotics.
* Accelerator communication involves exchanging continuous streams of feature maps, tensors, point clouds, video frames, and intermediate inference results at deterministic rates.
* Coherent die-to-die protocols are designed for processor-centric systems involving memory accesses and cache coherence.
* Accelerator traffic follows patterns different from coherent protocols, often involving large streaming transfers.
* Coherent protocols serialize communication into cache-line sized transactions, where protocol overhead consumes significant bandwidth when transferring large tensors.
* Non-coherent die-to-die approaches convert NoC packets to AXI transactions before crossing die boundaries, adding latency and complexity for multi-chiplet systems.
* The proposed solution involves allowing the native NoC to span multiple dies by transporting native packets directly across the interface.
* This architecture supports high-bandwidth, low-latency accelerator-to-accelerator data movement.
Executive Summary
The semiconductor industry is shifting from monolithic SoCs to chiplet-based architectures, while AI workloads are evolving toward physical AI systems requiring high-volume data flow between heterogeneous compute engines, such as in autonomous vehicles and robotics. Current coherent interconnect protocols are optimized for processor-centric memory accesses, which contrasts with the traffic patterns of physical AI systems that involve continuous streams of large data like feature maps and tensors. Accelerator communication exhibits a pattern resembling a dataflow network rather than symmetric multi-processor interaction, where bandwidth demands from accelerator-to-accelerator traffic are growing rapidly.
The text posits that existing coherent protocols introduce significant overhead when handling the streaming nature of AI traffic because they serialize communication into cache-line sized transactions, where protocol metadata consumes a large fraction of the available bandwidth for large data transfers. Non-coherent die-to-die approaches exist but require repeated protocol translation from NoC packets to AXI transactions, which adds latency and complexity as systems scale across multiple chiplets.
The proposed solution is to transport native, packetized traffic directly across die boundaries using a stable, invariant interface, allowing the NoC itself to span multiple dies. This approach directly addresses the demands of automotive AI by facilitating efficient, high-bandwidth data movement between accelerators while maintaining architectural stability despite evolving internal NoC implementations.
Full Take
The core tension presented is between established architectural paradigms—coherent protocols optimized for memory coherence versus stream-based accelerators—and the emerging demands of physical AI systems which prioritize massive, deterministic data flow. The analysis reveals a fundamental mismatch where existing solutions force inefficient protocol translation, suggesting that optimizing communication for processor state (cache lines) actively hinders optimization for dataflow and streaming throughput in modern heterogeneous systems.
The proposed shift towards an invariant interface emphasizes the necessity of decoupling the internal implementation details from the external connectivity layer. This is not merely an engineering convenience; it addresses a systemic limitation where incremental evolution within one subsystem (the NoC implementation) forces disruptive, expensive changes across the entire system boundary when data streams must cross those boundaries. The implication is that architectural stability—the ability for independent evolution of internal logic without breaking external compatibility—is a prerequisite for scaling complex, safety-critical systems like automotive AI.
The pursuit of this invariant interface speaks to a pattern where functional requirements (e.g., real-world perception demands) must dictate the communication architecture, rather than being constrained by the legacy structure of the processing elements themselves. The skepticism arises in whether achieving this invariance introduces new points of complexity or slows down innovation; however, the text suggests that by defining an invariant interface for native packets, the system achieves scalability and efficiency gains commensurate with the demands of AI dataflow, rather than compromising functionality.
Bridge Questions: If the overhead of protocol translation is removed entirely, what are the specific design constraints imposed on the physical layer to ensure true protocol-independent routing across varying die topologies? How does the concept of an "invariant interface" affect the long-term viability of hardware abstraction layers when chiplet vendors pursue divergent internal protocol evolutions? What is the verifiable cost-benefit ratio between the latency reduction offered by native packet transport and the complexity introduced by managing distributed, evolving routing information?
Sentinel — Human
This text is a dense, technically precise argument proposing a novel interconnect solution based on understanding the mismatch between traditional processor communication and modern AI dataflow requirements.
