
Same pattern on every program I’ve been near. CPU cluster: argued about for a quarter. Accelerator: properly staffed, own lead, own schedule. The interconnect tying it all together gets whatever’s left once the block list freezes. Then timing closure lands and suddenly it’s the only thing on the agenda.
Decide It While You Still Can
A network on chip is an architecture call. It is not a back-end chore you hand downstream after IP selection. The trouble is sequencing. Lock memory topology, let the floorplan acquire intent, settle your IP list, and you’ve quietly removed most of the fabric’s options before anyone sat down to discuss the fabric. None of those felt like interconnect decisions at the time. They were.
So put it in the same meeting where you’re fighting about memory channels and where the heavy traffic actually goes. Teams that do this lose fewer weekends to re-floorplanning in month eight.
Datasheet Bandwidth Tells You Almost Nothing
Aggregate numbers look comfortable. They usually are, right up until the ugly moment when four masters all want the same controller inside the same window.
Model the traffic early. Rough models are fine, you’re hunting for spikes and starvation, not three significant figures. One design I sat in on had generous headroom on paper and a single latency-sensitive block still blew its deadline, because the worst case had never been simulated. That was a respin. An expensive one, and entirely avoidable.
QoS Is a Day-One Argument
Masters are not interchangeable. Display can’t tolerate jitter. CPU wants latency. A DMA engine mostly wants throughput and genuinely doesn’t mind queuing.
Treat them as equals and eventually something urgent ends up stuck behind something that could have waited all afternoon. Define your classes while the architecture is still soft. Map each initiator on purpose, not by default. Then go verify that the arbitration in your NoC interconnect actually honours that under saturation, because nominal support and real behaviour when the fabric is jammed are not the same animal.
Physics Gets a Vote
Long wires lose. At advanced nodes the distance between a clean topology diagram and something routable keeps widening.
That tidy mesh? Possibly unbuildable once the hard macros sit where they’re going to sit. Bring physical design in before topology is settled, not after. Link width, pipeline depth, the shape of the thing, negotiate all of it against an actual floorplan instead of choosing on a whiteboard and throwing it over the wall.
Build for the Part Nobody Has Scoped Yet
Very little tapes out once. There’s a cost-reduced variant, a next-gen, some customer asking for a version with half the accelerators.
Fabric hand-tuned to exactly one configuration will fight you on every one of those. Keep it parameterized. Leave headroom you don’t currently need. And leave the clever micro-optimizations alone where they sit on things likely to move. The area you concede is almost always cheaper than six weeks of rework on the derivative.
Interface Checks Are Not System Checks
Block-level verification catches protocol violations. Fine. It does not catch deadlock, livelock, or the ordering problems that only surface when the whole fabric is under real pressure.
Those want system stimulus, formal on the paths where deadlock is plausible, and stress cases built deliberately to jam things up. Worth remembering, too, that these are the nastiest bugs to find in silicon. Intermittent. Workload-dependent. Close to impossible to reproduce on demand for whoever draws the short straw on debug.
Short Version
Settle the fabric early. Model real traffic. Assign QoS deliberately. Let the floorplan have an opinion. Leave headroom. Verify at system level. None of it is glamorous. All of it beats finding out post-tapeout.