PCIe, lanes & device links
Packet-based serial links connect devices through a negotiated topology.
A physically long slot can still have only a few electrically connected lanes.
Connects the host topology.
Find the bottleneck
Low arithmetic intensity hits the memory ceiling. More data reuse can move the workload toward the compute ceiling.
What this model includes
An ideal upper bound with fixed peak compute and bandwidth. Ignores latency, overhead, cache-level traffic, and instruction mix.
What happens inside
Train and route a link
PCIe is a point-to-point interconnect with lanes carrying traffic in both directions. Devices negotiate supported speed and width. A root complex, switches, and endpoints form the topology. Configuration space identifies devices; BARs expose resource windows. Physical connector size does not guarantee wired lane count.
Account for useful bandwidth
Packets carry requests and completions, with encoding and protocol overhead. Generation, width, payload size, and shared upstream links constrain throughput. A PCIe memory write or read has device-specific ordering and access rules. Newer generations change signaling and framing, so a single formula does not fit all generations.
What this means for your code
Low-level engineer
Inspect negotiated speed/width, BAR mappings, MSI vectors, and DMA capabilities. MMIO needs proper access primitives and ordering.
Software developer
Check topology before assuming peak transfer rates. A GPU, SSD, and network card may share an upstream bandwidth limit.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.