A 100,000-GPU cluster can waste expensive compute without a single GPU failing. One congested or flapping link can hold thousands of accelerators at a collective communication barrier, billing power while useful work stops. That makes the network part of the compute system. It also explains why shortening a few centimeters of copper inside a switch has become a board-level decision with data-center economics attached.

Asianometry's account of the AI bandwidth wall gets the mechanism right: switch silicon has scaled faster than the electrical path to faceplate optics. Co-packaged optics, or CPO, moves electro-optical conversion beside the switch ASIC. The useful question for an infrastructure buyer is narrower than whether light will replace copper. It is where the electrical channel has become expensive enough that giving up a replaceable pluggable module is a rational trade.

The wall is measured in joules and lost GPU time

A conventional optical switch has three distinct paths. Bits leave the switch ASIC through a serializer/deserializer, cross package bumps and a printed circuit board, pass through a connector, and enter a pluggable transceiver. The transceiver converts the electrical signal into light. At the far end, the process runs backward.

The board trace looks trivial next to kilometers of fiber. At 200 Gb/s per electrical lane, it is difficult. Loss, reflections, crosstalk, and connector discontinuities close the signal eye. Equalizers, retimers, and digital signal processors recover it by spending power. More cooling then removes that power. The loop gets worse as lane rate and radix rise.

NVIDIA's engineering description puts a concrete scale on its own design comparison: a traditional 200 Gb/s channel can lose about 22 dB before optical conversion, versus roughly 4 dB when its optical engine sits beside the ASIC. The company reports about 30 watts per pluggable interface and 9 watts in its CPO design. Those figures explain the direction of the benefit. They should not be copied into another vendor's business case because module type, reach, coding, cooling, and test boundary change the result.

Broadcom provides a second data point. Its 51.2 Tb/s Bailly platform packages eight 6.4 Tb/s optical engines with a Tomahawk 5 switch and claims 70 percent lower optical-interconnect power than pluggables. Broadcom also says pluggable optics can exceed half of switch-system power at 51.2 Tb/s. Again, that is a vendor measurement, but it tells buyers where to look. The relevant denominator is the complete switch at an identical delivered capacity.

AI makes this more valuable than it was in ordinary cloud networks. Training uses all-reduce, all-gather, and other collective operations that synchronize many workers. A delayed message can stall the group. The cost of a link is therefore not limited to watts and purchase price:

network cost = equipment + energy + cooling + stranded accelerator time + service interruption

That last pair can dominate. A small improvement in job completion time across thousands of accelerators can pay for a more expensive fabric. A nominally efficient link that is hard to recover after failure can erase the saving.

What CPO changes inside the box

CPO removes most of the lossy electrical channel between the switch and optics. A photonic integrated circuit contains modulators, waveguides, and photodetectors. An electronic companion die drives the modulators and reads the detectors. The pair sits on the switch package substrate or a closely connected interposer. Fiber exits the package toward the front panel.

This architecture creates four practical gains:

  1. Lower electrical-channel power. A shorter path needs less equalization and may avoid a retimer.
  2. Higher shoreline density. Fiber connectors replace an array of hot, bulky pluggable cages at the faceplate.
  3. Fewer active parts per link. Removing connectors and DSP stages can improve reliability if the package itself has high yield.
  4. A path beyond switch CPO. The same integration methods can eventually put optical I/O beside GPU and CPU packages, where scale-up fabrics need dense links.

The term still covers several physical arrangements. On-board optics puts the engine on the PCB. Near-package optics moves it closer through a high-speed connection. Full CPO places it on the same substrate as the host ASIC. Treat those as different products, since their electrical reach, repair boundary, and thermal behavior differ.

Standards are progressing but not finished. The Optical Internetworking Forum's implementation agreements include a co-packaging framework, a 3.2 Tb/s module agreement, and an external-laser form factor. That helps vendors agree on module and light-source interfaces. It does not make one supplier's optical engine, package, management software, and repair procedure interchangeable with another's.

TSMC's COUPE approach shows why packaging matters. A mature photonics process can make waveguides and detectors while a more advanced process makes the electronics; three-dimensional bonding shortens the connection between them. The company's 2025 annual report describes continued work on COUPE and its integration with advanced packaging. This is an ecosystem bet involving foundry processes, known-good-die testing, fiber attach, lasers, thermal control, and switch design. It is not a transceiver moved a few centimeters.

The serviceability bill arrives later

A technician can pull a failed pluggable transceiver in minutes. A failed optical engine soldered beside a high-value switch ASIC is a package or board event. CPO has to win enough power, density, or reliability to cover that loss of modular repair.

External lasers are one answer. Silicon is a poor light emitter, so CPO systems commonly generate light in a separate module and route it into the package. The OIF co-packaging framework notes that an external source can improve repairability because failed lasers can be replaced without replacing the switch. It also lets designers keep temperature-sensitive lasers away from a hot ASIC. The price is an optical power-distribution network with coupling loss, connectors, monitoring, and failover logic.

NVIDIA says its external-laser design uses four times fewer lasers than traditional optics and keeps them in a controlled, field-replaceable module. Its announced Spectrum-X and Quantum-X photonics switches claim 3.5 times better power efficiency, ten times greater resiliency, and 1.3 times faster deployment. Buyers should read those as hypotheses to validate in their topology. "Resiliency" may combine fewer components, redundant lasers, software recovery, and modeled fleet behavior. It is not the observed failure rate of every future installation.

Heat is the second bill. Switch ASICs dissipate hundreds of watts. Photonic resonators shift wavelength with temperature, and laser efficiency and lifetime depend on a stable thermal environment. A CPO package must remove heat without bending the package, drifting optical alignment, or forcing continuous tuning that consumes the power it saved. Fiber attach also has micrometer-scale alignment requirements and has to survive assembly and field handling.

A broad technical review of CPO identifies thermal management, optical power delivery, link budget, standardization, and manufacturability as connected problems. This is the right mental model. CPO moves complexity from the faceplate into packaging and system operations; it does not delete complexity.

Qualification also needs to follow the new integration boundary. A pluggable program can qualify the host and module separately against a common electrical interface. CPO couples switch silicon, optical engines, fiber attach, laser delivery, firmware, and cooling. Require lot-level yield data, accelerated thermal cycling, high-temperature operating life, vibration results, and optical-margin distributions rather than a single passing prototype. Ask which tests occur at wafer, known-good-die, package, board, and completed-system stages. A late optical failure can discard an otherwise good and expensive switch package, so yield belongs in total cost. If the vendor uses lane sparing or redundant engines to reach an availability target, price the lost capacity and define the alarm threshold that triggers replacement. This evidence also exposes interoperability limits: a standard external laser connector does not guarantee that another supplier's laser power, control protocol, or warranty will work with the installed switch.

Use a failure-domain table during procurement:

Failure Replaceable unit Required evidence
Laser degrades External laser module Redundancy, switchover time, alarm coverage
Fiber connection fails Cable or connector Cleaning, inspection, bend limits, spare strategy
One optical lane fails Engine, board, or package Lane sparing, degraded mode, repair time
Thermal drift rises Cooling or control loop Margin across inlet temperature and workload
Switch ASIC fails Whole switch Fabric reroute and job recovery behavior

A decision framework for pluggable, near-package, or CPO

Start with the constraint that blocks the planned cluster, then choose the least integrated architecture that clears it. CPO is a poor default and a strong targeted answer.

Stay with pluggables when deployments are small or heterogeneous, ports run below the density limit, operators value multi-vendor replacement, or hardware changes frequently. Pluggables also make sense when the switch power saving is small beside an underused GPU fleet. Fix scheduling and utilization before buying exotic optics.

Use linear pluggable or near-package designs when electrical-channel power is painful but serviceability still ranks high. These intermediate designs can remove some DSP or trace loss without binding the optical engine to the host ASIC. Check their reach and analog link margin carefully, since simplifying the module moves responsibility into the host.

Use CPO when faceplate density, per-port power, or reliability blocks a 51.2 Tb/s and faster fabric, and the organization can operate at the switch-level replacement boundary. Large homogeneous AI clusters are the clearest fit because they run links hard and can amortize validation across many identical units.

Score the options with weighted, measured criteria:

Criterion Measurement Why it matters
Useful efficiency Joules per completed training step or 1M tokens Captures compute stalls and network power
Bandwidth delivery Sustained collective bandwidth at target scale Peak port rate hides congestion
Availability Link-flap-free hours and job restarts AI jobs magnify short interruptions
Service Mean time to diagnose and restore CPO changes the replaceable unit
Thermal headroom Error rate across inlet and load range Photonics and hot ASICs share a package
Supply Qualified sources, lead time, spare coverage Integration can concentrate vendor risk
Upgrade value Capacity and usable lifetime A dense port that ages out early is expensive

Weighting should follow the facility. A power-constrained site may assign 30 percent to useful efficiency. A research cluster with changing topology may give service and upgrade value more weight. Do not let a vendor choose the weights after showing its strengths.

Test the fabric like an AI system

A 72-hour packet generator test is necessary and insufficient. The acceptance plan should combine physical-layer stress, fabric behavior, and a representative distributed workload.

First, baseline the complete power boundary. Measure switch input power, cooling allocation, and optical support equipment at idle and load. Report watts per delivered terabit at the actual reach. Include laser modules and management controllers.

Second, vary inlet temperature and traffic pattern. Uniform random traffic will not expose every hot spot. Run all-reduce and all-to-all patterns at the planned GPU count, message sizes, and oversubscription. Record tail latency, corrected errors, retransmissions, optical power margin, and thermal tuning power.

Third, inject failures. Remove a laser source, dirty or disconnect a fiber, disable a lane, reboot a switch, and force a cooling excursion within the supported range. Measure detection, reroute, job impact, and restoration. If redundancy is part of a ten-times-resiliency claim, its switchover belongs in the acceptance test.

Fourth, rehearse field work. Fiber density can turn a theoretically faster deployment into a labeling and cleaning problem. Time rack installation, inspection, cable replacement, laser replacement, board replacement, and configuration rollback with the technicians who will do the work.

Finally, run a costed pilot long enough to see intermittent faults. The output should be a small set of operating numbers: useful tokens or training steps per megawatt, link-flap-free job hours, mean repair time, and spare consumption. Those numbers can support a fleet decision. A claim about terabits per second cannot.

The next purchase gate is straightforward: require every finalist to run the same distributed job, at the same optical reach and inlet range, while the buyer measures power and recovers from injected failures. CPO should win that test before it wins the architecture diagram.