A vendor slide says "400 TOPS/W — 100× beyond GPUs." Is that a breakthrough, a
rounding convention, or a crime scene? By the end of this lesson you will be able to tell, and
quickly. TOPS/W is a fraction, which means there are exactly two ways to flatter it: inflate the
numerator (count operations generously) or deflate the denominator (count watts forgetfully) —
and photonic systems, with their unusually long tail of off-headline costs, offer more room for
forgetfulness than any hardware since the analog computers of the 1950s. None of this requires
bad faith; it only requires a boundary drawn where the drawing is easiest. Benchmarking is the
craft of redrawing that boundary around the whole machine, normalising for precision,
and then asking the question that no peak number answers: can the machine actually be fed? For
that last part we will reuse a tool you already own — the
Start with the defensible part. An
The mischief starts after that. "Effective" ops: some claims multiply the count
by a sparsity factor (ops that were skipped are ops that were "done"), or count a low-precision
op as several high-precision ones. Peak vs delivered:
Now the forgetful half of the fraction. A complete photonic accelerator draws power in six
places, and this course has priced every one of them: the laser (at wall-plug
efficiency — a 20%-efficient laser turns 1 W of needed light into 5 W at the socket, per the
| Line item | Power | In the brochure? |
|---|---|---|
| Light itself (optical MACs) | ≈ 0.8 W optical | always |
| Laser at 20% wall-plug | 4 W | sometimes |
| DACs + ADCs (128 × 25 mW) | 3.2 W | rarely |
| Drivers + TIAs | 2 W | rarely |
| Thermal tuning (4032 phases) | 10 W | almost never |
| Digital control + SRAM | 5 W | almost never |
Brochure arithmetic:
Even an honestly-measured peak means nothing if data cannot arrive fast enough — and here
photonics meets an unglamorous truth: its memory system is electronic. Place the accelerator on
a
One more axis, then the checklist. Photonic marketing loves latency ("inference at the speed of light") and photonic physics half-supports it: light does cross a mesh in picoseconds. But end-to-end latency includes DAC settling, detector integration, ADC conversion and the digital round trip — typically nanoseconds to microseconds — and quoted TOPS figures usually assume saturating batch sizes that are themselves a latency cost. Throughput at batch 512 and latency at batch 1 are different products; insist on knowing which one is being sold. Here is the full audit, compressed:
Six questions, thirty seconds each. Most extraordinary claims do not survive the second one.
The single most common omission in photonic-computing claims — common enough to deserve its own
box — is the conversion chain. The tell is a per-op energy quoted in attojoules or single-digit
femtojoules: that is the optical energy, and it is real, but the
Hardware marketing has always raced ahead of hardware. Supercomputing spent decades policing
the gap between peak FLOPS and LINPACK-delivered FLOPS; the megahertz wars sold clock frequency
as if it were performance until Pentium 4 made the fallacy expensive; and the first generation
of AI accelerators taught everyone that "INT8 TOPS" and "FP16 TOPS" differ by design choice,
not achievement. The photonic twist is only that the gap between the flattering boundary and the
honest one is wider — because lasers, converters and tuning are physically separate components
that can be left off a die photo with a clear conscience. The countermeasure is also old:
standard workloads, measured at the wall, reported with accuracy. MLPerf did it for digital
accelerators; photonics will have earned its seat at that table on the day a vendor publishes a
wall-plug MLPerf number — an event worth watching for, and, as the