Information Is Physical
Rolf Landauer spent his career at IBM repeating a three-word slogan until physics took it seriously:
"Information is physical." There is no such thing as a disembodied bit. Every 0 or 1
anyone has ever stored rides on some physical degree of freedom — a
ball in one of
two valleys, a patch of magnetisation, an ink mark, a neuron. And if bits are physical,
then the laws of physics apply to them: bits gravitate, bits take up space and — the point of this
lesson — bits carry entropy. Not "something a bit like entropy". The actual
thermodynamic quantity, measured in joules per kelvin, that appears in the second law. The exchange
rate is one of the most beautiful small formulas in science: one bit of missing information is
exactly k \ln 2 of entropy.
Two entropies, one formula
You have met two quantities called "entropy", born eighty years apart in different subjects. Claude
Shannon's information
entropy measures missing information about a message, in bits:
H \;=\; -\sum_i p_i \log_2 p_i \qquad \text{(bits)}.
Gibbs' statistical-mechanical entropy measures missing information about which
microstate
a physical system occupies, in joules per kelvin:
S \;=\; -k \sum_i p_i \ln p_i \qquad \text{(J/K)},
with k \approx 1.38 \times 10^{-23}\ \mathrm{J/K} Boltzmann's constant.
Stare at the two formulas: they are the same expression. The only differences are the
logarithm's base and a constant out front — and changing the base of a logarithm is
multiplying by a constant, since \ln x = \ln 2 \cdot \log_2 x. So for any
probability distribution whatsoever,
S \;=\; k \ln 2 \; \cdot \; H.
Thermodynamic entropy is Shannon entropy priced in physical units. One bit of uncertainty —
two equally likely alternatives — is worth
k \ln 2 \;\approx\; 9.57 \times 10^{-24}\ \mathrm{J/K}
of honest, second-law-obeying entropy. This is not an analogy or a pun on the word "entropy". Both
formulas count the same thing: how many equally likely possibilities are hiding behind what
you know. Shannon counts them in doublings (bits); Gibbs counts them in
e-foldings and multiplies by k for historical
reasons — because entropy was discovered through steam engines, via heat and temperature, a century
before anyone realised it was about information.
Seeing it in the wells
Go back to the double-well memory and think like a statistical mechanic: entropy is
k times the log of the accessible region of phase space.
Take a large ensemble of identical memory cells. If you know each cell's bit, every cell's
particle sits in a region you can point to — one well's worth of phase space. If you
don't know the bits, the ensemble is spread across both wells: the accessible
region is twice as large, and k \ln(2\Omega) = k\ln\Omega + k \ln 2 —
exactly one bit more entropy. The "which well?" uncertainty is not a metaphor for entropy. It
is entropy, sitting in the memory, as physical as the entropy of steam.
Worked conversions
Let's put numbers on the exchange rate. A register of n unknown,
independent bits has H = n bits of Shannon entropy, hence thermodynamic
entropy S = n\,k \ln 2. And when entropy is bought or sold at temperature
T, the going price of heat is T\,\Delta S —
so the natural energy scale of one bit at temperature T is
kT \ln 2:
kT \ln 2 \;\Big|_{T = 300\ \mathrm{K}} \;\approx\; 2.87 \times 10^{-21}\ \mathrm{J}
\;=\; 2.87\ \mathrm{zJ} \;\approx\; 0.018\ \mathrm{eV}.
Memorise that number — it is the protagonist of the next three lessons. Some conversions to calibrate
your intuition:
| information | entropy S | S in J/K |
| 1 bit unknown | k ln 2 | 9.57 × 10⁻²⁴ |
| 1 byte unknown | 8 k ln 2 | 7.66 × 10⁻²³ |
| 1 TB drive of unknown data | 8 × 10¹² k ln 2 | 7.66 × 10⁻¹¹ |
| melting 1 g of ice | ≈ 1.3 × 10²³ k ln 2 | ≈ 1.22 |
Notice the last row: the entropy released by melting a single gram of ice equals the entropy of about
10^{23} unknown bits — ten thousand full hard drives. Information-bearing
entropy is a fantastically small sliver of a warm object's total entropy. That is why nobody
noticed it for a century, and why it took a thought experiment with a single molecule (next lesson)
to force the issue.
The bridge, drawn
The formula S = k \ln 2 \cdot H holds for lopsided bits too. A bit that is
1 with probability p has Shannon entropy
H(p) = -p \log_2 p - (1-p)\log_2(1-p), peaking at one full bit when
p = \tfrac12 and vanishing when the outcome is certain. Its physical
entropy is the same curve, squashed by the factor \ln 2 \approx 0.693
into units of k:
In the late 1940s Shannon had his formula but no name for it. He considered "information"
(overloaded) and "uncertainty" (vague), and asked John von Neumann. Von Neumann's reply, as Shannon
told it: "Call it entropy, for two reasons. First, your function is already used in statistical
mechanics under that name. Second — and more important — nobody really knows what entropy is, so in a
debate you will always have the advantage." The story is probably polished (Shannon told it a
few different ways over the years), but the first reason was dead right, and this lesson is the
payoff: the shared name turned out to mark a shared identity. Von Neumann's joke had a half-life of
about a decade — by 1961 Landauer was busy showing precisely what entropy is, at least the
information-bearing kind.
It can, and it does — and that is real physics, not philosophy. Shannon entropy measures the
observer's missing information, so the very same memory register carries entropy
k \ln 2 per bit to Bob, who never saw the data, and zero
to Alice, who wrote it. Before you object that a physical quantity can't be subjective, note what the
difference buys: entropy determines extractable work, and Alice really can extract work
that Bob cannot — knowing which well the particle is in, she can gently carry it downhill in a way
Bob, forced to plan for both possibilities at once, has no protocol for. "Knowledge" here is nothing
mystical: it means Alice's own physical records are correlated with the memory, and that
correlation is a physical resource. Exactly how knowledge is cashed into work — at
kT \ln 2 per bit — is the next lesson's engine.
Where this is going
We now have the exchange rate: 1\ \text{bit} = k \ln 2 of entropy, worth
kT \ln 2 of heat at temperature T. Two
consequences follow, and they are the twin peaks of this module. If acquiring a bit of information
lowers your uncertainty about a system, you should be able to convert that bit into
kT \ln 2 of work — the
Szilard
engine does exactly this. And if erasing a bit discards
k \ln 2 of entropy from your memory, the second law must send it
somewhere else — that bill is
Landauer's
principle.