<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.2.0">Jekyll</generator><link href="https://www.jakobhansen.org/feed.xml" rel="self" type="application/atom+xml" /><link href="https://www.jakobhansen.org/" rel="alternate" type="text/html" /><updated>2026-08-05T20:01:05-07:00</updated><id>https://www.jakobhansen.org/feed.xml</id><title type="html">Jakob Hansen</title><subtitle></subtitle><entry><title type="html">Comma pumps, impossible musical figures, and holonomy</title><link href="https://www.jakobhansen.org/2026/08/05/comma-pumps/" rel="alternate" type="text/html" title="Comma pumps, impossible musical figures, and holonomy" /><published>2026-08-05T00:00:00-07:00</published><updated>2026-08-05T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2026/08/05/comma-pumps</id><content type="html" xml:base="https://www.jakobhansen.org/2026/08/05/comma-pumps/">&lt;p&gt;Tuning musical instruments is weird and beautiful and fascinating and
complicated and annoying. For as long as people have been thinking about how
music works, they’ve been wrestling with various incarnations and shadows of
something that feels like an impossibility theorem. One of the earliest such
results, which dates back to both ancient Greece and ancient China, is the fact
that \(2^{19}\) is almost but not quite equal to \(3^{12}\). The ratio \(3^{12} /
2^{19} = \frac{531441}{524288} \approx 1.01364\) is called the &lt;em&gt;Pythagorean
comma&lt;/em&gt;. It’s the frequency difference between a stack of 12 perfectly tuned
fifths (with frequency ratio \(\frac{3}{2}\)) and 7 perfect octaves (with
frequency ratio \(\frac{2}{1}\)). The Pythagorean comma is about 23.46 cents. If
you go to a piano and pluck out a sequence of twelve perfect fifths, you end up
at a note seven octaves above where you started, and you’ll hit all twelve notes
of the chromatic scale once along the way; this is the &lt;em&gt;circle of fifths&lt;/em&gt;. But
if these perfect fifths were tuned exactly, we would end up a Pythagorean comma
above the note we would expect. (And in fact, this is one reason why we have
enharmonic names for notes. In Pythagorean tuning, B# is not the same as C,
because B# is a stack of 12 perfect fifths above C.) Another way to measure this
frequency ratio is in cents, which is a logarithmic scale with 1200 cents per
octave, so that 100 cents is the size of an equal-tempered semitone.&lt;/p&gt;

&lt;h2 id=&quot;comma-pumps&quot;&gt;Comma pumps&lt;/h2&gt;

&lt;p&gt;One of the downstream effects of this inconsistency (noticed at least by the
17th century) is the existence of comma pumps. A comma pump is a chord
progression that forces the pitch center to shift by a comma if performed in
just intonation. If you perform a sequence of chords, with each one tuned
correctly, you’ll return to the original chord at the end but the pitches will
all be slightly higher or lower than in the original chord. Online sources don’t
usually do a great job of explaining how this works, either making it sound too
complex or mystical, or ignoring the things that make it work. The thing that
makes this forced (rather than just evidence that you’ve gone out of tune) is
&lt;em&gt;common-tone voice leading&lt;/em&gt;. If you play one chord followed by another, and
those chords have some notes in common, you should probably play the shared
notes at exactly the same pitch. If those chords have fixed tunings (e.g. major
triads are correctly tuned when the frequencies of their pitches form the ratio
\(4:5:6\)), the frequencies of all the remaining notes are determined by their
relationship with the shared notes. So if you have a sequence of chords where
each pair shares at least one note, and expect all the chords to be justly
tuned, you have no choice about where the tuning ends up at the end.&lt;/p&gt;

&lt;p&gt;As an example, let’s look at the chord progression Cmaj-Fmaj-Dmin-Gmaj-Cmaj
(I-IV-ii-V-I in C major). We’ll consider the ratio of the root of each chord to
our initial C as we move along the progression:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Fmaj (F-A-C) shares the pitch C with Cmaj (C-E-G), so the root of
Fmaj needs to be a perfect fifth below or a perfect fourth above C; that’s a
ratio of \(\frac{4}{3}\).&lt;/li&gt;
  &lt;li&gt;Dmin (D-F-A) shares both F and A with Fmaj, so its root
should be a minor third below F, a ratio of \(\frac{5}{6}\).&lt;/li&gt;
  &lt;li&gt;Gmaj (G-B-D) shares
the pitch D with Dmin, so its root needs to be a perfect fourth above D, another
\(\frac{4}{3}\).&lt;/li&gt;
  &lt;li&gt;Finally, Cmaj shares the pitch G with Gmaj, so its root needs to
be a perfect fifth below G, a ratio of \(\frac{2}{3}\).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Multiplying these all together we have \(\frac{4}{3} \cdot \frac{5}{6} \cdot
\frac{4}{3} \cdot \frac{2}{3} = \frac{2^5 \cdot 5}{2 \cdot 3^4} =
\frac{80}{81}\) (-21.51 cents). The C we came back to at the end of the
progression is slightly lower in pitch than the one we started with! The
difference here is called the &lt;em&gt;syntonic comma&lt;/em&gt;, usually defined as the
difference between four perfect fifths \(\left(\frac{3}{2}\right)^4\) and two
octaves plus a major third \(2^2 \cdot \frac{5}{4}\). Here’s what that sounds
like:&lt;/p&gt;

&lt;audio controls=&quot;&quot; src=&quot;/assets/commapump/syntonic_pump.wav&quot;&gt;
&lt;/audio&gt;

&lt;p&gt;The descent is not subtle (a longer progression would probably make it harder to
detect), but it’s hard to decide exactly where things go wrong. It’s like the
reverse of the Penrose staircase: instead of descending a continuous downhill
path and somehow returning back to the same spot, you follow a path that’s
supposed to loop back to the beginning and find yourself slightly lower than you
started.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/commapump/impossible_staircase.svg&quot; alt=&quot;&quot; class=&quot;center-image two&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;holonomy&quot;&gt;Holonomy&lt;/h2&gt;

&lt;p&gt;If you’re like me, impossible figures should make you
&lt;a href=&quot;https://www.iri.upc.edu/people/ros/StructuralTopology/ST17/st17-05-a2-ocr.pdf&quot;&gt;think&lt;/a&gt;
&lt;a href=&quot;https://math.stackexchange.com/questions/2382058/what-is-the-true-relationship-between-impossible-figures-and-cohomology&quot;&gt;about&lt;/a&gt;
&lt;a href=&quot;https://arxiv.org/abs/2507.01226&quot;&gt;algebraic&lt;/a&gt;
&lt;a href=&quot;https://arxiv.org/abs/2602.09313v1&quot;&gt;topology&lt;/a&gt;. If comma pumps are like musical
impossible figures, where’s the topology? Here’s one approach: given a set of
chords, we can form a voice leading graph where we have one vertex for each
chord, and an edge between any two chords that share a note. (Or, if you want to
be pickier, you can require them to share two notes.) Paths in this graph are
admissible chord progressions, and a cycle is a chord progression that returns
to the original chord. If we now assign each chord a tuning, we can annotate
each (oriented) edge by the frequency ratio between their roots induced by the
choice of tuning.&lt;sup id=&quot;fnref:0&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:0&quot; class=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; This is a gadget with many different names. You could call it
a &lt;a href=&quot;https://en.wikipedia.org/wiki/Gain_graph&quot;&gt;gain graph&lt;/a&gt; with gain group
\(\mathbb{Q}^\times_{+}\). Or you could think of this data as providing an element
of \(C^1(G;\mathbb{Q}^\times_+)\). Or, somewhat less precisely but more
evocatively, you could think of this as a flat connection for a discrete fiber
bundle with structure group \(\mathbb{Q}^\times_+\). You could even construct a
cellular sheaf from this data, although it turns out that’s less convenient to
work with.&lt;/p&gt;

&lt;p&gt;The point is that now every loop in the graph (i.e. any chord progression) has
an associated holonomy (i.e. a change in root pitch frequency) given by taking
the product of the root ratios for every edge it traverses. This gives us a map
\(\Phi: \pi_1(G) \to \mathbb{Q}^\times_+\); since the target is abelian this map
factors through \(H_1(G;\mathbb{Z})\) and is determined by its values there, so we
may as well just descend to homology.  Any chord progression that is not a comma
pump is in the kernel of this holonomy map. So \(H_1(G;\mathbb{Z}) / \ker \Phi\)
classifies the possible comma pumps for a given family of chords.&lt;/p&gt;

&lt;p&gt;This algebraic fact gives us some additional insight into how many comma pumps
can exist. Just intonation tuning is often performed in a “prime limit,” which
restricts the prime factors that can appear in the numerator and denominator of
a pitch ratio. 5-limit just intonation allows for factors of 2, 3, and 5. This
means that all frequency ratios involved live in a subgroup of
\(\mathbb{Q}^\times_+\), one isomorphic to \(\mathbb{Z}^3\), with one factor for
each prime. The mathematical tuning community, for its own whimsical reasons,
has decided to call this representation the group of
&lt;a href=&quot;https://en.xen.wiki/w/Monzo&quot;&gt;&lt;em&gt;monzos&lt;/em&gt;&lt;/a&gt;, and writes them as kets: for example,
\(\lvert-4\; 4\; {-1}\rangle\) represents \(\frac{81}{80}\) as a 5-limit monzo.
It’s also common to ignore octave differences, meaning we take the quotient
\(\mathbb{Q}^\times_+ / \langle 2 \rangle\), or just drop the first factor of
\(\mathbb{Z}^n\). Let’s call this monzo subgroup \(\Lambda\), since we can think of
it as a lattice.&lt;/p&gt;

&lt;p&gt;By construction, the map \(H_1(G;\mathbb{Z}) / \ker \Phi \to \Lambda\) is
injective. So there are at most \(\text{rank } \Lambda\) independent comma pumps.
A chord family tuned in 5-limit JI can have at most two independent comma pumps,
a family tuned in 7-limit JI can have at most three, and so on. Note that \(\Phi\)
is generally not surjective. If it were, it would mean that there is a single
chord progression that when followed, results in a root pitch that is, say, a
perfect fifth above the starting point. (The quotient \(\Lambda / \text{im }
\Phi\) can be thought of as a
&lt;a href=&quot;https://en.xen.wiki/w/Mathematical_theory_of_regular_temperaments&quot;&gt;&lt;em&gt;temperament&lt;/em&gt;&lt;/a&gt;:
it collapses the commas into unisons. This is one of the central pieces of
mathematical tuning theory.)&lt;/p&gt;

&lt;h2 id=&quot;geometric-realizations&quot;&gt;Geometric realizations&lt;/h2&gt;

&lt;p&gt;At least sometimes, we can find a nice family of 2-cells whose boundary spans
\(\ker \Phi\), giving us a cell complex \(X\) with 1-skeleton \(G\) such that
\(H_1(X;\mathbb{Z})\) is the space of comma pumping chord progressions. The
generating cells should be bounded by something like minimal comma-free chord
progressions. This is the case for the triads of the major diatonic scale.&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; They form a graph like this:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/commapump/diatonic_voice_leading_strip.svg&quot; alt=&quot;Diagram of the triads of the major scale, forming a Möbius strip&quot; class=&quot;center-image full-width&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Each of the triangles in this graph can be filled in (just multiply the ratios
around the boundary, taking into account orientation). Gluing the ends together
gives us a Möbius strip, which has \(H_1(X;\mathbb{Z}) = \mathbb{Z}\). So there is
at most one class of comma pumps within this family of chords. In fact, our
comma pump from earlier is here: it’s I-IV-ii-V-I in C major. Multiplying ratios
along the path gives us (up to octave equivalence and inversion) the syntonic
comma \(\frac{80}{81}\). Every comma pump progression you can generate with this
set of chords moves the root pitch by a syntonic comma (or a power thereof). For
instance, you can follow the circle of fifths-like progression
I-IV-vii°-iii-vi-ii-V-I. This progression never repeats a chord, but it pumps
\(\left(\frac{80}{81}\right)^2\). This is because it’s a cycle that follows the
boundary of the Möbius strip, which is twice a generator of \(H_1\).&lt;/p&gt;

&lt;p&gt;What if we look at all major and minor triads on the chromatic scale? To keep
the drawing simple, we’ll only allow voice leading between chords with two
shared notes.&lt;sup id=&quot;fnref:2&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2&quot; class=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt; (Note that we are assuming enharmonic equivalence here, which is again
not typically how just intonation theory works.)&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/commapump/diatonic_tonnetz_torus.svg&quot; alt=&quot;&quot; class=&quot;center-image full-width&quot; /&gt;&lt;/p&gt;

&lt;p&gt;This is the dual graph of Euler’s
&lt;a href=&quot;https://en.wikipedia.org/wiki/Tonnetz&quot;&gt;Tonnetz&lt;/a&gt;, a triangular lattice where
vertices correspond with notes and 2-simplices correspond with chords.  There
are two root intervals available (three if you count the unison between the
major and minor chord on the same root): a minor third and a major third. The
natural dual cell structure combined with enharmonic and octave equivalence makes this into a
torus. This again lets us conclude that there are at most two independent comma
pumps available. And there are two independent comma pumps:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Amaj-Amin-Fmaj-Fmin-C#maj-C#min-Amaj pumps \(\frac{128}{125} = \lvert 7\; 0 \;
{-3}\rangle\) (41.06 cents). This is the &lt;a href=&quot;https://en.wikipedia.org/wiki/Diesis&quot;&gt;(lesser)
diesis&lt;/a&gt; or augmented comma, the difference
between a stack of three major thirds and an octave.&lt;/li&gt;
  &lt;li&gt;Cmaj-Cmin-Ebmaj-Ebmin-F#maj-F#min-Amaj-Amin-Cmaj pumps \(\frac{648}{625} =
\lvert 3\; 4\; {-4} \rangle\) (62.57 cents), which is the &lt;a href=&quot;https://en.xen.wiki/w/648/625&quot;&gt;greater
diesis&lt;/a&gt; or diminished comma. This is the
difference between a stack of four minor thirds and an octave.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The monzos are linearly independent, and we can see on the diagram that the
corresponding comma pumps are independent generators of \(H_1\).&lt;/p&gt;

&lt;p&gt;These comma pumps are not too surprising—the chord roots move uniformly in
major or minor thirds. But where’s the syntonic comma? Well, \(\frac{81}{80} =
\frac{648}{625} / \frac{128}{125}\). (Or equivalently in cents, \(21.51 = 62.57 -
41.06\).) Because \(\Phi\) is a group homomorphism, we can construct a chord
progression that pumps the syntonic comma by addition in \(H_1\): subtract the
progression that pumps the augmented comma from the progression that pumps the
diminished comma. Then we can produce a simplified homologous cycle by homotopy
across 2-cells. One example progression would be
Cmaj-Emin-Gmaj-Bmin-Dmaj-Dmin-Fmaj-Amin-Cmaj.&lt;/p&gt;

&lt;h2 id=&quot;what-else&quot;&gt;What else?&lt;/h2&gt;

&lt;p&gt;There are a lot of other interesting questions to ask inside this framework;
maybe I’ll do some of them next:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;More exotic chord families.&lt;/strong&gt; For instance, what happens when you introduce
  chord families with harmonic sevenths (as in barbershop music)? Chords from
  symmetric scales like the diminished scale?&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Alternate chord tunings.&lt;/strong&gt; Performers may want to select tunings on the fly
  to avoid the possibility of a comma pump. We could model this by adding a
  complex mapping down onto our chord complex, whose fibers correspond with
  alternate tunings, with the central question being whether we can lift a cycle
  in the base complex to a holonomy-free one.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Inverse problems.&lt;/strong&gt; What must be true in order to produce a chord family
  with no comma pumps, or a family that has progressions pumping a specific
  interval?&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Torsion.&lt;/strong&gt; All the holonomy we see here generates a free subgroup of
  \(\mathbb{Q}^\times_+ / \langle 2 \rangle\). The quotient doesn’t introduce any
  torsion here, but if we lift to \(\mathbb{R}^\times_+ / \langle 2 \rangle\)
  there could be commas that when stacked equal an octave. The most likely way to
  get this is by tempering intervals to an EDO (equal division of the octave)
  scale that doesn’t temper out a particular comma.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:0&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;There’s a subtlety here: once we have assigned tunings to chords, there’s a
possibility that intervals fail to match between two chords that share more than
one note. We can either require that we assign consistent tunings, or remove
edges where the tunings are inconsistent between the chords. &lt;a href=&quot;#fnref:0&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;There’s a bit of fussiness around the diminished vii° chord. There isn’t a
standard just tuning of this chord, and to make this work out as a Möbius strip,
I’ve chosen 25:30:36 (two stacked pure minor thirds), which is definitely not
the tuning most people would reach for. If you use another tuning, the vii°-ii
transition isn’t allowed because the intervals don’t match, so the complex ends
up with a hole. This would normally introduce another generator to \(H_1\), but it
turns out that another more complex 2-cell shows up to kill that generator. &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;If we allow voice leading with only one shared note, every 2-cell in the
complex becomes a \(K_6\). Filling in all the higher-dimensional simplices gives a
space that is homotopy equivalent to a torus, so all the same conclusions still
hold, and the comma pumps are actually significantly simpler. &lt;a href=&quot;#fnref:2&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name></name></author><summary type="html">Tuning musical instruments is weird and beautiful and fascinating and complicated and annoying. For as long as people have been thinking about how music works, they’ve been wrestling with various incarnations and shadows of something that feels like an impossibility theorem. One of the earliest such results, which dates back to both ancient Greece and ancient China, is the fact that \(2^{19}\) is almost but not quite equal to \(3^{12}\). The ratio \(3^{12} / 2^{19} = \frac{531441}{524288} \approx 1.01364\) is called the Pythagorean comma. It’s the frequency difference between a stack of 12 perfectly tuned fifths (with frequency ratio \(\frac{3}{2}\)) and 7 perfect octaves (with frequency ratio \(\frac{2}{1}\)). The Pythagorean comma is about 23.46 cents. If you go to a piano and pluck out a sequence of twelve perfect fifths, you end up at a note seven octaves above where you started, and you’ll hit all twelve notes of the chromatic scale once along the way; this is the circle of fifths. But if these perfect fifths were tuned exactly, we would end up a Pythagorean comma above the note we would expect. (And in fact, this is one reason why we have enharmonic names for notes. In Pythagorean tuning, B# is not the same as C, because B# is a stack of 12 perfect fifths above C.) Another way to measure this frequency ratio is in cents, which is a logarithmic scale with 1200 cents per octave, so that 100 cents is the size of an equal-tempered semitone.</summary></entry><entry><title type="html">Covariance pooling works because it’s nonlinear</title><link href="https://www.jakobhansen.org/2026/05/25/covariance-pooling/" rel="alternate" type="text/html" title="Covariance pooling works because it’s nonlinear" /><published>2026-05-25T00:00:00-07:00</published><updated>2026-05-25T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2026/05/25/covariance-pooling</id><content type="html" xml:base="https://www.jakobhansen.org/2026/05/25/covariance-pooling/">&lt;p&gt;Researchers at Goodfire recently proposed &lt;a href=&quot;https://www.goodfire.ai/research/covariance-pooling&quot;&gt;a new method for probing transformer
internal activations&lt;/a&gt; that
they call covariance pooling. They frequently want to take advantage of the
structured internal representations a model has learned in order to make
predictions or inferences about data. When working with sequences, often the
level of granularity you want to analyze is different from the level of
granularity of the sequence. When working with language, you typically want to
ask questions about sentences, paragraphs, or documents rather than words or
tokens. When working with DNA sequences, you want to ask questions about genes
rather than individual base pairs. But the model’s internal representation of a
sequence of tokens is a sequence of vectors. Different sequences have different
lengths, and whatever method you use to probe these representations needs to
handle sequences of arbitrary length.&lt;/p&gt;

&lt;p&gt;The default approach either to use the representation of the last token in the
sequence, or average the representations of all tokens in the sequence (mean
pooling). Using the last token often works surprisingly well, likely because the
autoregressive training objective encourages the model to represent significant
amounts of information about a sequence in the representation of its final
token, but it’s often finicky and clearly suboptimal. Mean pooling can at least
in principle incorporate information from vectors across the entire sequence,
but it’s also a fairly dumb method.&lt;/p&gt;

&lt;p&gt;Dumb methods are sometimes good, because they’re both interpretable and unlikely
to overfit. But some things just aren’t represented cleanly in a way that these
simple pooling methods can capture. So people have been looking for more
sophisticated ways to probe across sequences. One reasonably popular option is
attention probing, where you use a linear probe to identify which tokens in a
sequence to aggregate. (Surprisingly, &lt;a href=&quot;https://blog.eleuther.ai/attention-probes/&quot;&gt;this doesn’t always do better than just
using the last token&lt;/a&gt;.)&lt;/p&gt;

&lt;p&gt;Covariance pooling is a little different: the idea is to use the second moment
of the activations in the sequence rather than the first moment. This is, in
effect, applying the map \(x \mapsto xx^T\) to the activations before mean
pooling. (They use a learned linear compression of the output, so it’s actually
\(x \mapsto Lx(Rx)^T\) for two learned projections \(L\) and \(R\).)&lt;/p&gt;

&lt;p&gt;There’s a fairly obvious question: what happens if you apply the feature map
after mean pooling? The covariance pooling is doing two things here: aggregating
information in a way that is sensitive to localized co-occurrence of features,
as well as applying a nonlinear feature map before training a linear probe. How
much of the benefit is due to the nonlinearity?&lt;/p&gt;

&lt;p&gt;To find out, I tried a few binary classification tasks on Gemma 3 270M&lt;sup id=&quot;fnref:0&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:0&quot; class=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;, focused
on settings with at least moderately long sequences (a couple hundred words).
These are:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://huggingface.co/datasets/stanfordnlp/imdb&quot;&gt;IMDB reviews&lt;/a&gt;: sentiment analysis on movie reviews&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://huggingface.co/datasets/TimSchopf/medical_abstracts&quot;&gt;Medical Abstracts&lt;/a&gt;: identify the type of disease discussed in an abstract from a medical paper&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://huggingface.co/datasets/zapsdcn/hyperpartisan_news&quot;&gt;Hyperpartisan News&lt;/a&gt;: identify whether a news article takes an extremely partisan viewpoint or is neutral&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://huggingface.co/datasets/google/civil_comments&quot;&gt;Civil Comments&lt;/a&gt;: toxicity detection in internet comments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here’s the accuracy of each probe. I didn’t put a lot of effort into
hyperparameter optimization&lt;sup id=&quot;fnref:a&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:a&quot; class=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;, so take these as rough lower bounds,
particularly for the fancier probe types.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Dataset&lt;/th&gt;
      &lt;th&gt;Mean pool&lt;/th&gt;
      &lt;th&gt;Covariance pool&lt;/th&gt;
      &lt;th&gt;Mean pool with quadratic feature map&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;IMDB&lt;/td&gt;
      &lt;td&gt;0.907±0.005&lt;/td&gt;
      &lt;td&gt;0.911±0.007&lt;/td&gt;
      &lt;td&gt;0.904±0.005&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Medical Abstracts&lt;/td&gt;
      &lt;td&gt;0.871±0.003&lt;/td&gt;
      &lt;td&gt;0.881±0.009&lt;/td&gt;
      &lt;td&gt;0.877±0.006&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hyperpartisan News&lt;/td&gt;
      &lt;td&gt;0.840±0.014&lt;/td&gt;
      &lt;td&gt;0.938±0.000&lt;/td&gt;
      &lt;td&gt;0.935±0.017&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Civil Comments&lt;/td&gt;
      &lt;td&gt;0.769±0.005&lt;/td&gt;
      &lt;td&gt;0.828±0.008&lt;/td&gt;
      &lt;td&gt;0.768±0.008&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;IMDB and Medical Abstracts are essentially a wash between the methods.  For
Hyperpartisan News both covariance pooling and quadratic feature maps did better
than plain mean pooling. This makes sense if models represent partisanship
linearly along a left-right axis: a linear classifier can’t bend the extremes
around into a horseshoe, but second-order feature maps can. Conversely, for
Civil Comments, quadratic feature maps are about as good as mean pooling, while
covariance pooling does better.&lt;/p&gt;

&lt;p&gt;So yes, at least sometimes, the improvements from covariance pooling come purely
from the nonlinear feature map and not to capturing local co-occurrence
information. Which shouldn’t be too surprising: nonlinear classifiers are more
powerful than linear ones. A major reason we use linear classifiers is because
they’re simple and interpretable. If you want the best performance possible,
it’s often better to just fine-tune a base model for that purpose rather than
try to intercept the concept somewhere in an LLM’s residual stream. But if you
want more insight into what the underlying model is doing, linear probes have
the advantage of exactly corresponding with directions in the residual stream.
This means you can easily compare the directions of two probes or convert them
to steering vectors. That’s harder with nonlinear probes&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;—their main
contribution to interpretability is just giving evidence that a concept is
represented in some way at some location in the model. Interesting, but less
readily extensible to other tasks.&lt;/p&gt;

&lt;p&gt;A related aside: any time you have a feature map, you’ve got to be thinking
about kernels, right? This is a simple feature map, but we could look at the
kernel it defines; it’s just the inner product of the output features. So if you
have sequences \(X\) and \(Y\), represented as \(D \times M\) and \(D \times N\)
matrices, it’s&lt;/p&gt;

\[k(X, Y) = \left\langle \frac{1}{M} X X^T, \frac{1}{N} Y Y^T \right\rangle = \frac{1}{MN} \text{tr}(XX^TYY^T) = \frac{1}{MN} \text{tr}(X^TYY^TX) = \frac{1}{MN}\|Y^TX\|_F^2.\]

&lt;p&gt;This takes \(O(MND)\) multiplications, compared with the \(O((M+N)D^2)\)
multiplications for the covariance pooled feature map (ignoring the
projections). When \(D\) is bigger than the sequence length, the kernel
evaluation is cheaper than the feature map. But, of course, with the kernel
approach you need to evaluate the kernel for every pair of sequences, rather
than a feature map once for each sequence. So situations where the kernel
approach is more efficient than the direct feature map are likely rare.&lt;/p&gt;

&lt;p&gt;There are much more sophisticated kernels for sets of vectors, like the
&lt;a href=&quot;https://people.cs.uchicago.edu/~risi/papers/bhatta.pdf&quot;&gt;Bhattacharyya kernel&lt;/a&gt;.
This also computes the covariance matrix of each set, along with the mean, then
computes the Bhattacharyya coefficient \(\int_{\mathbb{R}^D}
\sqrt{p_X(z)p_Y(z)}\,dz\) between multivariate Gaussian distributions with these
parameters. This boils down to a gnarly formula that involves determinants and
inverses of covariance matrices. This is probably a nicer similarity measure in
some sense, and it gives you an infinite dimensional feature map, but it also
seems like a giant pain. No wonder they invented transformers.&lt;/p&gt;

&lt;p&gt;It’s worth noting that even the Bhattacharyya kernel still doesn’t treat the
sequence as a sequence, since the order of the vectors doesn’t matter. There are
also &lt;a href=&quot;https://jmlr.org/papers/v20/16-314.html&quot;&gt;kernels that take order into
account&lt;/a&gt;, but they’re even more
obscure. If we’re interested in more high-powered LLM probing techniques, kernel
methods might be a place to look for inspiration.&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:0&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;I tried doing some DNA sequence probing tasks like in the original post,
but these were a giant pain to get working, and I’m not familiar enough with the
domain to be confident I’m not doing something extremely dumb. &lt;a href=&quot;#fnref:0&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:a&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Some specifics:&lt;/p&gt;
      &lt;ul&gt;
        &lt;li&gt;Activations are from layer 9 of Gemma 3 270M (residual stream dimension 640), scaled to average norm 1.&lt;/li&gt;
        &lt;li&gt;All datasets are subsampled to select long sequences and to balance labels. Training sets range in size from around 500 to 5000, and test sets range from 65 to around 3000 samples.&lt;/li&gt;
        &lt;li&gt;Probes were trained with independent initializations and data ordering from five different seeds.&lt;/li&gt;
        &lt;li&gt;Probes are trained for 100 epochs on all datasets.&lt;/li&gt;
        &lt;li&gt;Linear probes have a learning rate of 1e-1 (surprisingly high, this is one of the few hyperparameters I tweaked), the others 1e-3.&lt;/li&gt;
        &lt;li&gt;Nonlinear probes have weight decay of 1e-2 applied.&lt;/li&gt;
        &lt;li&gt;No reconstruction loss is applied for covariance pooling or quadratic feature maps.&lt;/li&gt;
        &lt;li&gt;The quadratic feature map is computed exactly the same way as the covariance pooled feature maps, just after mean pooling instead of beforehand.&lt;/li&gt;
        &lt;li&gt;For the covariance pooling and quadratic feature map, the intermediate dimension (i.e. the number of rows in \(L\) and \(R\)) is 64.&lt;/li&gt;
      &lt;/ul&gt;
      &lt;p&gt;&lt;a href=&quot;#fnref:a&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Though not impossible, see, e.g., &lt;a href=&quot;https://arxiv.org/abs/2502.03708&quot;&gt;Toward universal steering and
monitoring of AI models&lt;/a&gt; for an approach to
extracting steering vectors from (a specific class of) nonlinear probes. &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name></name></author><summary type="html">Researchers at Goodfire recently proposed a new method for probing transformer internal activations that they call covariance pooling. They frequently want to take advantage of the structured internal representations a model has learned in order to make predictions or inferences about data. When working with sequences, often the level of granularity you want to analyze is different from the level of granularity of the sequence. When working with language, you typically want to ask questions about sentences, paragraphs, or documents rather than words or tokens. When working with DNA sequences, you want to ask questions about genes rather than individual base pairs. But the model’s internal representation of a sequence of tokens is a sequence of vectors. Different sequences have different lengths, and whatever method you use to probe these representations needs to handle sequences of arbitrary length.</summary></entry><entry><title type="html">When you shouldn’t trust einsum</title><link href="https://www.jakobhansen.org/2025/10/15/einsum/" rel="alternate" type="text/html" title="When you shouldn’t trust einsum" /><published>2025-10-15T00:00:00-07:00</published><updated>2025-10-15T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2025/10/15/einsum</id><content type="html" xml:base="https://www.jakobhansen.org/2025/10/15/einsum/">&lt;p&gt;When building neural networks, you’re often working with a bunch of operations that are basically matrix multiplication but with inputs that aren’t arranged or oriented in exactly the right way. One way to deal with this is to pull up all the various tensor shape manipulation operations available in a library like PyTorch, and fiddle with all the dimensions of your tensors until the operation you want is a matrix multiplication. This is a pain, and no one likes writing or reading code like&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;I&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;J&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;K&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;L&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;shape&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;permute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;([&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;flatten&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;L&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;P&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;shape&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;transpose&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;flatten&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;z&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;@&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;z&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reshape&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;J&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;K&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;L&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;P&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;If the author is thoughtful, these operations will at least be commented to annotate the expected tensor shapes at each point along the way. And there are libraries like &lt;a href=&quot;https://github.com/patrick-kidger/jaxtyping&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;jaxtyping&lt;/code&gt;&lt;/a&gt; that let you record that information in a somewhat more formal and verifiable way. But it’s still messy and annoying.&lt;/p&gt;

&lt;p&gt;So a lot of people prefer to use &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;torch.einsum()&lt;/code&gt;, or even more powerful libraries like &lt;a href=&quot;https://einops.rocks&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einops&lt;/code&gt;&lt;/a&gt; to write these sorts of transformations. With &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einsum()&lt;/code&gt;, the whole mess above is just&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;z&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;torch&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;einsum&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;ijkl, mip -&amp;gt; jklmp&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This is much nicer and (so I thought) probably better optimized than whatever index shuffling code I might happen to write.&lt;/p&gt;

&lt;p&gt;But there are situations where you might try a different approach to computing a matrix multiplication. I recently was using a product of the form&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;z&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;torch&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;einsum&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;ijk, ikl -&amp;gt; jl&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;In my case, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x&lt;/code&gt; might have a shape like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(8, 16384, 2048)&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;y&lt;/code&gt; might have a shape like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(8, 2048, 1024)&lt;/code&gt;. One way to implement this is to convert it to a strict matmul:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;transpose&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;flatten&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;flatten&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;z&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;@&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;But another way is to treat it as a batched matmul and then reduce over the first index:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;z&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;@&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;nb&quot;&gt;sum&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dim&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;One might think that &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einsum()&lt;/code&gt; would consider these (and other) options and select whichever is the most efficient. This is not the case. PyTorch’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einsum()&lt;/code&gt; implementation tries to push as much work onto the GEMM kernel as possible, so it &lt;a href=&quot;https://github.com/pytorch/pytorch/blob/v2.9.0/aten/src/ATen/native/Linear.cpp#L262&quot;&gt;always works as follows&lt;/a&gt;:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Permute dimensions of the two tensors so that the contracting indices are at the end for the first tensor and the beginning for the second tensor.&lt;/li&gt;
  &lt;li&gt;Flatten both tensors into two dimensions.&lt;/li&gt;
  &lt;li&gt;Run a GEMM.&lt;/li&gt;
  &lt;li&gt;Reshape and permute the resulting tensor to have the expected output shape.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;There’s a reasonable argument that this is the most efficient way to compute a tensor contraction in general. It avoids allocating any intermediate buffers that would otherwise need to be constructed in a multi-step reduction, and GEMM kernels are very well optimized for basically any backend.&lt;/p&gt;

&lt;p&gt;There’s a problem, though: if the input tensors are not contiguous, flattening them can require a copy. For some input shapes, that copy can be larger than the intermediate buffer that might be required. (For instance, if &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x&lt;/code&gt; had shape &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(2, 128, 1024)&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;y&lt;/code&gt; had shape &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(2, 1024, 128)&lt;/code&gt; the intermediate tensor would be &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(2, 128, 128)&lt;/code&gt;, much smaller than a copy of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x&lt;/code&gt;.)&lt;/p&gt;

&lt;p&gt;In my case, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x&lt;/code&gt; was actually a slice &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;X[:8,]&lt;/code&gt; along the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i&lt;/code&gt; dimension from a larger tensor &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;X&lt;/code&gt;, and thus not contiguous. This might have been okay, since even then the copy of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x&lt;/code&gt; was smaller than the intermediate tensor. But this was happening in an autograd context. This meant that the copy of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x&lt;/code&gt; needed to stick around until the backwards pass was finished, so that its gradient could be populated and then propagated back to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;X&lt;/code&gt;. This doesn’t happen when you use the second method: the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sum()&lt;/code&gt; call doesn’t need to save anything for the backward pass, and the batched matmul only saves a view of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;X&lt;/code&gt; for the backward pass, consuming no extra memory. As the cherry on top, this &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einsum()&lt;/code&gt; operation was actually happening for a bunch of different overlapping slices of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;X&lt;/code&gt;, making a separate copy of each. Overall, this was adding tens of gigabytes of memory consumption to a forwards + backwards pass of my model. Switching to the two-step reduction drastically reduced this memory consumption.&lt;/p&gt;

&lt;p&gt;There are two morals to this story. One is that every abstraction is leaky, and you never know when you’ll need to understand what’s beneath it. In 95% of cases &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einsum()&lt;/code&gt; is smarter than me, and I never bothered to look beneath the hood. That meant I had no idea what might be going on. The second is that if you care about performance, you need to look at performance profiles. The path to my realizing what was going here on stared by wondering why every &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;einsum()&lt;/code&gt; call in my execution trace had a child &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aten::clone()&lt;/code&gt;. PyTorch has pretty good memory and execution tracing tools; if you need your GPU to go fast, &lt;em&gt;use them&lt;/em&gt;.&lt;/p&gt;</content><author><name></name></author><summary type="html">When building neural networks, you’re often working with a bunch of operations that are basically matrix multiplication but with inputs that aren’t arranged or oriented in exactly the right way. One way to deal with this is to pull up all the various tensor shape manipulation operations available in a library like PyTorch, and fiddle with all the dimensions of your tensors until the operation you want is a matrix multiplication. This is a pain, and no one likes writing or reading code like</summary></entry><entry><title type="html">Conditions for the existence of approximations to the constant sheaf</title><link href="https://www.jakobhansen.org/2025/06/12/approximations-to-the-constant-sheaf/" rel="alternate" type="text/html" title="Conditions for the existence of approximations to the constant sheaf" /><published>2025-06-12T00:00:00-07:00</published><updated>2025-06-12T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2025/06/12/approximations-to-the-constant-sheaf</id><content type="html" xml:base="https://www.jakobhansen.org/2025/06/12/approximations-to-the-constant-sheaf/">&lt;p&gt;This is a result I proved a while back as part of a larger project that never
went anywhere. It turns out there might be people interested in the result for
its own sake, so this seems a reasonable place to get it written down.&lt;/p&gt;

&lt;p&gt;In &lt;a href=&quot;https://www.jakobhansen.org/publications/thesis.pdf&quot;&gt;my thesis&lt;/a&gt;, I
introduced the idea of an approximation to a cellular sheaf, and gave some
randomized algorithms for constructing spectrally good approximations based on
the literature for graph sparsification. While these algorithms could give
bounds on the total dimension of the stalks of the approximating sheaf, they
didn’t have any control over what happened locally—there might be critical
parts of the complex where stalks are forced to stay the same dimension. So I
was interested in anything that might help better understand the constraints on
the stalk dimensions of a sheaf approximation, particularly in the case of
approximations to constant sheaves.&lt;/p&gt;

&lt;p&gt;For completeness, I’ll give a definition here: a \(k\)-approximation to the
constant sheaf \(\underline{V}\) on a complex \(X\) is a sheaf \(\Fc\) together
with a sheaf morphism \(a: \underline{V} \to \Fc\) for which the following hold:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;\(a\) is an isomorphism on stalks over cells of dimension \(\leq k\)&lt;/li&gt;
  &lt;li&gt;\(a\) is surjective on stalks over cells of dimension \(&amp;gt; k\)&lt;/li&gt;
  &lt;li&gt;\(H^k a\) is an isomorphism (note that condition 1 implies that \(H^i a\) is an isomorphism for \(i &amp;lt; k\)).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is most immediately interesting when \(k = 0\), so that an approximation to
the constant sheaf on a graph \(G\) has the same vertex stalks but may have
lower-dimensional edge stalks, while still having the same space of global
sections. The constant sheaf on a graph gives a local algorithm for verifying
the constancy of a \(0\)-cochain; this algorithm requires communication of a
\(\dim V\)-dimensional vector over each edge. An approximation to the constant
sheaf also provides such an algorithm, but potentially requiring less
communication per edge.&lt;/p&gt;

&lt;p&gt;If you choose the natural basis, the \(0\)-Laplacian of an approximation to the
constant sheaf \(\underline{\R^k}\) on a graph (or really any sheaf that admits
a map \(\underline{\R^k} \to \Fc\)) is a matrix-weighted graph Laplacian. This
is the obvious Laplacian matrix for a graph where each edge has a \(k \times k\)
symmetric positive semidefinite weight matrix—the off-diagonal blocks \(L_{ij}
= -W_{ij}\) are negatives of the corresponding weight matrices, and the diagonal
blocks \(L_{ii} = \sum_{i \sim j} W_{ij}\) are sums of the weight matrices for
edges incident to a node. These types of Laplacians have been studied in control
theory settings for reasons that were always a little opaque to me. Like a lot
of academic engineering literature, it’s hard for me to tell if the applications
described are genuine or just descriptions of conceivable-but-not-realistic
scenarios where precisely this mathematical object would be called for. To be
clear, I’m not one to judge here—while I describe directions for applicability
of sheaves, I’ve seldom felt that I had a clear-cut application that wouldn’t be
possible without sheaf theory. Actually, this result is a little nice in that
regard, since the sheaf language does help prove something about matrix-weighted
graphs that I don’t see a nice way to prove otherwise.&lt;/p&gt;

&lt;p&gt;Anyway, the setup here is that we’ve chosen a dimension \(k\) for the stalks of
the constant sheaf as well as a fixed graph (or cell complex). We’ll also choose
a &lt;em&gt;dimension vector&lt;/em&gt; \(d\) defining the dimensions of the stalks of \(\Fc\) over
each edge \(e\). The question is for which dimension vectors it is possible to
construct an approximation to the constant sheaf \(\underline{\R^k}\). Here’s the answer:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proposition.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Let \(G\) be a graph, and let \(d\) be a dimension vector for the edges of
\(G\). The following are equivalent:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;There exists a 0-approximation to the constant sheaf \(a: \underline{\R^k} \to \mathcal F\) on \(G\) with dimension vector \(d\)&lt;/li&gt;
  &lt;li&gt;For every partition \(\mathcal P\) of the vertices of \(G\),
\[\sum_{e \in E(\mathcal P)} d_e \geq k(\left| {\mathcal P}\right| - 1)\]&lt;/li&gt;
  &lt;li&gt;There exists a collection of \(k\) (not necessarily distinct) spanning trees of \(G\) such that each edge \(e\) of \(G\) is contained in at most \(d_e\) trees in the collection.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Proof.&lt;/strong&gt; We’ll show (1) \(\Rightarrow\) (2) by counting dimensions. Consider the
quotient graph \(G / \mathcal P\) which has one node for each set 
\(U \in \mathcal P\) and an edge between \(U\) and \(V\) if there exist 
\(u \in U, v \in V\) such that there is an edge between \(u\) and \(v\) in \(G\). 
We can take the pushforward \(p_* \underline{\Fc}\) over the quotient map \(p: G \to G / \mathcal P\).
For \(U \in V(G / \mathcal P)\), we have \(p_*\underline{\Fc}(U) = H^0(U ; \underline{\Fc})\). This has dimension at least \(k\), so 
\(C^0(G/\mathcal P; p_*\Fc) \geq \abs{\mathcal P}k\).
Meanwhile, the edge stalks of \(p_*\underline{\Fc}\) are just direct sums of edge stalks of \(\Fc\), so 
\(C^1(G/\mathcal P; p_*\Fc) = \sum_{e \in E(\mathcal P)} d_e\). 
The pushforward of a sheaf preserves its space of global sections, so \(\dim H^0(G/\mathcal P;p_* \Fc) = k\).
Since \(H^0 = \ker \delta\), we must have \(\dim H^0 \geq \dim C^0 - dim C^1\), so&lt;/p&gt;

&lt;p&gt;\[\abs{\mathcal P} k - \sum_{e \in E(\mathcal P)} d_e \leq \dim C^0(G/\mathcal P; p_*\Fc) - \dim C^1(G/\mathcal P;p_*\Fc) \leq k.\]&lt;/p&gt;

&lt;p&gt;We can use a tree-packing result proved independently by Tutte&lt;sup id=&quot;fnref:tutte&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:tutte&quot; class=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; and
Nash-Williams&lt;sup id=&quot;fnref:nash-williams&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:nash-williams&quot; class=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; for the implication (2) \(\Rightarrow\) (3). 
They showed that there exists a set of \(k\) edge-disjoint
spanning trees in a multigraph \(G\) if and only if for every partition
\(\mathcal P\) of \(V(G)\), \(E(\mathcal P) \geq k(\abs{\mathcal P} - 1)\). Replacing
each edge of \(G\) with \(d_e\) edges yields an equivalence with condition (2).&lt;/p&gt;

&lt;p&gt;Finally, (3) \(\Rightarrow\) (1) by considering for each spanning tree \(T_j\)
the inclusion map \(\iota: T_i \to G\) and pushing the constant sheaf
\(\underline{\R}\) on \(T_j\) forward to get a collection of \(k\) sheaves
\(\Fc_j = \iota_* \underline{\R}\) on \(G\). Then
each \(\Fc_j\) is an approximation to \(\underline{\R}\) on \(G\), so that 
\(\Fc = \bigoplus_j \Fc_j\) is an approximation to \(\underline{\R^k} = \underline{\R}^{\oplus k}\) 
on \(G\). The requirement on the spanning tree packing
ensures that the dimension vector of \(\Fc\) is bounded above by \(d\); taking a direct
sum with a sheaf supported only on edges can then make the dimension vector
exactly \(d\). \(\:\:\blacksquare\)&lt;/p&gt;

&lt;p&gt;A matroid-theoretic generalization of the Tutte/Nash-Williams theorem proven by
Edmonds&lt;sup id=&quot;fnref:edmonds&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:edmonds&quot; class=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt; allows us to generalize this to higher-dimensional complexes,
although I’m not sure if anyone cares:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Let \(X\) be an \(n\)-dimensional regular cell complex, and let \(d\) be a dimension
vector on \(X\) with \(d_\sigma = k\) for \(\dim(\sigma) &amp;lt; n\). The following are equivalent:&lt;/p&gt;
  &lt;ol&gt;
    &lt;li&gt;There exists an \((n-1)\)-approximation \(a: \underline{\R^k} \to \Fc\) on \(X\)
with dimension vector \(d\).&lt;/li&gt;
    &lt;li&gt;For every subcomplex \(Q\) of \(X\) with \(Q^{n-1} = X^{n-1}\), 
\[\sum_{\substack{\dim \sigma = n \ \sigma \in X \setminus Q}}d_\sigma
 \geq k \left( \dim {H}^{n-1}(Q;\R) - \dim H^{n-1}(X;\R) \right)\]&lt;/li&gt;
    &lt;li&gt;There exists a collection of \(k\) spanning \(n\)-acycles of \(X\), such that for
each \(n\)-cell \(\sigma\) of \(X\), \(\sigma\) is contained in at most \(d_\sigma\)
of the acycles.&lt;/li&gt;
  &lt;/ol&gt;
&lt;/blockquote&gt;

&lt;p&gt;But the people who might actually be interested in this are concerned with
matrix-weighted graphs, so I’ll translate it into that language:&lt;/p&gt;

&lt;p&gt;Let \(G\) be a graph, and suppose we want to assign \(k \times k\) weight matrices \(W_{ij}\) to produce a matrix-weighted Laplacian \(L\). Then the following are equivalent:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;There exists a choice of weight matrices \(W_{ij}\) with \(\rank(W_{ij}) = r_{ij}\) such that \(\dim \ker L = k\)&lt;/li&gt;
  &lt;li&gt;For every partition \(\mathcal P\) of the vertices of \(G\),
\[\sum_{i\sim j | \mathcal P_i \neq \mathcal P_j} r_{ij} \geq k(\left| {\mathcal P}\right| - 1)\]&lt;/li&gt;
  &lt;li&gt;There exists a collection of \(k\) spanning trees of \(G\) such that each edge \(i \sim j\) is contained in at most \(r_{ij}\) trees.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So what does this mean? In one sense, this theorem might be kind of
disappointing. It’s not hard to show that the spanning tree condition implies
that an approximation to the constant sheaf is possible, since it basically
ignores all of the extra degrees of freedom we get from being able to choose
bases other than the standard basis for \(\R^k\). You might hope that by
cleverly choosing subspaces to project onto, it might be possible to get by with
lower-dimensional edge stalks (or lower-rank weight matrices) than you might if
you restricted yourself to axis-aligned projections. But this isn’t the case:
the obvious sufficient condition turns out to also be necessary.&lt;/p&gt;

&lt;p&gt;Now, there’s still a reasonable hope that even if varying the basis doesn’t buy
you anything combinatorially, it might still buy you something spectrally: for a
given dimension vector, it might be that it’s possible to get a bigger spectral
gap with general matrix weights than with axis-aligned projections. This is
similar to the argument I give in &lt;a href=&quot;https://www.jakobhansen.org/publications/mwgraphs.pdf&quot;&gt;Expansion in Matrix-Weighted
Graphs&lt;/a&gt; that
better-than-Ramanujan matrix-weighted expander graphs might exist.&lt;/p&gt;

&lt;p&gt;I also have a hunch that existence of an approximation to the constant sheaf is
a generic condition, so that if it’s possible to construct one, it’s also easy.
At one point I convinced myself that I had proven this, but I can’t manage to
find anywhere I wrote this down. The general idea, though, is that unless some
sort of catastrophic alignment of constraints happens, matrices will generically
have the maximum rank possible.&lt;/p&gt;

&lt;p&gt;What I like most about this result is that it has an essential and nontrivial
use of sheaf theory. Using the framework of approximations to the constant sheaf
made proving (1) \(\Rightarrow\) (2) straightforward. I’m not sure there’s a
nice way to prove the equivalent statement working purely in terms of matrices.
I’ll take that as a win for algebraic topology.&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:tutte&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;W. T. Tutte, “&lt;a href=&quot;https://londmathsoc.onlinelibrary.wiley.com/doi/abs/10.1112/jlms/s1-36.1.221&quot;&gt;On the Problem of Decomposing a Graph into n Connected Factors&lt;/a&gt;,” Journal of the London Mathematical Society, vol. s1-36, no. 1, pp. 221–230, 1961, doi: 10.1112/jlms/s1-36.1.221. &lt;a href=&quot;#fnref:tutte&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:nash-williams&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;C. St. J. A. Nash-Williams, “&lt;a href=&quot;https://academic.oup.com/jlms/article-abstract/s1-36/1/445/828412?login=false&quot;&gt;Edge-Disjoint Spanning Trees of Finite Graphs&lt;/a&gt;,” Journal of the London Mathematical Society, vol. s1-36, no. 1, pp. 445–450, Jan. 1961, doi: 10.1112/jlms/s1-36.1.445. &lt;a href=&quot;#fnref:nash-williams&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:edmonds&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;J. Edmonds, “&lt;a href=&quot;https://nvlpubs.nist.gov/nistpubs/jres/69B/jresv69Bn1-2p73_A1b.pdf&quot;&gt;Lehman’s switching game and a theorem of Tutte and Nash-Williams&lt;/a&gt;,”  Journal of Research of the National Bureau of Standards-B. Mathematics and Mathematical Physics, vol. 69B, no. 1 and 2, p. 73, Jan. 1965, doi: 10.6028/jres.069B.005. &lt;a href=&quot;#fnref:edmonds&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name></name></author><summary type="html">This is a result I proved a while back as part of a larger project that never went anywhere. It turns out there might be people interested in the result for its own sake, so this seems a reasonable place to get it written down.</summary></entry><entry><title type="html">Representational limitations of knowledge graph embeddings</title><link href="https://www.jakobhansen.org/2025/05/08/knowledge-graph-embeddings/" rel="alternate" type="text/html" title="Representational limitations of knowledge graph embeddings" /><published>2025-05-08T00:00:00-07:00</published><updated>2025-05-08T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2025/05/08/knowledge-graph-embeddings</id><content type="html" xml:base="https://www.jakobhansen.org/2025/05/08/knowledge-graph-embeddings/">&lt;p&gt;Knowledge graph embeddings are probably not that interesting these days. It was never quite clear to me exactly what they were supposed to accomplish (something to do with automated fuzzy reasoning?) but it seems like most of those use cases have been superseded by LLMs. But when I was working on knowledge graph embeddings for a minute, this was something that bugged me.&lt;/p&gt;

&lt;p&gt;The idea behind knowledge graph embeddings is pretty well encapsulated by the whole word2vec story about analogical reasoning. You know the deal, \(\text{king} - \text{man} + \text{woman} \approx \text{queen}\).&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; The idea is that the vector \(\text{woman} - \text{man}\) approximately represents some semantic relationship like “is the female equivalent of.” You can make this idea much more explicit, and learn embeddings that are meant to directly encode some set of known relationships between entities. This means you start with some knowledge graph—a set of entities \(\mathcal{E}\), a set of relations \(\mathcal{R}\), and a set \(\mathcal{T}\) of factual triplets \((x, r, y)\), with \(x, y \in \mathcal{E}\) and \(r \in \mathcal{R}\). If \((x, r, y) \in \mathcal{T}\) we know that the relation \(r\) holds between entities \(x\) and \(y\). An equivalent way to frame this, probably a bit more familiar to mathematicians, is to think of \(\mathcal{R}\) as literally a collection of relations on the set \(\mathcal{E}\).&lt;/p&gt;

&lt;p&gt;There are maybe a surprising number of knowledge graphs that you could start from. The Semantic Web was once the future of knowledge. All of the information in the world would be linked together with URIs and RDF triplets! You could have a software agent that would query the World Wide Web for you and find answers to complicated questions!&lt;sup id=&quot;fnref:2&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2&quot; class=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;  Anyway, there’s still a good amount of this kind of information sitting around. It’s the quintessence of the circa-2010 Internet, and still appeals to a certain sort of person. Why limit yourself to writing Wikipedia articles when you can convert them into precisely formulated factual statements on &lt;a href=&quot;https://www.wikidata.org/wiki/Wikidata:Main_Page&quot;&gt;Wikidata&lt;/a&gt;?&lt;/p&gt;

&lt;p&gt;Anyway, the point of view of a knowledge graph as a set of relations on the entity set naturally raises the question of what properties those relations satisfy. Some of these sorts of properties are immediately familiar to anyone who took a sufficiently introductory class on naive set theory:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;reflexivity&lt;/li&gt;
  &lt;li&gt;(anti)symmetry&lt;/li&gt;
  &lt;li&gt;transitivity&lt;/li&gt;
  &lt;li&gt;(co)injectivity&lt;sup id=&quot;fnref:3&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3&quot; class=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A lot of the relations we might want to model in a knowledge graph naturally satisfy some of these properties. This is even potentially useful if we have an incomplete knowledge graph; properties like transitivity can help us infer relations that we don’t have a direct record of. For instance:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;The relation “is a friend of” is (hopefully?) symmetric&lt;/li&gt;
  &lt;li&gt;The relation “is a subordinate of” is antisymmetric&lt;/li&gt;
  &lt;li&gt;The relation “is a type of” is transitive&lt;/li&gt;
  &lt;li&gt;The relation “is an author of” is not injective (books can have more than one author) or coinjective (authors can write more than one book)&lt;/li&gt;
  &lt;li&gt;The relation “is married to” is (usually) injective and coinjective (i.e. it defines a bijection on its domain)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But it’s actually surprisingly tricky to come up with knowledge graph embedding schemes where the representations are clearly able to satisfy these kinds of properties. Let’s look at &lt;a href=&quot;https://papers.nips.cc/paper_files/paper/2013/hash/1cecc7a77928ca8133fa24680a88d2f9-Abstract.html&quot;&gt;TransE&lt;/a&gt;,  the embedding method that corresponds with the word2vec analogy arithmetic. Entities and relations are all represented as vectors in the same space, and we want the existence of a triple \((x,r,y)\) to imply that \(v_x + v_r \approx v_y\). Let’s make that \(\approx\) an \(=\) for now and think about what properties a relation encoded this way can exhibit.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The only reflexive relation represented in this way the identity relation.&lt;/li&gt;
  &lt;li&gt;The only symmetric relation represented in this way is also the identity.&lt;/li&gt;
  &lt;li&gt;Every relation represented in this way is strictly antisymmetric (except the identity).&lt;/li&gt;
  &lt;li&gt;The only transitive relation represented in this way is the identity.&lt;/li&gt;
  &lt;li&gt;Every relation represented in this way is injective and coinjective.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I don’t think most of these constraints are things that can be swept under the rug by allowing the equality to be only approximate. It doesn’t help at all with symmetry or transitivity. You can kind of work around the enforced injectivity injectivity by clustering vectors representing different tail entities \(y'\) around \(v_x + v_r\). This works when you only have one relation to care about. But say we want to encode “\(x\) is an author of \(y\)” for a bunch of books, and also “\(y\) has a topic of \(z\)” for each of these books. Now every book this author writes has to be about the same topic (or cluster of topics)!&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://ojs.aaai.org/index.php/AAAI/article/view/8870&quot;&gt;TransH&lt;/a&gt; is another method that was designed to solve the problem that TransE can only represent bijective relations. The idea is to choose a subspace \(V_r\) for each relation \(r\) and project onto \(V_r\) before checking the approximate equality: \(P_{V_r} (v_x + v_r) \approx P_{V_r} v_y\). This is a lot nicer, and has a natural way of dealing with one-to-many and many-to-many relations. But it’s still limited in weird ways. If you want a reflexive, transitive, or symmetric relation, you need to set \(v_r = 0\), so any relation satisfying one of these properties satisfies all three and is in fact an equivalence relation.&lt;/p&gt;

&lt;p&gt;The main reason I’m interested in thinking about knowledge graph representations these days is to better understand how LLMs &lt;a href=&quot;https://arxiv.org/abs/2406.19501&quot;&gt;encode&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/abs/2407.14662&quot;&gt;relational&lt;/a&gt; &lt;a href=&quot;https://arxiv.org/abs/2308.09124&quot;&gt;knowledge&lt;/a&gt;. LLMs know facts like “Michael Jordan plays basketball”&lt;sup id=&quot;fnref:4&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:4&quot; class=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;, and can do precisely the sort of compositional reasoning with these facts that knowledge graph embeddings are supposed to enable. They can even do this with facts provided only in context and not in their training data. To do this, they’re working with fairly straightforward operations on vector representations of entities. There’s been a bit of work at trying to find what these representations look like, mostly using approaches that represent relations with individual matrices. The idea for these is that either \(A_r v_x \approx v_y\) or \(\langle A_r v_x, v_y \rangle \gg 0\) for true relations.&lt;/p&gt;

&lt;p&gt;The first approach has many of the same issues as TransE, since it’s still trying to represent a possibly one-to-many relation with a function. Matrix multiplication is significantly more expressive than vector addition, though—in particular it’s not always injective, so it can natively represent many-to-one relationships. Transitivity is still tricky, though. If we want \(A_r v_x \approx v_y\) and \(A_r v_y \approx v_z\) to mean \(A_r v_x \approx v_z\) we need \(A_r\) to be (approximately) idempotent, at least on a subspace that contains the embeddings of all relevant entities. But that means that it’s a projection matrix, and so everything except the first entity in a transitive chain of relations needs to have the same embedding. You run into similar issues if you want to represent reflexive relationships. There are more things you can do to increase expressivity—&lt;a href=&quot;https://arxiv.org/pdf/2110.03789&quot;&gt;my paper&lt;/a&gt; on this took things about as general as they could get while keeping everything linear—but the more expressive it gets, the easier it is to overfit as well.&lt;/p&gt;

&lt;p&gt;One lesson to take from this whole thing is that actually, linear representations kinda suck if you want to encode a large universe of concepts simultaneously. And this is why language models don’t operate linearly on static encodings of concepts. Internal token representations are contextual; it’s entirely possible that the existence of a relation between embeddings of two entities is only linearly detectable when the context implies that it may be important to extract that relation. For instance, when a model is prompted with “What is the capital of France?”, the embedding of “France” is different than when it’s prompted with “What is the GDP of France?”, and it may be that the representation of the “capital of” relation only works with the first. This throws a bit of a wrench in the idea that we might be able to easily decode all of the relations an LLM is aware of from its internal activations in a single context. So there’s definitely a lot more to be done to understand an LLM’s knowledge graph (whatever that actually means). I’m just not sure how much studying knowledge graph embeddings will actually help.&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;There’s some doubt about whether word2vec was actually learning anything nearly this sophisticated. It turns out that \(\text{man}\) and \(\text{woman}\) were already pretty close, as were \(\text{king}\) and \(\text{queen}\), so maybe all this equation was picking up on is that \(\text{king} \approx \text{queen}\). Is this maybe an early &lt;a href=&quot;https://arxiv.org/pdf/2104.07143&quot;&gt;interpretability illusion&lt;/a&gt;? &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Man, I miss the Internet of the early 2010s. The promise and vision of the future was so much cooler than any future we were plausibly going to end up with. &lt;a href=&quot;#fnref:2&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;OK, maybe “coinjective” isn’t a term most people learn in an introduction to proofs course. But it’s useful, just meaning that the converse relation is injective. &lt;a href=&quot;#fnref:3&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:4&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;This specific example is weirdly common in literature on the topic of what LLMs know and how they know it. &lt;a href=&quot;#fnref:4&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name></name></author><summary type="html">Knowledge graph embeddings are probably not that interesting these days. It was never quite clear to me exactly what they were supposed to accomplish (something to do with automated fuzzy reasoning?) but it seems like most of those use cases have been superseded by LLMs. But when I was working on knowledge graph embeddings for a minute, this was something that bugged me.</summary></entry><entry><title type="html">Something perplexing about UMAP</title><link href="https://www.jakobhansen.org/2025/03/18/umap-puzzle/" rel="alternate" type="text/html" title="Something perplexing about UMAP" /><published>2025-03-18T00:00:00-07:00</published><updated>2025-03-18T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2025/03/18/umap-puzzle</id><content type="html" xml:base="https://www.jakobhansen.org/2025/03/18/umap-puzzle/">&lt;p&gt;If you want to visualize some large high-dimensional dataset, odds are you’re
going to try UMAP. After its initial release in 2018, it quickly displaced t-SNE
(which had itself displaced, I don’t know, diffusion maps?) as the nonlinear
dimensionality reduction method everyone uses by default, particularly in
machine learning (for instance, dealing with embeddings or neural network
activations) and biology (e.g. for looking at single-cell gene expression
profiles).&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;There are a handful of parameters that UMAP users tend to adjust when creating
embeddings, with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; probably being the most common. &lt;a href=&quot;https://umap-learn.readthedocs.io/en/latest/parameters.html#n-neighbors&quot;&gt;The UMAP
documentation&lt;/a&gt;
highlights the need to explore different values of this parameter in order to
get good results. If &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; is too low, it’s possible for the resulting
embedding to look like a whole bunch of tiny clusters with no relationship to
each other. If it’s too high, computing the embedding can take a really long
time, and some of the local structure in the dataset can be lost.&lt;/p&gt;

&lt;p&gt;UMAP builds its embedding by building a nearest-neighbors graph on the points,
assigning weights to each edge in that graph based on a normalized version of
the distances between points, and then optimizes the embeddings in the target
space so that their weighted nearest-neighbors graph is as similar as possible
to that of the original data. (UMAP’s creator sometimes calls the graph a “fuzzy
simplicial set,” but there’s not really much lost by just thinking of it as a
weighted graph.) It makes sense, then, that the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; parameter has a
pretty big influence on the resulting embedding. If there are two high-density
regions of the data separated by a lower-density (or completely empty) region
and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; is low, it’s quite possible that all the nearest neighbors of
every point in each region lie inside the same region, meaning that the
neighborhood graph becomes disconnected, and the embedding algorithm has no idea
where to place the clusters in relation to each other. Increasing &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt;
makes this less likely to happen.&lt;/p&gt;

&lt;h2 id=&quot;a-puzzle&quot;&gt;A puzzle&lt;/h2&gt;

&lt;p&gt;So far so good. But let’s look a little closer using some synthetic data. We’ll
arrange a bunch of clusters along a line, and vary the number of points in each
cluster. As long as the number of points in each cluster is less than
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt;, there should be at least some information about the relationship
between clusters in the neighborhood graph, and we should therefore see it in
the embedding. But here’s what happens:
&lt;img src=&quot;/assets/umap_puzzle/default_clusters.png&quot; alt=&quot;UMAP behavior with n_neighbors = 15&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
The information about the position of the clusters on the line starts to
disappear around cluster size 5!&lt;/p&gt;

&lt;p&gt;These embeddings used the default value of 15 for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt;. What happens if
we increase it to 30? 
&lt;img src=&quot;/assets/umap_puzzle/more_nbrs_clusters.png&quot; alt=&quot;UMAP behavior with n_neighbors = 30&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
Hmm, looks like the larger-scale information sticks around a little longer, but
still breaks down by cluster size 7. Maybe if we double &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; again?
&lt;img src=&quot;/assets/umap_puzzle/even_more_nbrs_clusters.png&quot; alt=&quot;UMAP behavior with n_neighbors = 60&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
It seems that doubling &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; only increases the number of data points we
can handle per cluster by 1.&lt;/p&gt;

&lt;h2 id=&quot;whats-going-on-here&quot;&gt;What’s going on here?&lt;/h2&gt;

&lt;p&gt;It turns out that this is a consequence of the way that UMAP normalizes
distances in the graph. UMAP wants to assign a weight \(w_{ij} = \exp(-d_{ij})\)
to each edge in the graph. However, since the raw distances in a point cloud
could be very different between points, it’s important to do some density
correction. First, we’ll always assign the edge from \(i\) to its nearest
neighbor weight 1. This means subtracting the distance to the nearest neighbor
(we’ll call this \(d_i\)) from all distances, or equivalently dividing all
weights by the largest weight. Then we’ll ensure that the total weight of every
neighborhood is the same. This is done by finding a scaling factor
\(\sigma_i\) so that&lt;/p&gt;

\[\sum_{j \in N(i)} \exp\left(-\frac{d_{ij} - d_i}{\sigma_i} \right) = W\]

&lt;p&gt;for some target neighborhood weight \(W\) (constant across neighborhoods).
Equivalently, we find some exponent \(\eta\) so that \(\sum_{j \in N(i)}
w_{ij}^\eta = W\). The UMAP paper (without any explanation, or even a mention
that this is a choice) sets \(W = \log_2(k) + 1\), where \(k\) is the number of
neighbors found in the graph. (To be precise, it tries to get the value of the
sum excluding the nearest neighbor of \(i\) to equal \(\log_2(k)\), which is
equivalent.) &lt;sup id=&quot;fnref:2&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2&quot; class=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;Let’s look at what happens when \(i\) has a more than one neighbor at distance
\(d_i\). The first normalization step sets all these distances to zero, which
makes perfect sense. But notice that now the corresponding terms in \(\sum_{j \in
N(i)} w_{ij}\) are all equal to 1, and the next scaling step can’t modify them.
In particular, what if there are more than \(\log_2(k) + 1\) neighbors at the same
distance?&lt;/p&gt;

&lt;p&gt;The UMAP implementation finds the appropriate value of \(\sigma\) by &lt;a href=&quot;https://github.com/lmcinnes/umap/blob/release-0.5.8/umap/umap_.py#L197-L242&quot;&gt;a simple
bisecting
search&lt;/a&gt;,
which works because the objective function is increasing in \(\sigma\). But when
the equation has no solution, this search is terminated after 64 (!) steps,
giving a value of \(\sigma = 2^{-64}\). There’s actually a check for this
situation, and if the value of \(\sigma\) is too small, it’s replaced by
\(10^{-3}\) times the average distance in the graph. This means that the weights
actually affected by \(\sigma\) end up looking like \(\exp\left(-\frac{d_{ij} -
d_i}{ 10^{-3} d_{ij} }\right)\), which is basically \(\exp(-10^3)\). This is
tiny, so much so that it underflows to 0 in 32-bit floating point (which the
UMAP implementation uses).  So if \(i\) has at least \(\log_2(k) + 1\) neighbors
at equal distances, all of the edges to its more distant neighbors end up
getting a weight of zero! The normalization ends up effectively giving us a
graph with \(\log_2(k) + 1\) neighbors instead of \(k\).&lt;/p&gt;

&lt;p&gt;This argument depends on \(i\) having a bunch of neighbors at exactly the same
distance, but the phenomenon is pretty robust to small variations in distance
for the nearby neighbors. The above examples actually have clusters sampled from
uniform distributions of width 0.05, while the distance between cluster
centroids is 1. Increasing the spatial dispersion of the clusters makes the
default UMAP settings perform somewhat better, as expected, since it makes the
difference in distances between near and far neighbors smaller.&lt;/p&gt;

&lt;p&gt;I hacked support for setting \(W\) independently from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; into the
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;umap&lt;/code&gt; library. If we set \(W\) to 10 and keep &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors = 15&lt;/code&gt;, here’s what
the plots look like:
&lt;img src=&quot;/assets/umap_puzzle/tweaked_clusters.png&quot; alt=&quot;UMAP behavior with W = 10&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
The problem is more or less fixed.&lt;/p&gt;

&lt;h2 id=&quot;how-much-does-this-matter&quot;&gt;How much does this matter?&lt;/h2&gt;

&lt;p&gt;For the typical uses of UMAP, probably not that much. Most of the data that
people analyze with UMAP is big enough and has varied enough distances that the
kinds of sharp density transitions that cause this problem aren’t very common.
At least, this is probably the case for data coming from some continuous metric
space. But UMAP also supports a wide variety of more discrete metrics like the
Hamming or Jaccard distances. With these metrics, tied distances are much more
common. In fact, I first ran into this issue in the wild (without knowing it)
while using the &lt;a href=&quot;https://www.stat.berkeley.edu/~breiman/RandomForests/cc_home.htm#prox&quot;&gt;random forest proximity
metric&lt;/a&gt;
(basically, the distance between two data points is the number of trees in the
forest where the points end up in different leaves). In these situations it’s
possible to get quite large sets of points at distance zero from each other.
Probably the right thing to do there is deduplicate the points before trying to
do an embedding that relies on a fixed number of neighbors per point, but only
effectively having \(log_2(k)\) neighbors when you thought you had \(k\)
certainly doesn’t help.&lt;/p&gt;

&lt;p&gt;Another time when this might matter is when looking at small datasets. There’s
&lt;a href=&quot;https://github.com/lmcinnes/umap/issues/94&quot;&gt;an issue from 2018&lt;/a&gt; on the UMAP
GitHub repository complaining that the embeddings of points sampled from a
normal distribution don’t change much as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; increases. Tweaking the
neighborhood weight target does seem to alleviate that issue somewhat. Here’s
what embeddings look like with \(W = \log_2(k) + 1\):
&lt;img src=&quot;/assets/umap_puzzle/default_umap_standard_normal_500.png&quot; alt=&quot;UMAP embedding standard normal data&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
And here they are with \(W = \frac{2}{3} k\):
&lt;img src=&quot;/assets/umap_puzzle/tweaked_umap_standard_normal_500.png&quot; alt=&quot;UMAP embedding standard normal data with W = 2/3 * n_neighbors&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
It’s subtle, but the embeddings with the adjusted neighborhood weights continue
to smooth out a bit as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors&lt;/code&gt; increases.&lt;/p&gt;

&lt;p&gt;This type of behavior seems to be most prevalent in smaller datasets. To use the
excruciatingly standard example of embedding MNIST: if we subsample MNIST to
2000 images and run UMAP with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n_neighbors = 60&lt;/code&gt;, there’s a pretty significant
difference between the two neighborhood normalizations.
&lt;img src=&quot;/assets/umap_puzzle/sample_mnist.png&quot; alt=&quot;Comparing embeddings of an MNIST sample&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
On the full dataset, the difference is minimal; the version with \(W =
\frac{2}{3} k\) maybe looks a little fuzzier.
&lt;img src=&quot;/assets/umap_puzzle/full_mnist.png&quot; alt=&quot;Comparing embeddings of the full MNIST dataset&quot; class=&quot;center-image two&quot; style=&quot;max-width: 100%&quot; /&gt;
Again, in most cases this isn’t going to make a big difference. Still, this does
seem like a bit of an unforced error. I’m a fan of &lt;a href=&quot;https://arxiv.org/abs/2007.08902&quot;&gt;this
paper&lt;/a&gt; that investigated the factors that make
UMAP look different from t-SNE look different from Laplacian eigenmaps.
(Spoiler: it seems to come down to the relative strength of attractive forces
between neighbors and repulsive forces between all pairs of points.) One of the
ablations the authors did was replacing all the graph weights with 1, which had
minimal qualitative effects in the examples they looked at (including, yes,
MNIST). So all the complicated weight processing that caused this issue may not
be buying us much anyway.&lt;/p&gt;

&lt;p&gt;Really, the main thing to take away from this is that the results of any
dimensionality reduction method are never a perfect representation of the
relationships in the data. There is no perfect representation. Every method of
data analysis, supervised or unsupervised, is a way of prying out some insight
into a hidden underlying truth about the data. The better you understand the
techniques you use, the better you’ll be able to grasp that hidden truth.&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Why did UMAP win out so quickly? I think mostly two things (and really
    mainly the first): the standard UMAP implementation was at the time &lt;em&gt;much&lt;/em&gt;
    faster than the standard t-SNE implementation, which really matters for
    something you’re primarily using as a visualization tool. (Nowadays optimized
    implementations of both algorithms are pretty similar in performance.) 
    UMAP embeddings also tend to look more interesting and varied than t-SNE
    plots. The stereotypical t-SNE plot looks like a collection of blobs
    smushed together into a circle, while the stereotypical UMAP plot is
    full of teardrops and spirals and all kinds of dynamic-looking shapes.
    Another way to say this is that UMAP embeddings exhibit very high
    variations in density of points. This turns out to be useful if you’re
    using them as inputs to a density-based clustering algorithm like
    HDBSCAN. &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;My hunch is that it comes from trying to do something sort of like t-SNE’s
    perplexity normalization of the weights for each neighborhood. There the weights
    are normalized to sum to 1 (i.e. be a probability distribution), and the
    corresponding bandwidth parameter \(\sigma\) is set so that the entropy of the
    distribution is equal to \(\log_2(P)\). The number of neighbors used in the
    computation is usually chosen to scale linearly with the perplexity \(P\), so if
    you squint, t-SNE’s normalization looks like \(\log_2(k)\). &lt;a href=&quot;#fnref:2&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name></name></author><summary type="html">If you want to visualize some large high-dimensional dataset, odds are you’re going to try UMAP. After its initial release in 2018, it quickly displaced t-SNE (which had itself displaced, I don’t know, diffusion maps?) as the nonlinear dimensionality reduction method everyone uses by default, particularly in machine learning (for instance, dealing with embeddings or neural network activations) and biology (e.g. for looking at single-cell gene expression profiles).1 Why did UMAP win out so quickly? I think mostly two things (and really &amp;#8617;</summary></entry><entry><title type="html">Sheaves and Probability, Part 3</title><link href="https://www.jakobhansen.org/2021/04/02/sheaves-probability-3/" rel="alternate" type="text/html" title="Sheaves and Probability, Part 3" /><published>2021-04-02T00:00:00-07:00</published><updated>2021-04-02T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2021/04/02/sheaves-probability-3</id><content type="html" xml:base="https://www.jakobhansen.org/2021/04/02/sheaves-probability-3/">&lt;p&gt;(&lt;a href=&quot;/2021/03/17/sheaves-probability&quot;&gt;Part 1&lt;/a&gt;, &lt;a href=&quot;/2021/03/26/sheaves-probability-2&quot;&gt;Part 2&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;The distributed algorithm for inference I described last time is kind of
awkward, and not very well linked with the rest of the field. Olivier Peltre
has a better one. It’s connected with the well-known message-passing
algorithm for this sort of calculation. Unfortunately, while I think I have a
decent idea of the general idea of what’s going on, I don’t have a very
detailed understanding of all the moving pieces, so some of this exposition
might be just a little bit wrong.&lt;/p&gt;

&lt;p&gt;We will need to give up a bit of our nice simplicial structure to
make everything work properly. We will instead work directly with a cover of
our set of random variables \(\mathcal I\) by a collection of sets \(\mathcal
V\), where we require that \(\mathcal V\) be closed under intersection. This
requirement makes \(\mathcal V\) a partially ordered set under the relation
\(\alpha \leq \beta\) if \(\beta \subseteq \alpha\). (The change in direction
is to ensure that our sheaves stay sheaves and our cosheaves stay cosheaves.)
The state spaces for each \(\alpha\) still form a sheaf (i.e., a covariant
functor) over \(\mathcal V\); and so we also get the cosheaf \(\mathcal A\)
of observables and the sheaf \(\mathcal A^*\) of functionals. (This is
probably terrible terminology for anyone not comfortable with cellular
sheaves, since typically we want to think of a presheaf as a contravariant
functor. Sorry.)&lt;/p&gt;

&lt;p&gt;We can move back to the simplicial world by taking the nerve \(\mathcal
N(\mathcal V)\) of \(\mathcal V\): this is the simplicial complex whose
\(k\)-simplices are nondegenerate chains of length \(k+1\) in \(\mathcal V\).
In particular, vertices of \(\mathcal{N}(\mathcal V)\) correspond to elements
of \(\mathcal V\), while edges correspond to pairs \(\alpha \leq \beta\). A
sheaf or cosheaf on \(\mathcal V\) extends naturally to
\(\mathcal{N}(\mathcal V)\). The stalk over a simplex \(\alpha \leq \cdots
\leq \beta\) is \(\mathcal F(\beta)\), and the restriction maps between 0-
and 1-simplices are \(\mathcal F(\alpha) \to \mathcal F(\alpha \leq \beta)\),
\(x_\alpha \mapsto \mathcal F_{\alpha \leq \beta} x_\alpha\), and \(\mathcal
F(\beta) \to \mathcal F(\alpha \leq \beta)\), \(x_\beta \mapsto x_\beta\).&lt;/p&gt;

&lt;p&gt;Extending a sheaf or cosheaf to \(\mathcal N(\mathcal V)\) preserves homology
and cohomology (defined in terms of derived functors and all that), so this
is not an unreasonable thing to do. The interpretations of \(H_0(\mathcal
N(\mathcal V);\mathcal A)\) and \(H^0(\mathcal N(\mathcal V);\mathcal A^*)\)
are also the same. Homology classes of \(\mathcal A\) correspond to
Hamiltonians defined as a sum of local terms, and cohomology classes of
\(\mathcal A^*\) (when normalized) correspond to consistent local marginals
(or pseudomarginals).&lt;/p&gt;

&lt;p&gt;We now come to the part of the framework I have always found most confusing,
because it’s seldom very well explained. Peltre (and others, probably) makes
a distinction between the &lt;em&gt;interaction potentials&lt;/em&gt; \(h_\alpha\) which are
summed to get a Hamiltonian on \(S\), and the collection of &lt;em&gt;local
Hamiltonians&lt;/em&gt; \(H_\alpha\), obtained by treating each subset \(\alpha\) as
if it were an isolated system and summing the potentials \(h_\beta\) for
\(\beta \subseteq \alpha\). Both of these can be seen as 0-cochains of
\(\mathcal A\), and they are not really homologically related. There is an
invertible linear map \(\zeta: C_0(\mathcal{N}(\mathcal V);\mathcal A) \to
C_0(\mathcal{N}(\mathcal V);\mathcal A)\) which sends a collection of
interaction potentials to its family of local Hamiltonians. The formula is
straightforward: \((\zeta h)_\alpha = \sum_{\beta \subseteq \alpha} \mathcal
A_{\alpha \leq \beta} h_\beta\). The inverse of \(\zeta\) is the Möbius transform \(\mu\)
given by \((\mu H)_\alpha = \sum_{\beta \subseteq \alpha} c_{\beta}
\mathcal{A}_{\alpha \leq \beta} H_\beta\). The coefficients \(c_{\beta}\) are
the Möbius numbers of the poset \(\mathcal V\). Notice that these formulas
refer to the poset relations, not to the incidence of simplices in \(\mathcal
N(\mathcal V)\), which is a bit frustrating—the simplicial structure doesn’t
seem to really be doing much here. However, \(\zeta\) and \(\mu\) can be
extended to chain maps, and hence descend to homology. There are analogous
operations on cochains of \(\mathcal A^*\), adjoint to the operations on the
chains of \(\mathcal A\).&lt;/p&gt;

&lt;p&gt;The action of \(H^0(\mathcal N(\mathcal V);\mathcal A^*)\) on \(H_0(\mathcal
N(\mathcal V);\mathcal A)\) calculates the expected value of the Hamiltonian
\(H\) given by a class of interaction potentials \(h_\alpha\) with respect to
a consistent set of marginals. If a chain represents a collection of local
Hamiltonians, then this pairing doesn’t compute an expectation. Instead, if
\(p \in H^0\) and \(H \in C_0\), we have to take \(\langle p, \mu \cdot
H\rangle\) to get the expectation. By adjointness, this is \(\langle \mu\cdot
p, H\rangle\).&lt;/p&gt;

&lt;p&gt;For reasons I don’t understand very deeply, these operations play a key role
in formulating message passing dynamics. My current understanding is that it
has to do with the gradients of the Bethe entropy. The Bethe entropy is
defined on a section of \(\mathcal A^*\) by the same inclusion-exclusion
procedure as we did before, but using the Möbius coefficients to perform
inclusion-exclusion: \(\check{S}(p) = \sum_{\alpha} c_\alpha S(p_\alpha)\).
Think of this as computing local (i.e. marginalized) entropy and then
applying the Möbius transform.&lt;/p&gt;

&lt;p&gt;The Bethe free energy is the functional we minimize to find the approximate pseudomarginals:
\(\langle p, h \rangle - \check{S}(p)\). Adding the constraint that \(p\) be consistent (\(\delta p = 0\)) and normalized (\(\langle p_\alpha, \mathbf{1}_\alpha \rangle = 1\)), we get the Lagrangian condition&lt;/p&gt;

\[h + \mu (\log p + \mathbf{1}) + \partial \lambda + \tau = 0\]

&lt;p&gt;which we can then convert to&lt;/p&gt;

\[-\log p - \mathbf{1} = \zeta(h + \partial \lambda + \tau).\]

&lt;p&gt;The Lagrange multiplier \(\tau\) is an arbitrary 0-chain which is constant on
each vertex, so its zeta transform also satisfies the same condition, and we
can roll the \(\mathbf 1\) into it. \(\lambda\) is an arbitrary 1-chain.
Ultimately, we have the requirement that&lt;/p&gt;

\[-\log p = \zeta(h + \partial \lambda) + \tau.\]

&lt;p&gt;This holds on the level of cochains, and of course the cochain \(p\) must also be
in \(H^0(\mathcal{N}(\mathcal V),\mathcal A^*)\). So in order to be a
critical point for the Bethe free energy functional, \(-\log p\) must be
obtained as the zeta transform of something homologous to the given local
potential \(h\), up to some normalizing additive constant. Solving the
marginalization problem amounts to a search for a local potential \(h'\)
homologous to the given \(h\) that makes \(p = \exp(-\zeta h')\) a section of
\(\mathcal A^*\). Note that every section of \(\mathcal A^*\) is normalized
to some constant sum, so the normalization constraint isn’t all that
interesting. Further, if \(p = \exp(-\zeta h')\), it is automatically
nonnegative.&lt;/p&gt;

&lt;p&gt;One of the key insights now is that if we define a differential equation on
\(h\) where the derivative of \(h\) is always in the image of \(\delta\), the
evolution of the state will be restricted to the homology class of the
initial condition. If we define an operator on \(C_0(\mathcal{N}(\mathcal
V);\mathcal A)\) which vanishes when \(\exp(-\zeta h')\) is a section of
\(\mathcal A^*\), we’ll be able to use it to implement just such a differential equation.&lt;/p&gt;

&lt;p&gt;Let \(\Delta = \partial \circ \mathcal D \circ \zeta\), where \(\mathcal D:
C_0(\mathcal N(V);\mathcal A^*) \to C_1(\mathcal N(\mathcal V); \mathcal A)\)
is defined by 
\((\mathcal D H)_{\alpha \leq \beta} = -\log(\mathcal A^*_{\alpha \leq \beta}\exp(-H_\alpha)) +H_\beta \in \mathcal A(\alpha \leq \beta) = \mathcal A(\beta)\). This operator clearly vanishes when \(\exp(-\zeta
h)\) is a section of \(\mathcal A^*\), since \(\mathcal D \circ \zeta\) does.
In fact, \(\Delta\) vanishes at precisely these
points. This operator \(\Delta\) is somewhat mysterious, but if you squint
hard enough it looks kind of like a Laplacian. In fact, its linearization
around a given 0-chain appears to be a sheaf Laplacian.&lt;/p&gt;

&lt;p&gt;Thus the system of differential equations \(\dot{h} = -\Delta h\) has
stationary points corresponding precisely to the critical points of the Bethe
entropy. If you discretize this ODE with a naive Euler scheme with step size
1, you get a discrete-time evolution equation, and this implies dynamics on
local marginal estimates \(p_\alpha \propto \exp(-(\zeta h)_\alpha)\). It
turns out that these dynamics on \(p\) are precisely the standard
message-passing dynamics, typically described in terms of marginalization,
sums, and products. This is a really neat result. And despite the
complications involved in defining everything, I find this more
comprehensible than the other expositions of generalized message-passing
algorithms I’ve read. (I’m looking at you, &lt;a href=&quot;https://people.eecs.berkeley.edu/~wainwrig/Papers/WaiJor08_FTML.pdf&quot;&gt;Wainwright and
Jordan&lt;/a&gt;.)
I’d still like to understand this better, though. What exactly is the
relationship between the operator \(\Delta\) and a sheaf Laplacian? Is there
a sheaf-theoretic way to understand the max-product algorithm for finding
maximum-likelihood elements rather than marginal distributions? Can we use
(possibly conically constrained) cohomology of \(\mathcal A^*\) to say
anything about the pseudomarginal problem?&lt;/p&gt;</content><author><name></name></author><summary type="html">(Part 1, Part 2)</summary></entry><entry><title type="html">Sheaves and Probability, Part 2</title><link href="https://www.jakobhansen.org/2021/03/26/sheaves-probability-2/" rel="alternate" type="text/html" title="Sheaves and Probability, Part 2" /><published>2021-03-26T00:00:00-07:00</published><updated>2021-03-26T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2021/03/26/sheaves-probability-2</id><content type="html" xml:base="https://www.jakobhansen.org/2021/03/26/sheaves-probability-2/">&lt;p&gt;(See &lt;a href=&quot;/2021/03/17/sheaves-probability&quot;&gt;Part 1&lt;/a&gt; for background.)&lt;/p&gt;

&lt;p&gt;Once the structure of a graphical model (or whatever else you want to call
it) is determined, and we know how to represent it using sheaves and
cosheaves, what do we do with it? Typically, we know the local Hamiltonian
\(H\) and we want to know the probabilities of certain events. Of course, we
don’t usually care about the probabilities of individual global states, but
about the marginal probabilities of local states, or some other statistics
that we can derive from those marginals. For instance, in the Ising model, we
might want to know the distribution of the proportion of local sites that
have \(+1\) spin instead of \(-1\). We don’t need the whole distribution to
calculate this, just the marginal distributions for each site.&lt;/p&gt;

&lt;p&gt;A local Hamiltonian \(H = \sum_\alpha h_\alpha\) produces local probability
potential functions \(\phi_\alpha = \exp(- h_\alpha)\). The resulting
probability distribution on states is defined by \(p_H(x) = \frac{1}{Z}
\prod_{\alpha} \exp(- h_\alpha)\), where \(Z\) is a normalizing factor
called the &lt;strong&gt;partition function&lt;/strong&gt;. \(Z = \sum_{x} \prod_{\alpha} \exp(-
h_\alpha(x_\alpha))\). The only hard part of calculating the probability
distribution is calculating the partition function, because it requires
summing over the exponentially large global state space. The goal is to find
a way to calculate the marginals for each set \(\alpha\) without computing
the whole partition function.&lt;/p&gt;

&lt;p&gt;One way to reframe things is to note that the probability distribution is the solution to an optimization problem. The distribution \(p_H\) is the minimizer (for fixed \(H\)) of the &lt;em&gt;free energy&lt;/em&gt; functional \(\mathbb{F}(p,H) = \mathbb{E}_{p}[H] - S(p)\), where \(S(p)\) is the entropy of the distribution \(p\). 
To see this, we think of \(H\) and \(p\) as big vectors, and form the Lagrangian&lt;/p&gt;

&lt;p&gt;\[\mathcal L(p,\lambda,\mu) = \mathbb{E}_p[H] - S(p) + \lambda (\langle \mathbf{1},p\rangle -1) - \langle \mu, p\rangle.\]&lt;/p&gt;

&lt;p&gt;The expectation is linear in \(p\), and \(S\) is concave, so this is a convex optimization problem. The &lt;a href=&quot;https://en.wikipedia.org/wiki/Karush%E2%80%93Kuhn%E2%80%93Tucker_conditions&quot;&gt;KKT optimality conditions&lt;/a&gt; require that the derivative of \(\mathcal L\) with respect to \(p\) be zero, that \(\langle \mathbf{1}, p\rangle = 1\), \(p \geq 0\), \(\mu \geq 0\), and \(\mu(x)p(x) = 0\) for all \(x\). Differentiating \(\mathcal L\) with respect to \(p\), we get
\[H + (\log(p) + 1) + \lambda \mathbf{1} - \mu = 0,\] 
which gives
\[p(x) = \exp(- H(x) - (\lambda + 1) + \mu(x)) = \exp(- H(x))\exp(-\lambda -1)\exp(\mu(x)).\]&lt;/p&gt;

&lt;p&gt;Since \(\mu(x)p(x) = 0\), we can ignore the term \(\exp(\mu(x))\), since it
is 1 unless \(p(x)\) is zero. Finally, we see that \(\exp(-\lambda-1)\) must
be the normalizing factor \(1/Z\), since \(p\) must be normalized and
changing \(\lambda\) is the only way to make that happen.&lt;/p&gt;

&lt;p&gt;This reformulation alone doesn’t solve the problem, since we still have to
optimize over the space of all possible distributions on \(S\). But remember:
when \(H = \sum_\alpha h_\alpha\) we can calculate \(\mathbb{E}_p[H]\)
locally via the action of \(H^0(X; \mathcal A^*)\) on \(H_0(X; \mathcal A)\).
Unfortunately, the same is not (in general) true of the entropy \(S(p)\).
Naively, we need to actually extend our local marginals to a global
distribution in order to calculate its entropy, which is exactly what we were
trying to avoid. Even worse, a pseudomarginal \(\mu \in H^0(X; \mathcal
A^*)\) might not even correspond to any probability distribution on the total
space of states \(S\). Even if it does, that probability distribution is far
from unique. Assuming \(\mu\) actually corresponds to the local marginals of
some distribution, we can solve this uniqueness problem by assuming that it
corresponds to the distribution with maximal entropy.&lt;/p&gt;

&lt;p&gt;Even checking whether \(\mu\) is a true marginal distribution is NP-hard in general. (The hardness depends on how the covering sets in \(\mathcal V\) are arranged. If \(X\) is a tree, every pseudomarginal is a true marginal, for example.)&lt;/p&gt;

&lt;p&gt;The Bethe-Kikuchi approximation solves these two problems by ignoring them and introducing one of its own.&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;First, forget the problem of finding a real marginal distribution. All we’re looking for are pseudomarginals. We’ll just hope they correspond to marginals of a real distribution. In practice this doesn’t seem to be too big a deal.&lt;/li&gt;
  &lt;li&gt;Second, we will have to replace the entropy with a locally calculated approximation.&lt;/li&gt;
  &lt;li&gt;The new problem is that the Bethe approximation to the entropy is no longer a concave function, and hence we can’t guarantee uniqueness of the obtained distribution.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The definition of the Bethe-Kikuchi entropy of a pseudomarginal is motivated
by an inclusion-exclusion process. A first approximation might be to
calculate the entropy for each local pseudomarginal \(p_\alpha\), for each
maximal covering set \(\alpha\). This is a good start. But, because the sets
may overlap, we’ve double-counted the entropy associated with the common
variables in, say, \(\alpha \cap \beta\). So we subtract this entropy. But these
sets might overlap, so we need to add back in the extra entropy we removed.
Since we removed it 3 times, once for each face of \([\alpha\gamma]\), we
need to add two back in. Eventually we get to&lt;/p&gt;

&lt;p&gt;\[\check{S}(p) = \sum_{\alpha} S(p_\alpha) - \sum_{\alpha,\beta} S(p_{\alpha}) + 2\sum_{\alpha,\beta,\gamma} S(p_{\alpha\beta\gamma}) - \cdots = \sum_{k=1}^n \sum_{\alpha_1,\ldots\alpha_k} (-1)^{k-1}kS(p_{\alpha_1\cdots\alpha_k}).\]&lt;/p&gt;

&lt;p&gt;Note that this definition implicitly assumes that \(p\) is a section of
\(\mathcal A^*\). Coming up with a formula that is defined on
\(C^0(X;\mathcal A^*)\) is a bit messy. One way to do it is to let
\(\check{S}_{\alpha}(p_\alpha) = S(p_\alpha) + \sum_{\alpha \trianglelefteq
\sigma} (-1)^{\dim \sigma}\frac{\dim \sigma}{\dim \sigma + 1}S(A^*_{\alpha
\trianglelefteq \sigma} p_\alpha)\) for each vertex \(\alpha\) of \(X\). Then
we let \(\check{S}(p) = \sum_\alpha \check{S}_\alpha(p_\alpha)\). This splits
the computation of the term corresponding to a \(k\)-simplex of \(X\) equally
between its \(k+1\) vertices.&lt;/p&gt;

&lt;p&gt;The Bethe-Kikuchi entropy is often described in a slightly different but
equivalent way, better adapted to situations where we don’t use the
simplicial structure. For this, we convert the cover \(\mathcal V\), which we
previously assumed to have no sets contained in each other, to its closure
\(\overline{\mathcal V}\) under intersection, and take its reversed poset of
inclusion. The &lt;em&gt;Möbius numbers&lt;/em&gt; of this poset are an assignment of integers
\(c_\alpha\) to each element \(\alpha \in \overline{\mathcal V}\) such that
\(\sum_{\beta\subseteq \alpha} c_\beta = 1\) for every \(\alpha\). This
definition captures the inclusion-exclusion principle behind the
Bethe-Kikuchi approximation. If we let \(\check{S}(p) = \sum_\alpha c_\alpha
S(p_\alpha)\), we get an equivalent definition with less redundancy and a
natural localized definition on 0-cochains, although we are forced to work
with cochains of \(\mathcal A\) defined on \(\overline{\mathcal V}\).&lt;/p&gt;

&lt;p&gt;However we decide to calculate \(\check{S}(p)\), the approximate marginal
inference problem is what one might call a &lt;em&gt;homological program&lt;/em&gt;:&lt;/p&gt;

&lt;p&gt;\[\min_{p} \langle p, h \rangle - \check{S}(p) \text{ s.t. } p \in H^0(X;\mathcal A^*), p \geq 0, p \text{ normalized}\]&lt;/p&gt;

&lt;p&gt;The Bethe-Kikuchi entropy is not concave, so this is not a convex
optimization problem. But it is localizable as discussed earlier, so we can
try some of our &lt;a href=&quot;/publications/distopt.pdf&quot;&gt;distributed homological
programming&lt;/a&gt; techniques without the optimality
guarantees. The general idea is to replace the constraint \(p \in
H^0(X;\mathcal A^*)\) with the constraint \(L_{\mathcal A^*} p = 0\). The
objective function is \(F(p) = \sum_{\alpha} \langle p_\alpha,
h_\alpha\rangle - \check{S}_\alpha(p_\alpha)\), and we can construct a local
Lagrangian&lt;/p&gt;

&lt;p&gt;\[\mathcal L(p,\lambda,\mu,\tau) = \sum_\alpha \mathcal L_\alpha(p,\lambda_\alpha,\mu_\alpha, \tau_\alpha),\]&lt;/p&gt;

&lt;p&gt;with 
\(\mathcal{L}_\alpha(p,\lambda_\alpha,\mu_\alpha,\tau_\alpha) = \langle p_\alpha, h_\alpha \rangle - \check{S}_\alpha (p_\alpha) + \langle \lambda_\alpha, (L_{\mathcal{A}^*} p)_\alpha\rangle + \mu_\alpha (\langle \mathbf{1}, p_\alpha \rangle - 1) + \langle \tau_\alpha,p_\alpha \rangle\).&lt;/p&gt;

&lt;p&gt;The local Lagrangian \(\mathcal{L}_\alpha\) only depends on the values of
\(p_\beta\) where \(\beta\) is a neighboring vertex to \(\alpha\). The
primal-dual dynamics on \(\mathcal L\)—gradient descent on \(p\) and ascent
on the dual variables—is also locally determined, and (hopefully) converges
to a critical point of the optimization problem. You can then discretize the
continuous-time dynamics and work out what the messages passed between nodes
of \(X\) look like, but that’s a lot of work and not particularly
enlightening. While it’s interesting that you can come up with a distributed
algorithm using the general nonsense of homological programming, I’m not sure
this is actually a fruitful approach. The resulting algorithm doesn’t look
anything like the well-studied message-passing algorithms for marginal
inference. And it has some obvious deficits. For instance, if \(X\) is a
tree, the standard message-passing algorithms converge to an exact solution
in finitely many steps, while convergence of this optimization algorithm is
asymptotic at best (and possibly not guaranteed to happen).&lt;/p&gt;

&lt;p&gt;There is a better approach, developed by Olivier Peltre, that makes
connections between this sheaf-theoretic perspective and message passing
algorithms. His perspective interprets message-passing as a discrete-time
approximation to the flow of a nonlinear Laplacian-like operator on 0-chains
of \(\mathcal A\). Part 3 will outline this framework and its implications.&lt;/p&gt;</content><author><name></name></author><summary type="html">(See Part 1 for background.)</summary></entry><entry><title type="html">Sheaves and Probability, Part 1</title><link href="https://www.jakobhansen.org/2021/03/17/sheaves-probability/" rel="alternate" type="text/html" title="Sheaves and Probability, Part 1" /><published>2021-03-17T00:00:00-07:00</published><updated>2021-03-17T00:00:00-07:00</updated><id>https://www.jakobhansen.org/2021/03/17/sheaves-probability</id><content type="html" xml:base="https://www.jakobhansen.org/2021/03/17/sheaves-probability/">&lt;p&gt;Let me pretend to be a physicist for a moment. Consider a system of \(n\)
particles, where each particle \(i\) can have a state in some set \(S_i\).
The total state space of the system is then \(S = \prod_i S_i\), which grows
exponentially as \(n\) increases. If we want to study probability
distributions on the state space, we very quickly run out of space to
represent them. Physicists have a number of clever tricks to represent and
understand these systems more efficiently.&lt;/p&gt;

&lt;p&gt;For instance, the time evolution of the system is oftened determined by
assigning each state an energy and following some sort of Hamiltonian dynamics.
That is, we have a function \(H: S \to \Reals\) giving the energy of each
state. But specifying the energy of exponentially many states is
exponentially hard. One way to solve this problem is to define \(H\)
&lt;em&gt;locally&lt;/em&gt;. The simplest way to do this is to say \(H = \sum_i h_i\), where
\(h_i\) depends only on the state of particle \(i\). But of course this means
that there are no interactions between the particles. We can increase the
scope of locality, letting \(H = \sum_{\alpha} h_\alpha\), where each
\(\alpha\) is a subset of particles, and \(h_\alpha\) depends only on the
states of the particles in \(\alpha\).&lt;/p&gt;

&lt;p&gt;In a thermodynamic setting, the probability of a given state \(\mathbf{s}\)
is typically proportional to \(\exp(-\beta H(\mathbb{s}))\), where \(\beta\)
is an inverse temperature parameter, making states with lower energies more
probable. When \(H = \sum_\alpha h_\alpha\), this becomes \(\prod_\alpha
\exp(-\beta h_\alpha(\mathbb{s}))\). This is relatively easy to compute for
any single state \(\mathbf{s}\). The hard part is finding the constant of
proportionality. The actual probability is \(p(\mathbf{s}) =
\frac{1}{Z(\beta)} \prod_\alpha \exp(-\beta h_\alpha(\mathbb{s}))\), where
\(Z\) is the &lt;em&gt;partition function&lt;/em&gt;, computed by summing over all possible
states.&lt;/p&gt;

&lt;p&gt;We can think about this problem from a purely probabilistic perspective as
well; it doesn’t hinge on the thermodynamics. Consider a set of
non-independent discrete random variables \(X_i\) for \(i \in \mathcal I\),
whose joint distribution is \(p(\mathbf{x}) \propto
\prod_{\alpha}\psi_{\alpha}(\mathbf{x}_\alpha)\), where each \(\psi_\alpha\)
is a nonnegative function defined for the random variables in some subset
\(\alpha\) of \(\mathcal I\). Denote the collection of subsets \(\alpha\)
used in this product by \(\mathcal V\).&lt;/p&gt;

&lt;p&gt;Why is this interesting? Here are a few examples:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;The Ising Model&lt;/strong&gt;. This is one of the earliest statistical models for
magnetism. We have a lattice of atoms, each with spin \(X_i \in \{\pm 1\}\).
The joint probability of finding the atoms in a given spin state is
\(p(\mathbf{x}) \propto \exp(\sum_i \phi_i x_i + \sum_{i\sim j} \theta_{ij}
x_ix_j) = \prod_i \psi_i(x_i) \prod_{i \sim j}\psi_{ij}(x_i,x_j)\). We can
write this in the Hamiltonian form as well: \(H = -\sum_{i} \phi_i x_i -
\sum_{i \sim j} \theta_{ij} x_i x_j\).&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;LDPC Codes.&lt;/strong&gt; These are a ubiquitous tool in coding theory, which designs
ways to transmit information that are robust to error. We encode some
information in redundant binary &lt;em&gt;code words&lt;/em&gt;, which are then transmitted. The
code is designed so that if a few bits are corrupted during transmission, it
is possible to recover the correct code word.
LDPC codes are a particularly efficient type of code. Here we treat the code
words as realizations of a tuple of random variables \((X_1,\ldots,X_n)\),
valued in \(\mathbb{F}_2\). The probability distribution used will be simple,
and will be supported solely on the set of acceptable codewords. For an
\((n,k)\) LDPC code, a code word has length \(n\), and satisfies a collection
of \(n-k\) constraints of the form \(\sum_{i \in \alpha} x_i = 0\). In other
words, the set of codewords is the kernel of an \((n-k)\times n\) matrix with
\(\mathbb{F}_2\) entries.&lt;/p&gt;

    &lt;p&gt;The local potential functions \(\psi_\alpha\) are simple: they are equal to 1
  when \(\sum_{i \in \alpha} x_i = 0\) and 0 otherwise. Note that the
  corresponding local Hamiltonian would be infinite for non-codewords. The
  problem of finding the most likely code word given a received tuple is an
  inference problem with respect to this (or a closely related) distribution.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Graphical Models.&lt;/strong&gt; In various statistical or machine learning tasks, it can
be useful to work by specifying a joint distribution that factors according
to some local decomposition of the variables. In general, there should be a
term for each set of variables that is closely related. In a
spatially-related task, these might be random variables associated with
nearby points in space. The clusters might also correspond to already-known
facts about a domain, like closely-related proteins in an analysis of a
biological system.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are a number of interesting facts about probability distributions that
factor in this way. One is a form of conditional independence. Given a
factorization \(p(\mathbf{x}) \propto \prod_{\alpha}
\psi_{\alpha}(\mathbf{x}_\alpha)\), we can produce a graph \(G_{\mathcal V}\)
whose vertices are associated with the random variables \(X_i\) and whose
edges are determined by adding a clique to the graph for each subset \(\alpha
\in \mathcal V\).&lt;/p&gt;

&lt;p&gt;Another, slightly more topological way to construct this graph is as the
1-skeleton of a Dowker complex, where the relation is the inclusion relation
between the set of indices \(I\) and the set \(\mathcal V\) of subsets
\(\alpha\) used in the functions \(\psi_\alpha\).&lt;/p&gt;

&lt;p&gt;If we observe a set \(S\) of random variables, and the removal of the
corresponding set of nodes separates two vertices \(u,v\) in the graph, the
corresponding random variables \(X_u, X_v\) are conditionally independent
given the observations. This property makes the collection of random
variables a &lt;strong&gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Markov_random_field&quot;&gt;Markov random
field&lt;/a&gt;&lt;/strong&gt; for the graph. It
is a fascinating but nontrivial fact that these two properties are
equivalent: a set of random variables (with strictly positive distributions)
has a distribution that splits into factors corresponding to the cliques of
\(G_{\mathcal V}\) if and only if it is a Markov random field with respect to
\(G_{\mathcal V}\). This result is the &lt;a href=&quot;https://en.wikipedia.org/wiki/Hammersley%E2%80%93Clifford_theorem&quot;&gt;Hammersley-Clifford
theorem&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The interpretation of the graph \(G_{\mathcal V}\) as the 1-skeleton of a
Dowker complex hints at some deeper relationships between Markov random
fields, factored probability distributions, and topology. My goal here is to
bring some of these relationships into clearer focus, with a good deal of
help from &lt;a href=&quot;https://opeltre.github.io&quot;&gt;Olivier Peltre&lt;/a&gt;’s PhD thesis &lt;a href=&quot;https://arxiv.org/abs/2009.11631&quot;&gt;Message
Passing Algorithms and Homology&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Let’s construct a simplicial complex and some sheaves from the scenario we’ve
been considering. Start with a set \(\mathcal I\) indexing random variables
\(\{X_i\}_{i \in \mathcal I}\), with \(S_i\) the codomain of the random
variable (here just a finite set) for each \(i \in \mathcal I\). Take a cover
\(\mathcal V\) of \(\mathcal I\)—i.e., a collection of subsets of
\(\mathcal I\) whose union is \(\mathcal I\). Let \(X\) be the Cech nerve of
this cover. That is, \(X\) a simplicial complex with a 0-simplex \([\alpha]\)
for each \(\alpha \in \mathcal V\), a 1-simplex \([\alpha, \beta]\) for each
pair \(\alpha,\beta \in \mathcal V\) with nonempty intersection (and we may
as well assume \(\alpha,\beta\) are distinct), a 2-simplex
\([\alpha,\beta,\gamma]\) for each triple with nonempty intersection,
etcetera.&lt;/p&gt;

&lt;p&gt;We’ll use a running example of an Ising-type model on a very small graph. Let
\(\mathcal I = \{1,2,3,4\}\), \(\mathcal{V} = \{\{1,2\},\{2,3\},\{3,4\}\}\),
and \(S_i = \{\pm 1\}\). \(X\) is then a path graph with 3 nodes and 2 edges.
A warning: \(X\) is not the same as \(G_{\mathcal V}\), which has 4 nodes and
3 edges. Each vertex of \(X\) corresponds to an edge of \(G_{\mathcal V}\).
\(X\) and \(G_{\mathcal V}\) (or, really, the clique complex of \(G_{\mathcal
V}\)) are dual in a sense: maximal simplices of the clique complex of
\(G_{\mathcal V}\) correspond to vertices of \(X\).&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/probabilitysheaves/isingcover.svg&quot; alt=&quot;Variables, covering, and states&quot; class=&quot;center-image&quot; /&gt;&lt;/p&gt;

&lt;p&gt;There is a natural cellular sheaf of sets on \(X\), given by letting
\(\mathcal{F}([\alpha_1,\ldots,\alpha_k]) = \prod_{i \in \bigcap \alpha_j}
S_i\). That is, over the simplex corresponding to the intersection of the
sets \(\alpha_1\cap\cdots\cap \alpha_k\), the stalk is the set of all
possible outcomes for the random variables contained in that intersection.
The restriction maps of \(\mathcal F\) are given by the projection maps of
the product. For instance \(\mathcal{F}_{[\alpha] \trianglelefteq
[\alpha,\beta]}\) is the projection \(\prod_{i \in \alpha} S_i \to \prod_{i
\in \alpha \cap \beta} S_i\). Let’s denote \(v_1 = [\{1,2\}]\), \(v_2 =
[\{2,3\}]\), \(v_3 = [\{3,4\}]\), and \(e_{12} = [\{1,2\},\{2,3\}]\), \(e_{23} = [\{2,3\},\{3,4\}]\).
Then \(\mathcal F(v_i) = \{(\pm 1,\pm 1)\}\) for all \(i\) and \(\mathcal
F(e_{ij}) = \{\pm 1\}\). The restriction map \(\mathcal F_{v_1
\trianglelefteq e_{12}}\) sends \((x,y) \mapsto y\). We call this sheaf
\(\mathcal F\) the sheaf of (deterministic) states of the system. Its
sections correspond exactly to global states in \(S = \prod_i S_i\).&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/probabilitysheaves/isingstatesheaf.svg&quot; alt=&quot;The Ising state sheaf&quot; class=&quot;center-image&quot; /&gt;&lt;/p&gt;

&lt;p&gt;We can construct a cellular cosheaf of vector spaces on \(X\) by letting each
stalk be the vector space \(\Reals^{\mathcal F(\alpha)}\). The extension maps
of this cosheaf are those induced by the functor \(\Reals^{(-)}\). You might
call this “cylindrical extension”: when \(f: \mathcal F([\alpha]) \to
\mathcal F([\alpha, \beta])\), a function \(h: \mathcal F([\alpha, \beta])
\to \Reals\) pulls back to a function \(h^*: \mathcal F([\alpha]) \to
\Reals\) by \(h^*(x) = h(f(x))\). We call the resulting cosheaf \(\mathcal
A\), the cosheaf of observables.&lt;/p&gt;

&lt;p&gt;In the Ising example, \(\mathcal A(v_i) = \Reals^{\mathcal F(v_i)} =
\Reals^{\{\pm 1\}\times \{\pm 1\}}\simeq \Reals^{2}\otimes \Reals^2\) (with
basis \(\{e_{\pm 1} \otimes e_{\pm 1}\}\)) and \(\mathcal A(e_{ij}) =
\Reals^{\mathcal F(e_{ij})} = \Reals^{\{\pm 1\}} \simeq \Reals^2\). The
extension map \(\mathcal A_{v_1 \trianglelefteq e_{12}}\) sends \(x\) to \((e_1 + e_2)
\otimes x\).&lt;/p&gt;

&lt;p&gt;You can probably already see that notation is a major struggle when working
with these objects. It’s a problem akin to keeping track of tensor indices.
One reason the sheaf-theoretic viewpoint is helpful is that, once defined, it
gives us a less notationally fussy structure to hang all this data from.&lt;/p&gt;

&lt;p&gt;The linear dual of the cosheaf \(\mathcal A\) is a sheaf \(\mathcal A^*\)
with isomorphic stalks and restriction maps given by “integrating” or summing
along fibers. If we take the obvious inner product on the stalks of
\(\mathcal A\) (the one given by choosing the indicator functions for
elements of \(\mathcal F(\sigma)\) as an orthonormal basis), these
restriction maps are the adjoints of the extension maps of \(\mathcal A\).
Probability distributions on \(\mathcal{F}([\alpha]) = \prod_{i \in \alpha} S_i\) naturally lie
inside \(\mathcal A^*([\alpha])\); the restriction map \(\mathcal A^*([\alpha])
\to \mathcal A^*([\alpha, \beta])\) computes marginalization onto a subset
of random variables.&lt;/p&gt;

&lt;p&gt;The probability distributions on \(\mathcal{F}([\alpha])\) are those elements
\(p\) of \(\mathcal A^*([\alpha])\) with \(p(\mathbf{1}) = 1\) and \(p(x^2)
\geq 0\) for every observable \(x\). Global sections of \(\mathcal A^*\)
which locally satisfy these constraints are sometimes known as
&lt;em&gt;pseudomarginals&lt;/em&gt;. They do not always correspond to probability distributions
on \(S = \prod_{i\in \mathcal{I}} S_i\). But they’re as close as we can get while only working
with local information. You can always get some putative global distribution,
but it might assign negative probabilities to some events, even if all the
marginals are nonnegative.&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; We can therefore think of \(\mathcal A^*\) (or
at least the subsheaf given by the probability simplex in each stalk) as a
sheaf of probabilistic states.&lt;/p&gt;

&lt;p&gt;Because \(\mathcal A^*\) is the linear dual of \(\mathcal A\), stalks of
\(\mathcal A^*\) act on stalks of \(\mathcal A\). If \(p \in \mathcal
A^*(\sigma)\) is a probability distribution on \(\mathcal{F}(\sigma)\), the
action on \(h \in \mathcal A(\sigma)\), \(\langle p, h\rangle\) takes the
expectation of \(h\) with respect to the probability distribution \(p\).
Further, this action commutes with restriction and extension maps: \(\langle
p_\sigma,\mathcal A_{\sigma \trianglelefteq \tau} h_\tau \rangle = \langle
\mathcal A^*_{\sigma \trianglelefteq \tau} p_\sigma, h_\tau\rangle\).&lt;/p&gt;

&lt;p&gt;Suppose \(h\) is a 0-chain of \(\mathcal A\); we can think of \(h\) as
defining a global Hamiltonian \(H\) on \(S\). That is, \(H = \sum_{\alpha \in
\mathcal V} h_\alpha \circ \pi_{\alpha}\), where \(\pi_\alpha: S \to
\mathcal{F}([\alpha])\) is the natural projection. The expectation of such an
observable \(H\) depends only on the marginals \(p_\alpha\) of the global
probability distribution \(P\) for each set \(\alpha\). So if \(p\) is the
0-cochain of marginals of \(P\), \(\mathbb{E}_P[H] = \langle p, h\rangle\).&lt;/p&gt;

&lt;p&gt;Further, if \(h\) and \(h'\) are homologous 0-chains, they define the same
global function \(H\); adding a boundary \(\partial x\) to \(h\) corresponds
to shifting the accounting of the energy associated with some random
variables from one covering set to another. We therefore naturally have an
action of \(H^0(\mathcal{A}^*;X)\) on \(H_0(\mathcal A; X)\). It computes the
expectation of a local Hamiltonian with respect to a set of pseudomarginals.
So in order to compute the expected energy if the system state follows a
certain probability distribution, we only need the marginals and the local
Hamiltonians. The locality in the systems we consider is captured in this
sheaf-cosheaf pair.&lt;/p&gt;

&lt;p&gt;So far, all we’ve done is represented the structure of a graphical model in
terms of a sheaf-cosheaf pair. The duality between marginal distributions and
local Hamiltonians is implemented by linear duality of sheaves and cosheaves.
&lt;a href=&quot;/2021/03/26/sheaves-probability-2&quot;&gt;Next time&lt;/a&gt; we’ll see if there’s anything we can actually do with this
language. Can we use it to understand or design inference algorithms?
(Answer: yes, sort of.)&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;It’s a fun exercise to use the discrete Fourier transform and the Fourier slice theorem to show this. &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;</content><author><name></name></author><summary type="html">Let me pretend to be a physicist for a moment. Consider a system of \(n\) particles, where each particle \(i\) can have a state in some set \(S_i\). The total state space of the system is then \(S = \prod_i S_i\), which grows exponentially as \(n\) increases. If we want to study probability distributions on the state space, we very quickly run out of space to represent them. Physicists have a number of clever tricks to represent and understand these systems more efficiently.</summary></entry><entry><title type="html">Hybrid Systems and Homotopy Theory</title><link href="https://www.jakobhansen.org/2021/03/04/hybrid-systems/" rel="alternate" type="text/html" title="Hybrid Systems and Homotopy Theory" /><published>2021-03-04T00:00:00-08:00</published><updated>2021-03-04T00:00:00-08:00</updated><id>https://www.jakobhansen.org/2021/03/04/hybrid-systems</id><content type="html" xml:base="https://www.jakobhansen.org/2021/03/04/hybrid-systems/">&lt;p&gt;Hybrid systems are a modeling approach for interacting discrete-time and continuous-time phenomena. Probably the best introduction is via a pair of canonical examples.&lt;/p&gt;

&lt;p&gt;First, consider the dynamics of a ball bouncing up and down on the ground. Most of the time its motion is governed by Newton’s laws. When it hits the ground, we have a modeling choice. We can try to model the deformation of the ball, the transfer of energy and resulting acceleration and deceleration (complicated) or use what we know about how balls bounce: the velocity vector gets turned around and reduced by some scaling factor (much easier). This second approach is the hybrid system model, because there are discontinuities in the state of the ball that happen at discrete times.&lt;/p&gt;

&lt;p&gt;As another example, let’s look at a thermostat. Most thermostats function by turning a heat source on and off—there is no intermediate output. The system state is just the temperature of the room. When the heater is on, this state evolves according to one differential equation (hopefully one that causes temperature to increase); when it is off it follows a different equation. The thermostat is able to switch between these vector fields in response to the state of the system.&lt;/p&gt;

&lt;p&gt;These two examples capture the two main features hybrid systems are designed for:&lt;/p&gt;
&lt;ol&gt;
  &lt;li&gt;the possibility for discontinuous changes in the state of the system (bouncing ball)&lt;/li&gt;
  &lt;li&gt;the possibility for discontinuous changes in the equations of motion for the system (thermostat)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A common formalism for hybrid systems comes from, essentially, gluing a finite state machine to some continuous state spaces and continuous-time dynamics.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Definition.&lt;/strong&gt; A &lt;em&gt;hybrid system&lt;/em&gt; is given by a directed graph \(G\), together with a smooth manifold \(S_v\) (possibly with boundary) for each vertex \(v\) of \(G\), endowed with a vector field \(X_v\), and for each edge \(e: u \to v\) of \(G\), an embedded submanifold \(G_e\) of \(S_u\) (called a guard) and a smooth map \(r_e: G_e \to S_v\) (called a reset map).&lt;/p&gt;

&lt;p&gt;The semantics aren’t hard to understand once you get the idea. The state of the system may be in any one of the spaces \(S_u\), and evolves according to the vector field \(X_u\). If the state hits one of the guards \(G_e\), it is instantaneously sent to the next space \(S_v\) via the reset map \(r_e\).&lt;/p&gt;

&lt;p&gt;For the bouncing ball, \(G\) is the graph with one vertex \(u\) and one edge \(e: u \to u\). The state space \(S_u\) is \(\Reals_{\geq 0} \times \Reals\), consisting of the (positive) position and (signed) velocity of the ball. The guard manifold \(G_e\) is \(0 \times \Reals_{\leq 0}\), capturing the events where the ball is on the ground with negative velocity. The reset map \(r_e\) sends \((0,v)\) to \((0,-cv)\), where \(0 &amp;lt; c &amp;lt; 1\) is the coefficient of restitution of the ball. The vector field \(X_u\) is determined by gravity: \((\dot{x},\dot{v}) = (v,-g)\).&lt;/p&gt;

&lt;p&gt;For the thermostat, there are two vertices in the graph, corresponding to the two possible states of the controller. \(S_{\text{on}} = S_{\text{off}} = \Reals\), and there is an edge from \(\text{off}\) to \(\text{on}\) and one from \(\text{on}\) to \(\text{off}\). The thermostat has a temperature \(T_{\text{on}}\) at which it turns on the heater, and a temperature \(T_{\text{off}}\) at which it turns it off. These two points are the guard manifolds, and somewhat confusingly, \(G_{\text{on}\to \text{off}} = T_{\text{off}}\), \(G_{\text{off}\to\text{on}} = T_{\text{on}}\). The reset maps are the natural inclusion, so the only difference is in the vector fields. I won’t try to describe these, because modeling heat flow is not my deal.&lt;/p&gt;

&lt;p&gt;Hybrid systems can display behavior (some might call it pathological) that is not exhibited by either discrete- or continuous-time systems. Let’s take another look at the ball. Because the coefficient of restitution is less than 1, the ball’s speed decreases every time it hits the guard manifold, and so each time its bounce height decreases. Because the acceleration due to gravity is constant, the time between bounces decreases geometrically. So the transition times when it hits the guard set accumulate at a certain point, and it’s hard to define what happens to the system after that time. 
This is called Zeno behavior, since the ball seems to actually carry out the sequence of actions Zeno argued you must in order to go from point A to point B: go half the distance, then go half the distance remaining, and so forth. It’s typically seen as a problem in a hybrid system, since it’s not typically physically realizable, and may lead to undefined behavior.&lt;/p&gt;

&lt;p&gt;A while back someone told me about a PhD thesis that somehow managed to apply the theory of model categories to hybrid systems. I couldn’t find anything about this at the time, but I recently uncovered it. The dissertation of legend is actually two: Aaron Ames wrote an &lt;a href=&quot;http://ames.caltech.edu/A%20categorical%20theory.pdf&quot;&gt;engineering PhD thesis&lt;/a&gt; and a &lt;a href=&quot;http://ames.caltech.edu/Hybrid%20model%20structures.pdf&quot;&gt;mathematics masters thesis&lt;/a&gt; covering two sides of the story. As is the norm in the world of engineering, this turned into a &lt;a href=&quot;http://ames.caltech.edu/AmesHSCC2005.pdf&quot;&gt;bunch&lt;/a&gt; &lt;a href=&quot;http://ames.caltech.edu/AmesACC2005.pdf&quot;&gt;of&lt;/a&gt; &lt;a href=&quot;http://ames.caltech.edu/HCatGraph.pdf&quot;&gt;other&lt;/a&gt; &lt;a href=&quot;http://ames.caltech.edu/AMSHybridModel.pdf&quot;&gt;papers&lt;/a&gt; (and those are just the ones about a very small segment of the thesis).&lt;/p&gt;

&lt;p&gt;There are a lot of interesting ideas in these theses, both about hybrid systems and other stuff. I was pleased to find that these constructions are basically cryptomorphic versions of cellular sheaves and cosheaves.&lt;/p&gt;

&lt;p&gt;A categorical hybrid system interprets the underlying graph \(G\) of the hybrid system as a category \(\mathcal{G}\) with one object for each edge and one object for each vertex. Each attachment between an edge and a vertex is realized by a morphism from the edge to the vertex. (This is not &lt;em&gt;quite&lt;/em&gt; the incidence category of a regular cell complex, since there may be self-loops.) We also want to keep track of the orientation of each edge, so there’s a bit of extra structure tagging each incidence morphism with the information of whether it goes toward the source of the edge or the target of the edge. A categorical hybrid system is just a functor \(\mathcal{X}: \mathcal{G} \to \mathbf{Man}\) to the category of smooth manifolds.&lt;/p&gt;

&lt;p&gt;To upgrade a traditional hybrid system into a categorical one, we let \(\mathcal{X}(v) = X_v\) and \(\mathcal{X}(e) = G_e\). The action of \(\mathcal{X}\) on morphisms is as follows. Each edge has two outgoing morphisms \(e_s\) and \(e_t\), going to its source and target vertices. \(\mathcal{X}(e_s)\) should be the inclusion map \(G_e \hookrightarrow X_{s(e)}\), while \(\mathcal{X}(e_t)\) should be the reset map \(r_e: G_e \to X_{t(e)}\).
The actual dynamics of the system are added to the vertex stalks afterwards; we think of \(\mathcal{X}\) as being a “hybrid state space” on which we can define any dynamics we please.&lt;/p&gt;

&lt;p&gt;I like thinking about these as cellular cosheaves, even though that’s not quite technically correct, both because I’m comfortable with that language, but also because you can kind of think of the execution of a hybrid system as a sort of holonomic walk over the underlying graph.
We can take the colimit of the functor \(\mathcal{X}\) (in a suitable category of topological spaces) to get a genuine space where the dynamics happen.
The colimit forgets some topological information about the discrete part of the system. If we want to preserve this, we can take the homotopy colimit.
It turns out that if the homotopy colimit of one of these categorical hybrid systems is contractible, Zeno behavior cannot occur (regardless of choice of vector fields).&lt;/p&gt;

&lt;p&gt;This result sounds very alluring, but I couldn’t actually find an explicit proof of it. The closest I found was an explanation that if the underlying graph of the hybrid system has no directed cycles, then Zeno behavior is impossible, because there can only be finitely many state transitions. If the homotopy colimit is simply connected, the underlying graph must have no cycles. After this explanation, the result no longer seems quite so deep—it’s something we could have deduced without any of the algebraic topology. In fact, by looking at the underlying graph itself, we actually get a finer invariant, because we can look for directed cycles rather than any sort of cycle.&lt;/p&gt;

&lt;p&gt;As an example, the thermostat cannot exhibit Zeno behavior as long as the on and off temperatures are separated by a positive distance and the vector fields are fixed. This is because it always takes the same amount of time to go between the temperatures. But both the colimit of the corresponding hybrid functor and the homotopy colimit are homeomorphic to \(S^1\). We can’t rule out Zeno behavior by the homotopy condition.
On the other hand, this criterion does correctly indicate that Zeno behavior is possible for the bouncing ball, but only if you take the homotopy colimit. The standard colimit is just a cone, while the homotopy colimit is a cylinder.&lt;/p&gt;

&lt;p&gt;The masters thesis is dedicated to exploring the homotopy theory of these categorical hybrid systems. Ames constructs a model category structure for the category of hybrid functors, which recovers the homotopy theory of the homotopy colimit. You can compute homotopy or homology using something like the Bousfield-Kan spectral sequence. The second page of this spectral sequence is essentially given by taking stalkwise homology of the cosheaf and computing cosheaf homology.&lt;/p&gt;

&lt;p&gt;There may be more to this than I’ve been able to uncover, though. A couple papers hint at using the actual vector fields of the hybrid system to define a sort of Morse homology, which might be a better way to identify non-Zeno behavior. But that project doesn’t seem to have gone anywhere. Instead, Ames is &lt;a href=&quot;http://ames.caltech.edu/index.html&quot;&gt;a successful researcher at Caltech&lt;/a&gt; building robots, with the only category theory or homological algebra relegated to diagrams on his website.&lt;/p&gt;

&lt;p&gt;I guess this is a lesson about using fancy math to solve real problems. It takes a lot of work—much more than just a dissertation—to take an idea (“hybrid systems are functors”) from abstract mathematics to anywhere near an application. And the incentives often aren’t there to carry a program like that through to completion, even if it might end up being useful.
I am, of course, unable to resist speculating about other high-powered tools that might be brought to bear on the problem. Trajectories of a hybrid system are directed, so maybe directed algebraic topology has something meaningful to say. Or maybe the theory of stratified spaces might be useful in understanding the colimit of one of these hybrid functors. But there are no guarantees, and without a clear motivating problem, it’s not likely that this line of research will ever, say, produce novel guarantees about the behavior of hybrid systems.&lt;/p&gt;

&lt;p&gt;I do think the analogy with cellular cosheaves helped me understand the structure of hybrid systems better. Maybe the real benefit of work like this is giving mathematicians a way to readily understand and compartmentalize what more applied researchers are doing.&lt;/p&gt;</content><author><name></name></author><summary type="html">Hybrid systems are a modeling approach for interacting discrete-time and continuous-time phenomena. Probably the best introduction is via a pair of canonical examples.</summary></entry></feed>