← Back to Writings

When the World Moves, the Prediction Should Follow

Equivariance in deep learning tells us how a neural network predictions should change when the input is transformed.

Date
March 1, 2026
Links
›Contents

Rotate a photograph of a dog and it remains a dog. Rotate its segmentation mask and the mask should rotate with it.

These two statements sound almost obvious, but the former ask the classifier to ignore the rotation, while the latter ask for its preservation in the output. The first property is called invariance, the second: equivariance.

Most neural networks are expected to learn these behaviours from examples. Show the model enough rotated dogs, translated molecules, or permuted point clouds and it may eventually infer which changes matter and which don't. This works, but it also mean spending data and more parameters on relearning structure that was already (the features of a rotated dogs stays the same).

Equivariant deep learning starts from the opposite position: if a transformation has a predictable effect on the task, then it should be encoded directly.

This article expands on a report I wrote about Equivariant Adaptation of Large Pretrained Models. The paper introduces EquiAdapt, a wrapper that makes a pretrained model equivariant by learning how to return each input to a canonical pose before prediction (Mondal et al., 2023). Its central finding is more general than the method itself: mathematical symmetry is useful only if it remains aligned with the distribution on which the model learned.

What should a neural network know before it sees any data?

Invariance cats

Invariance example.

Equivariance cats

Equivariance example.

Symmetry as a transformation

A symmetry is not merely a visual resemblance, but it is a transformation under which some relevant structure is preserved.

The natural language for these transformations is a group. A group (G,∘)(G,\circ) contains a set of elements and an operation satisfying four conditions:

  1. Closure: g1∘g2∈Gg_1\circ g_2\in G for all g1,g2∈Gg_1,g_2\in G.
  2. Associativity: (g1∘g2)∘g3=g1∘(g2∘g3)(g_1\circ g_2)\circ g_3=g_1\circ(g_2\circ g_3).
  3. Identity: there is an e∈Ge\in G such that e∘g=g∘e=ge\circ g=g\circ e=g.
  4. Inverse: every g∈Gg\in G has a g−1∈Gg^{-1}\in G satisfying g∘g−1=eg\circ g^{-1}=e.

The group describes how transformations compose. To describe what they do to data, we also need a group action. For a vector space X\mathcal X, a linear action is represented by

ρX:G→GL(X),\rho_X:G\rightarrow \mathrm{GL}(\mathcal X),

where GL(X)\mathrm{GL}(\mathcal X) is the set of invertible linear maps on X\mathcal X. A representation respects the group structure:

ρX(g1g2)=ρX(g1)ρX(g2),ρX(e)=I.\rho_X(g_1g_2)=\rho_X(g_1)\rho_X(g_2), \qquad \rho_X(e)=I.

Obviously, the same abstract group can act differently on different spaces. A rotation may rearrange image pixels in the input, rotate a mask in the output, or multiply a feature vector by a representation matrix.

Some groups appear repeatedly in machine learning:

GroupTransformationTypical data
CnC_nnn discrete planar rotationsImages with a finite set of orientations
DnD_nDiscrete rotations and reflectionsImages and regular polygons
SO(2)SO(2), SO(3)SO(3)Continuous rotations in 2D or 3DImages, molecules and point clouds
SE(d)SE(d)Rotations and translations in dd dimensionsRigid-body systems
E(d)E(d)Rotations, reflections and translationsPhysical and geometric data
SnS_nPermutations of nn elementsSets, graphs and particles

The correct group is determined by the task, not by the data type alone.

Invariance and equivariance

Let f:X→Yf:\mathcal X\rightarrow\mathcal Y be a model, with representations ρX\rho_X and ρY\rho_Y acting on the input and output spaces. The model is GG-equivariant if

f ⁣(ρX(g)x)=ρY(g)f(x)∀g∈G,  x∈X.f\!\left(\rho_X(g)x\right) = \rho_Y(g)f(x) \qquad \forall g\in G,\;x\in\mathcal X.

The equation says that transforming the input before applying the model is equivalent to applying the model first and then transforming its output; or, in other words the two operations commute.

The model is GG-invariant if

f ⁣(ρX(g)x)=f(x).f\!\left(\rho_X(g)x\right)=f(x).

Invariance is therefore a special case of equivariance in which the output carries the trivial representation:

ρY(g)=I∀g∈G.\rho_Y(g)=I \qquad \forall g\in G.
TaskDesired transformation law
Image classificationRotating the image should not change the class: invariant
Semantic segmentationRotating the image should rotate the pixel labels: equivariant
Object detectionThe category stays fixed while the bounding box transforms: mixed
Molecular energy predictionRigid motion should not change a scalar energy: E(d)E(d)-invariant
Molecular force predictionRotations should rotate force vectors: E(d)E(d)-equivariant
Set classificationReordering elements should not change the label: permutation-invariant
Node prediction on a graphReordering nodes should reorder node outputs: permutation-equivariant

This distinction matters because throwing away a transformation too early can be convenient but also destructive. In particular, equivariance retains the transformation in a controlled way; invariance removes it.

Orbits, stabilizers and what the model sees

For an input xx, its orbit is the set of all versions reachable through the group:

Ox={ρX(g)x:g∈G}.\mathcal O_x = \left\{\rho_X(g)x:g\in G\right\}.

An invariant model is constant across this orbit. An equivariant model is not constant, but its outputs across the orbit are completely determined by one prediction and the output action ρY\rho_Y.

Some transformations may leave xx unchanged. They form the stabilizer

Gx={g∈G:ρX(g)x=x}.G_x = \left\{g\in G:\rho_X(g)x=x\right\}.

Stabilizers are often trivial for an asymmetric object in a "generic" pose, but not for instance for a circle, sphere, regular polygon, etc. (no matter how you rotate a sphere, it looks exactly the same, so every rotation in SO(3)\mathrm{SO}(3) sits in its stabilizer). If ff is equivariant and g∈Gxg\in G_x, then

f(x)=f(ρX(g)x)=ρY(g)f(x).f(x) =f(\rho_X(g)x) =\rho_Y(g)f(x).

The output must preserve every symmetry of the input. In stabilizer notation,

Gx⊆Gf(x).G_x\subseteq G_{f(x)}.

This compact relation is the "neural-network version" of Curie's symmetry principle: a deterministic equivariant map cannot produce an output with less symmetry than its input. If a task requires one specific direction to be chosen from a perfectly symmetric object, the model needs extra information, randomness, or an explicit reference frame.

Why equivariance is useful

Equivariance is an inductive bias, meaning it restricts the functions a model is allowed to learn. That restriction can help in three ways.

  • First, it reduces sample complexity. If all rotated versions of a pattern must be processed consistently, the network does not need to learn each orientation. Parameters can be shared across the group orbit.

  • Second, it provides systematic generalization. Data augmentation encourages similar behaviour on transformed samples, but a mathematically equivariant architecture specifies such behaviour in its architecture.

  • Third, it can enforce physical consistency. For NN particles with positions X=(x1,…,xN)X=(x_1,\ldots,x_N), a scalar energy UU should satisfy U(RX+t)=U(X),U(RX+t)=U(X), while forces should satisfy

F(RX+t)=RF(X),R∈SO(3),  t∈R3.F(RX+t)=RF(X), \qquad R\in SO(3),\;t\in\mathbb R^3.

The coordinates frame is arbitrary, but physics is not. Encoding this distinction is particularly valuable in scientific machine learning, where a numerically plausible prediction can still violate a physical law or geometric constraint.

There is no universal benefit, however, since if a false symmetry is learned, then it produces a false constraint (a flipped 6 is a 9, yet is wrong invariance). Therefore, the important question is not whether a model should be equivariant, but to which group, at which layers, and for which outputs.

Four ways to learn a symmetry

Deep-learning systems usually obtain invariance or equivariance through augmentation, architectural design, group averaging, or canonicalization. These approaches encode the same prior with different guarantees and costs.

Data augmentation

For a loss ℓ\ell, group augmentation optimizes an orbit-averaged objective such as

Laug(θ)=E(x,y)∼DEg∼μG[ℓ ⁣(fθ(ρX(g)x),ρY(g)y)],\mathcal L_{\mathrm{aug}}(\theta) = \mathbb E_{(x,y)\sim\mathcal D} \mathbb E_{g\sim\mu_G} \left[ \ell\!\left(f_\theta(\rho_X(g)x),\rho_Y(g)y\right) \right],

where μG\mu_G is a distribution over transformations. This encourages the correct relation on sampled transformations, but it does not guarantee it. Moreover, the model can still behave inconsistently across training examples, and continuous groups have infinite augmented dataset.

Augmentation also changes the training distribution, which is harmless only when the augmented samples remain meaningful and aligned with the downstream task. For instance, training on 90-degree rotations of trains spends capacity on poses the model will almost never encounter while executing its tasks.

Architectural equivariance

An architecture can restrict every layer to commute with the group action. For a linear layer W:Vin→VoutW:V_{\mathrm{in}}\rightarrow V_{\mathrm{out}}, equivariance requires the relation

Wρin(g)=ρout(g)W∀g∈G.W\rho_{\mathrm{in}}(g) = \rho_{\mathrm{out}}(g)W \qquad \forall g\in G.

Where the permissible weights lie in the solution space of these linear constraints.

Ordinary convolution is the familiar example. If TaT_a translates a signal by aa and kk is a kernel, then

(k∗Tax)(u)=∫k(v)x(u−v−a) dv=Ta(k∗x)(u).(k*T_a x)(u) = \int k(v)x(u-v-a)\,dv = T_a(k*x)(u).

Convolution is therefore translation-equivariant. Group-equivariant CNNs replace translations with richer discrete groups, while steerable CNNs constrain kernels through group representations and irreducible components (Cohen and Welling, 2016; Weiler and Cesa, 2019). In three dimensions, Tensor Field Networks use spherical harmonics and representation coupling to obtain rotation and translation equivariance (Thomas et al., 2018).

This approach gives strong guarantees, but these specialized layers make existing architectures difficult to reuse. For a large pretrained backbone, this would likely mean traning a new model.

Group averaging

For a finite group and an arbitrary prediction model pp, the Reynolds operator constructs an invariant function

finv(x)=1∣G∣∑g∈Gp(ρX(g)x).f_{\mathrm{inv}}(x) = \frac{1}{|G|}\sum_{g\in G}p(\rho_X(g)x).

For equivariant outputs, each prediction must first be moved back into a common output frame:

feq(x)=1∣G∣∑g∈GρY(g−1)p(ρX(g)x).f_{\mathrm{eq}}(x) = \frac{1}{|G|} \sum_{g\in G} \rho_Y(g^{-1}) p(\rho_X(g)x).

This is the basis of frame averaging and Equi-Tuning. Under a squared-distance objective, the Reynolds operator can be seen as projecting a pretrained model onto the space of GG-equivariant functions (Puny et al., 2022; Basu et al., 2023).

The cost is direct: exact averaging requires ∣G∣|G| backbone evaluations. Continuous groups require integration with respect to the Haar measure,

feq(x)=∫GρY(g−1)p(ρX(g)x) dμ(g),f_{\mathrm{eq}}(x) = \int_G \rho_Y(g^{-1})p(\rho_X(g)x)\,d\mu(g),

which is usually approximated by Monte Carlo samples. For a large backbone, this makes inference extremely costly and lengthy.

Learned canonicalization

Canonicalization on the other hand selects a representative pose from each orbit. Let p:X→Yp:\mathcal X\rightarrow\mathcal Y be any prediction network and let

c:X→Gc:\mathcal X\rightarrow G

predict the transformation associated with an input. The canonical sample is

xcan=ρX(c(x)−1)x.x_{\mathrm{can}} = \rho_X(c(x)^{-1})x.

After prediction, an equivariant task restores the original output pose:

f(x)=ρY(c(x))p ⁣(ρX(c(x)−1)x).f(x) = \rho_Y(c(x)) p\!\left(\rho_X(c(x)^{-1})x\right).

For an invariant task, ρY\rho_Y is trivial and the final restoration disappears.

The construction becomes equivariant when the canonicalizer respects

c(ρX(g)x)=g c(x).c(\rho_X(g)x)=g\,c(x).

Indeed, the inverse c(ρX(g)x)−1=c(x)−1g−1c(\rho_X(g)x)^{-1}=c(x)^{-1}g^{-1} cancels the input transformation, while the output factor contributes ρY(g)\rho_Y(g). The main network pp itself does not need to be equivariant (Kaba et al., 2023).

EquiAdapt model functioning.

EquiAdapt model functioning.

The pretrained-model problem

Canonicalization appears to solve the architectural problem: attach a small network in front of a large pretrained backbone, transform the input into a canonical pose, and run the expensive model only once.

The difficulty is that an orbit has no inherently preferred "orientation/transformation". A canonicalizer may consistently rotate every upright image by 180∘180^\circ, even if it makes no sense, such as upside down cars. This is mathematically valid, but a backbone pretrained mainly on upright cars images may perform badly on the resulting distribution.

Naive learned canonicalization creates two related problems:

  • Alignment: the chosen canonical pose may not match the orientations preferred by the pretrained backbone.
  • Augmentation: early in training, a random canonicalizer can expose the backbone to extreme transformations that are unnecessary for the task.

EquiAdapt addresses both with a canonicalization prior. Let PD\mathbb P_{\mathcal D} denote the dataset's preferred distribution over GG and Pc(x)\mathbb P_{c(x)} the distribution predicted by the canonicalizer. The regularizer is

Lprior=Ex∼D[DKL(PD ∥ Pc(x))].\mathcal L_{\mathrm{prior}} = \mathbb E_{x\sim\mathcal D} \left[ D_{\mathrm{KL}} \left( \mathbb P_{\mathcal D} \,\|\, \mathbb P_{c(x)} \right) \right].

If the prior density q(R)q(R) is fixed and the canonicalizer predicts p(R∣c(x))p(R\mid c(x)), the entropy of the prior is constant. Minimizing the KL divergence is therefore equivalent to minimizing the cross-entropy

Lprior=−Ex∼DER∼q[log⁡p(R∣c(x))].\mathcal L_{\mathrm{prior}} = -\mathbb E_{x\sim\mathcal D} \mathbb E_{R\sim q} \left[ \log p(R\mid c(x)) \right].

For a discrete group such as CnC_n, an identity prior places all mass on ee:

q(g)=δg,e⟹Lprior=−Ex∼Dlog⁡pc(x)(e).q(g)=\delta_{g,e} \quad\Longrightarrow\quad \mathcal L_{\mathrm{prior}} = -\mathbb E_{x\sim\mathcal D} \log p_{c(x)}(e).

The model is not being told to ignore rotations. It is being encouraged to map the training distribution back to the identity pose, so the backbone continues to receive inputs resembling those on which it was pretrained.

For continuous rotations, EquiAdapt uses an isotropic matrix Fisher distribution on SO(n)SO(n):

p(R∣R^,s)=1n(s)exp⁡ ⁣(s Tr⁡(R^⊤R)),p(R\mid \widehat R,s) = \frac{1}{n(s)} \exp\!\left( s\,\operatorname{Tr}(\widehat R^\top R) \right),

where R^\widehat R is the mode and ss controls concentration. With a prior centred at the identity, the orientation-dependent part of the loss becomes

Lprior≡−λTr⁡(Rc(x))≡λ2∥Rc(x)−I∥F2,\mathcal L_{\mathrm{prior}} \equiv -\lambda\operatorname{Tr}(R_{c(x)}) \equiv \frac{\lambda}{2} \left\|R_{c(x)}-I\right\|_F^2,

where ≡\equiv denotes equality up to an additive constant. The trace is largest when the predicted rotation is closest to the identity.

In practice, discrete canonicalizers can output one logit per group element, apply a softmax, and select with an argmax using a straight-through gradient estimator. Continuous canonicalizers must output valid rotation matrices, often by predicting vector fields and orthonormalizing them. That additional geometry makes continuous optimization much less forgiving.

What the experiments show

The most revealing result is not simply that prior-regularized learned canonicalization improves rotated accuracy. It is that it recovers robustness without giving up most of the original-distribution performance.

On CIFAR10, the following models were fine-tuned from ImageNet-pretrained and evaluated both on the original test set and on a test set averaged over the eight rotations in C8C_8:

BackboneMethodOriginal accuracy ↑C8C_8-average accuracy ↑
ResNet50Vanilla96.97 ± 0.0157.77 ± 0.25
ResNet50Rotation augmentation94.91 ± 0.0790.11 ± 0.19
ResNet50Learned canonicalization93.29 ± 0.0192.96 ± 0.09
ResNet50C8C_8 augmentation95.76 ± 0.0794.36 ± 0.09
ResNet50Prior-regularized LC96.19 ± 0.0195.31 ± 0.17
ViTVanilla98.13 ± 0.0463.59 ± 0.48
ViTPrior-regularized LC96.14 ± 0.1495.08 ± 0.10

The vanilla models are accurate on familiar orientations and poor across the orbit. Naive canonicalization works but worse than prior-regularized one, which recovers much of the lost in-distribution accuracy while retaining transformation robustness (Mondal et al., 2023).

The canonical pose distribution supports the alignment explanation. On CIFAR10, the fraction of images mapped to the identity after training rose from 0.23 with naive learned canonicalization to 0.76 with prior regularization. The prior does not merely change the final classifier; it changes what the backbone is asked to see (Mondal et al., 2023).

The same trade-off appears in zero-shot instance segmentation on COCO:

BackboneCanonicalizerParametersOriginal mAP ↑C4C_4-average mAP ↑
Mask R-CNNNone0 M45.5727.67
Mask R-CNNSmall G-CNN0.2 M35.7735.77
Mask R-CNNG-WideResNet1.9 M44.5144.50
SAMNone0 M62.3458.78
SAMSmall G-CNN0.2 M59.2859.28
SAMG-WideResNet1.9 M62.1362.13

The equal original and rotated scores indicate near-perfect consistency, but it seems that the small canonicalizer becomes now the expressivity bottleneck, since increasing it from 0.2 million to 1.9 million parameters almost closes the Mask R-CNN gap. For SAM, the larger wrapper adds about 0.3% parameters and increases inference time by about 7.3%, far less than evaluating a 641-million-parameter backbone once per group element.

Point-cloud results show why this matters beyond images. Under arbitrary SO(3)SO(3) test rotations, a standard PointNet trained with only vertical-axis rotations falls from 85.9% to 19.6% ModelNet40 classification accuracy. Prior-regularized canonicalization with pretrained PointNet achieves 84.3 ± 1.2% without rotation augmentation. For DGCNN, the corresponding result is 90.2 ± 1.3%, compared with 33.8% for the conventionally trained model (Mondal et al., 2023).

Exact symmetry is not always available

Canonicalization has a structural ambiguity as seen before whenever we are dealing with a non-trivial stabilizer, since several group elements produce the same observed sample.

This motivates relaxed equivariance. Rather than requiring one exact transformation, the output may be correct up to the stabilizer coset:

∀g1∈G,  x∈X,∃g2∈g1Gxsuch thatf(ρX(g1)x)=ρY(g2)f(x).\forall g_1\in G,\;x\in\mathcal X, \quad \exists g_2\in g_1G_x \quad\text{such that}\quad f(\rho_X(g_1)x) = \rho_Y(g_2)f(x).

The distinction is not cosmetic, because for a symmetric input, demanding a unique canonical direction can create discontinuities.

Continuous groups add further difficulties:

  • Valid outputs must remain on a manifold such as SO(3)SO(3) rather than in unconstrained Euclidean space.
  • Estimating or differentiating probability densities on the group can be expensive and numerically unstable.
  • Monte Carlo group averaging introduces variance, while discretization replaces continuous symmetry with a finite approximation.

The original EquiAdapt experiments reported difficulty optimizing the continuous-rotation prior for images, even though the analogous construction was theoretically sounding and worked well for 3D point clouds.

Can canonicalization avoid equivariant networks?

EquiAdapt moves most equivariance constraints out of the backbone, but it does not make them disappear, we still need an equivariant canonicalizer, therefore this creates a circular-looking dependency, where to avoid a large equivariant architecture, we are obliged to build an additional equivariant network.

EquiOptAdapt reframes canonicalization as optimization over the orbit. Let u:X→Ru:\mathcal X\rightarrow\mathbb R be an ordinary, non-equivariant scoring network. Select the canonical transformation through

g∗(x)∈arg min⁡g∈Gu ⁣(ρX(g−1)x).g^*(x) \in \operatorname*{arg\,min}_{g\in G} u\!\left(\rho_X(g^{-1})x\right).

Every transformed input generates the same set of orbit scores, only re-indexed by the group action. If the minimum is unique, the selected group element changes predictably with the input; if it is not unique, the set of minimizers reflects the stabilizer.

The practical version learns embeddings sθs_\theta for transformed inputs and converts them to a canonicalization distribution. With reference vector vRv_R and temperature τ\tau,

Pc(g∣x)=exp⁡ ⁣(vR⊤sθ(ρX(g−1)x)/τ)∑g′∈Gexp⁡ ⁣(vR⊤sθ(ρX(g′−1)x)/τ).P_c(g\mid x) = \frac{ \exp\!\left(v_R^\top s_\theta(\rho_X(g^{-1})x)/\tau\right) }{ \sum_{g'\in G} \exp\!\left(v_R^\top s_\theta(\rho_X(g'^{-1})x)/\tau\right) }.

An identity prior again aligns the chosen pose with the pretrained backbone:

Lprior=−Ex∼Dflog⁡Pc(e∣x).\mathcal L_{\mathrm{prior}} = -\mathbb E_{x\sim\mathcal D_f} \log P_c(e\mid x).

The method learns canonical orientations with a non-equivariant scoring network and was reported to converge faster than EquiAdapt on the tested C4C_4 setting. Its evidence remains narrower, since the published experiments focus only on C4C_4 (Panigrahi and Mondal, 2024).

Measuring equivariance

Task accuracy does not measure equivariance, since a model can achieve similar average accuracy across rotations while producing inconsistent predictions for individual samples.

A direct normalized equivariance error is

εG(f)=Ex∼D, g∼μG[∥f(ρX(g)x)−ρY(g)f(x)∥2∥ρY(g)f(x)∥2+δ],\varepsilon_G(f) = \mathbb E_{x\sim\mathcal D,\,g\sim\mu_G} \left[ \frac{ \left\| f(\rho_X(g)x)-\rho_Y(g)f(x) \right\|_2 }{ \left\|\rho_Y(g)f(x)\right\|_2+\delta } \right],

where δ>0\delta>0 prevents division by zero. Exact equivariance gives εG(f)=0\varepsilon_G(f)=0 up to numerical and interpolation error.

A serious evaluation should report at least three quantities:

  1. Task performance on the original distribution.
  2. Task performance across transformed or out-of-distribution samples.
  3. Equivariance error comparing the two computational paths directly.

Where equivariance matters most

Equivariance is most valuable when the transformation law is known independently of the dataset, meaning mostly in core scientific fields:

  • Scientific machine learning: energies, forces, velocities and fields have precise behaviour under transformation, therefore E(n)E(n)-equivariant graph networks use these laws directly in message passing (Satorras et al., 2021)
  • Robotics and perception: camera pose changes coordinates, not what we are seen, so predictions must remain consistent across viewpoints and reference frames
  • Medical imaging: orientation and acquisition geometry may vary, while anatomical structures should transform predictably
  • Point clouds and sets: the ordering of points is arbitrary, so making permutation invariance or equivariance is a basic requirement
  • Pretrained-model adaptation: a small canonicalization module can add a missing symmetry without rebuilding an expensive backbone

Conclusion

Invariance means that a transformation should not affect the answer. Equivariance means that it should affect the answer in a known way.

That single distinction turns symmetry into a practical design principle. It explains weight sharing in convolutions, the consistency of geometric networks, and the possibility of adapting a large pretrained model without redesigning it. EquiAdapt also exposes the central complication: a mathematically valid coordinate system may still be the wrong one for the model behind it.

The goal is therefore not to make neural networks ignore geometry. It is to make their dependence on geometry explicit.

Useful Readings

References