Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance

Fabian A. Mikulasch1, Friedemann Zenke1,2
1 Friedrich Miescher Institute for Biomedical Research, Basel, Switzerland
2 Faculty of Science, University of Basel, Switzerland
arXiv preprint · September 2026

A general framework for world model learning

Predictive self-supervised learning (SSL) learns useful representations by predicting in latent space, but why? We prove two fundamental principles of predictive SSL:

  1. Predictive SSL can recover the true latent variables of the data, even if these variables have stochastic, nonlinear dynamics.
  2. This recovery works even if the data contains strong, unpredictable nuisance

The framework that allows for these proofs, which we term InfoLDM, is a generalization of existing SSL approaches.

InfoLDM maps observations with multiple possible futures into latent space, combining a predictive model with representation entropy.
InfoLDM maximizes the goal function $\mathcal F$: The likelihood of a predictive model constrains the latent distribution, while entropy keeps representations informative.

Like joint-embedding predictive architectures (JEPAs), InfoLDM predicts encoded observations without reconstructing the input. In fact, by choosing a specific latent predictive model and entropy estimator we can re-derive popular JEPA approaches:

Gaussian model + logdet entropy → VICReg-style
Choosing the latent model $p_\theta$ to be a Gaussian with predictor $\mu$ and estimating entropy through the log-determinant of representation covariance $\Sigma$ connects the objective to variance and covariance regularization. \[ p_\theta(z|z_c) = \frac{1}{Z} \exp (- \|z - \mu(z_c)\|^2 ) \] \[ H_{q_f}[z] \approx \text{log det}(\Sigma) \] \[ \Rightarrow \] \[ \mathcal{F} \propto - \left\langle\|z - \mu(z_c)\|^2\right\rangle_{q_f(z,z_c)} + \text{log det}(\Sigma) \]
von Mises-Fisher model + KDE entropy → SimCLR-style
Kernel density entropy estimation connects the objective to contrastive learning, such as SimCLR. \[ p_\theta(z|z_c) = \frac{1}{Z} \exp ( z^T \mu(z_c) ) \] \[ H_{q_f}[z] \approx - \sum_i \log \sum_j \exp ( z_j^T z_i ) \] \[ \Rightarrow \] \[ \mathcal{F} \propto \sum_i \log \frac{\exp ( z_i^T \mu({z_c}_i) )}{\sum_j \exp ( z_j^T z_i )} \]

Before discussing the formal framework and proofs in more detail, let's get an intuitive understanding of the results.

Identifying nonlinear dynamics under nuisance

We show that InfoLDM allows to find excellent representations and a predictive latent model for nonlinear, action dependent systems—even if the observations are corrupted by high dimensional nuisance. As an example we use the MuJoCo hopper, with the goal to recover physical hopper states (joint angles, positions, ...) from generated videos.

Clean reference animation of the MuJoCo Hopper moving.
Clean hopper environment (only for reference)
The Hopper with changing body colors and a multicolored noisy background, as observed by the model.
Observations with nuisance

InfoLDM receives the corrupted videos as input and learns a latent model. As comparison we also train generative models that reconstruct the input (additionally to latent prediction). Only the non-generative models find meaningful representations.

UMAP projections of physical variables, image data, generative representations, and latent predictive representations, colored by physical state.
Representation geometry. The latent predictive model recovers structure related to the physical state that is obscured in the observed images.

We can quantify this with linear readout probes.

Linear readout R squared for Hopper height, torso orientation, and joint states. Latent predictive representations show high recovery compared with generative and next-observation baselines.
Recovering the physical state. Linear probes can read out height, torso orientation, and joint states from the learned representations. We also compare with 'next observation' models, which aim to predict the next image.

Because generative models try to predict or encode every detail in the images, which is too challenging, they consistently fail. InfoLDM directly learns the latents dynamics of the nonlinear system and recovers the underlying variables.

We can also use the learned dynamics to simulate the system: we provide observations for 5 timesteps and query the latent model to generate future trajectories of the underlying physical variables.

Animated rollout of the Hopper system using the learned latent dynamics.
Simulating the system. Rolling out the learned latent dynamics lets us simulate how the physical state evolves over time. The model was trained on 8-step sequences and generalizes beyond that. Shaded areas denote 65% sampling intervals. Uncertainty accumulates over time.

Because InfoLDM is based on a probabilistic latent model it is straightforward to get calibrated uncertainty estimates, for example by sampling from it. This uncertainty combines effects of stochastic environment dynamics, observation nuisance, and a potentially imperfect model.

Distinguishing nuisance from signal stochasticity

Recovery under nuisance poses a conundrum: What is considered nuisance and what is only stochasticity in an actually relevant signal, and can the model distinguish those? Our formal results make this precise: Complete unpredictablity → nuisance; predictable signal with added, structured noise → recovered by the algorithm.

To show this we construct pairs of images in which selected factors are causally related, with stochastic variation between the two images. The remaining factors vary independently and act as nuisance.

In the object-signal setting shown below, object position, rotation, and hue are coupled across image pairs. Lighting and background are unrelated. InfoLDM recovers the object (signal) variables such that they can be linearly read out from the representation, while background (nuisance) variables are ignored.

Stochastic image pairs with correlated object position, rotation, and color, but uncorrelated lighting and background. Linear probes recover signal factors strongly and nuisance factors weakly.
Stochastic does not mean nuisance. The model recovers causally related factors even though the relation is stochastic.

The Signal is defined by the relationship between observations. In a second setting, we couple the environment factors instead of object factors, and InfoLDM recovers the former but not the latter.

InfoLDM: what is recovered, and how

How does InfoLDM allow us to prove this recovery under nuisance? Here is a short overview of the formal framework:

Graphical model showing context, signal, and nuisance generating observations, with encoders mapping observations and context into a predictive latent model.
Generative and learned models. Observations $x$ combine signal and nuisance; encoders map observations and context into representations connected by a predictive latent model.

In summary, $c$, $s$ and $n$ denote true latent context, signal, and nuisance variables; $z = f(x)$ is the encoded observation, and $z_c = f_c(x_c)$ is the encoded context. In the simulations above the context is the history and actions (hopper), or the related image, but it can be any relevant encoded conditioning variable. The encoders induce a joint distribution $q_f(z,z_c)$, while a learned latent model $p_\theta(z|z_c)$ describes that distribution. InfoLDM maximizes:

\[ \begin{aligned} \mathcal{F} &= \underbrace{- D_{\mathrm{KL}}[q_f(z,z_c)\,\|\,p_\theta(z,z_c)] }_{\text{LDM}} + \underbrace{\vphantom{ D_{\mathrm{KL}}[q_f(z|z_c)\,\|\,p_\theta(z|z_c)] }I_{q_f}[z; z_c]}_{\text{MI}} \\ &= \textcolor{#c32d76}{\left\langle \log p_\theta(z,z_c)\right\rangle_{q_f(z,z_c)}} + \textcolor{#625bc0}{H_{q_f}[z]} + \textcolor{#625bc0}{H_{q_f}[z_c]} \end{aligned} \]

Provable identification of the signal variable

We show that maximizing the InfoLDM goal function allows to recover the signal variable. Under the paper’s assumptions and with a Gaussian predictor, the true signal can be recovered from the learned representation by an affine map:

\[ s = Az + b \]

More generally, we prove this result for exponential family predictive models, where the recovery guarantee is on the sufficient statistics of the distribution.

Assumptions behind the guarantee
  1. Private nuisance: $n \perp\!\!\!\perp c \mid s$ — nuisance adds no context information given the signal.
  2. Exponential-family prediction: $p(s\mid c)$ and $p_\theta(z\mid z_c)$ are exponential-family conditionals with sufficient statistics $\tau_\star(s)$ and $\tau_z(z)$.
  3. Ideal optimum: $D_{\mathrm{KL}}(q_f\|p_\theta)=0$ and $I_{q_f}[z;z_c]=I[s;c]<\infty$ — exact matching and full predictive information.
  4. Context variation: $\operatorname{span}\{\alpha(c)-\alpha(c_0)\}=\mathbb{R}^{r_s}$ — natural parameters vary in every signal-statistic direction.
  5. Nondegenerate statistics: $\tau_\star(s)\mid c$ lies in no proper affine hyperplane — no redundant statistic dimensions.

At the optimum of InfoLDM:

\[ \tau_\star(s) = A\tau_z(z) + b \]

The proof shows that the LDM and MI terms serve two complimentary functions:

MI: what is retained

Maximizing predictive mutual information preserves the information shared with the context. At saturation, all prediction-relevant signal information is retained.

This alone does not identify the representation’s coordinates: invertible nonlinear transformations preserve mutual information.

LDM: how it is represented

Latent distribution matching constrains the encoded distribution to follow the learned model. These constraints determine how the retained information is represented.

We show: For exponential-family predictive distributions, InfoLDM identifies the true signal’s sufficient statistics through an affine readout of the learned statistics.

For the full proof and more details, read the paper posted on the arxiv.

BibTeX

@misc{mikulasch2026predictive,
  title={Predictive Self-Supervised Learning Provably Identifies
         Stochastic Signals under Nuisance},
  author={Fabian A. Mikulasch and Friedemann Zenke},
  year={2026},
  eprint={2609.37789},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2609.37789}
}