BIFROST: Bridging Invariant Feature Representation for Observation-space Sim2Real Transfer

Yunfu Deng, Josiah Hanna
University of Wisconsin–Madison
IEEE/RSJ International Conference on Intelligent Robots and Systems, 2026

Overview

The basis for attempting sim2real at all is the assumption that simulation and reality pose the same task. Despite differing in observations and dynamics, the two domains share a common low-dimensional structure, analogous to low-rank MDPs, which admit a compact representation (Agarwal et al.) that related tasks can share (Cheng et al.). Existing methods typically do not exploit this structure. They address each gap in isolation, with system identification and dynamics randomization for the dynamics gap and visual randomization and domain adaptation for the perceptual gap, and must compose or layer these modules when both gaps coexist.

BIFROST learns the shared structure directly from raw observations. It identifies corresponding simulated and real states by behavioral equivalence: two states from different domains are equivalent when acting identically from both leads to equivalent long-term outcomes, regardless of domain-specific differences in rendering or physics. This criterion, cross-domain bisimulation, operates on the behavioral consequences of both gaps rather than on an explicit model of each gap source, which turns both gaps into a single representation learning problem.

A shared encoder maps each domain's history, the observations and actions so far, into one latent space, and is trained so that behaviorally equivalent histories are mapped close together; the policy is then trained on this latent space. Sharing the encoder is not enough on its own: co-training on pooled data from both domains does not guarantee that equivalent histories are mapped nearby when the visual and dynamics gaps are large, so BIFROST enforces the alignment explicitly.

Method

Paired data. Collect 200 real trajectories by teleoperation and cut them into short segments. For each, reset the simulator to the configuration in the segment's first real frame (joint angles and object poses) and replay the same actions. The result is a sim history and a real history that share every action, step for step, so any divergence between them comes from differences in rendering and physics rather than in task-relevant state.

Shared history encoder. A single GRU \(\phi\) reads the history \(h_t=(o_0,a_0,\dots,o_t)\), every observation \(o\) (camera image, plus proprioception when available) and action \(a\) so far, from either domain, and outputs a latent state \(z_t=\phi(h_t)\). One frame cannot show velocity, friction or contact; a history can.

Reward predictor. \(\hat R(z,a)\) predicts the real robot's reward from either domain's latent; a sim latent is scored against the reward of its paired real step. This anchors the latent to the real task, keeps the encoder from overfitting to simulation, and later gives the policy its reward.

Latent dynamics model. \(\hat P(\cdot\mid z,a)\) predicts a Gaussian over the next latent, trained on each domain's own transitions. It makes the latent capture how the scene evolves, and gives the alignment a predicted future to compare.

Bisimulation alignment. For every pair, the dynamics model predicts where the sim history and the real history go next under the same action, and a Wasserstein-1 loss pulls the two predictions together, spread as well as mean. Histories that lead to the same future end up in the same place, whatever domain they came from.

Policy. With the encoder and the reward predictor frozen, SAC trains a policy \(\pi(a\mid z)\) in simulation on latent states, with \(\hat R\) as its reward. On the robot, the same frozen encoder feeds it real histories, and the policy runs as trained, with no fine-tuning.

BIFROST in three stages: pair real and simulated trajectories; train a shared history encoder with cross-domain bisimulation alignment, then train a policy on its latent states in the simulator; deploy the frozen encoder and the policy on the real robot.
Pair real and simulated trajectories, align their histories in one latent space, train the policy there in simulation, and deploy it on the robot unchanged.

Theoretical motivation

The reward predictor and the alignment each target one term of a cross-domain bisimulation metric. Write the simulator as the source MDP \(\mathcal M_{\mathrm{src}}=(\mathcal S_{\mathrm{src}},\mathcal A,P_{\mathrm{src}},R_{\mathrm{src}},\gamma)\) and the real robot as the target MDP \(\mathcal M_{\mathrm{tgt}}=(\mathcal S_{\mathrm{tgt}},\mathcal A,P_{\mathrm{tgt}},R_{\mathrm{tgt}},\gamma)\): the two share the action space \(\mathcal A\) and the discount \(\gamma\in[0,1)\), and each has its own states \(\mathcal S\), transition kernel \(P(\cdot\mid s,a)\) and reward \(R(s,a)\) for a state \(s\) and an action \(a\in\mathcal A\). For a simulated state \(s_1\in\mathcal S_{\mathrm{src}}\) and a real state \(s_2\in\mathcal S_{\mathrm{tgt}}\), the generalized bisimulation metric (GBSM) of Tao et al. is the distance \(d\) defined recursively by

$$d(s_1,s_2)=\max_{a\in\mathcal A}\Big\{\textcolor{#8A4B1F}{\underbrace{\big|R_{\mathrm{src}}(s_1,a)-R_{\mathrm{tgt}}(s_2,a)\big|}_{\text{reward predictor}}}+\gamma\,\textcolor{#2F5D8A}{\underbrace{W_1\big(P_{\mathrm{src}}(\cdot\mid s_1,a),\,P_{\mathrm{tgt}}(\cdot\mid s_2,a);\,d\big)}_{\text{bisimulation alignment}}}\Big\},$$

where \(W_1(\mu,\nu;d)\) is the Wasserstein-1 distance: the least expected cost of transporting a distribution \(\mu\) onto a distribution \(\nu\) when moving a unit of mass from \(x\) to \(y\) costs \(d(x,y)\). With \(V^*_{\mathrm{src}}(s)\) and \(V^*_{\mathrm{tgt}}(s)\) the optimal value of \(s\) in its own domain, the metric bounds the value gap between domains:

$$\big|V^*_{\mathrm{src}}(s_1)-V^*_{\mathrm{tgt}}(s_2)\big|\le d(s_1,s_2).$$

States close under \(d\) are worth nearly the same in sim and in real, so a policy learned on sim latents carries over to real ones. BIFROST drives a latent version of both terms to zero on every sim–real pair, with histories in place of states, which is what lets it work from raw pixels. For a paired sim history \(h^{\mathrm{src}}\) and real history \(h^{\mathrm{tgt}}\) with next action \(a\) and real reward \(r^{\mathrm{tgt}}\), and \(\mathbb E\) averaging over pairs,

$$\textcolor{#8A4B1F}{\mathcal L_{\mathrm{reward}}}=\mathbb E\Big[\big(\hat R(\phi(h^{\mathrm{src}}),a)-r^{\mathrm{tgt}}\big)^2+\big(\hat R(\phi(h^{\mathrm{tgt}}),a)-r^{\mathrm{tgt}}\big)^2\Big],$$

$$\textcolor{#2F5D8A}{\mathcal L_{\mathrm{align}}}=\mathbb E\Big[W_1\big(\hat P(\cdot\mid\phi(h^{\mathrm{src}}),a),\,\hat P(\cdot\mid\phi(h^{\mathrm{tgt}}),a)\big)\Big],$$

with \(W_1\) here using Euclidean distance between latents as its transport cost. A latent dynamics prediction loss \(\mathcal L_{\mathrm{LDP}}\) trains \(\hat P\) on each domain's own next latent, and the encoder, reward predictor and dynamics model are trained together on \(\mathcal L_{\mathrm{reward}}+\lambda_{\mathrm{LDP}}\mathcal L_{\mathrm{LDP}}+\lambda_{\mathrm{align}}\mathcal L_{\mathrm{align}}\), with weights \(\lambda\).

Latent space analysis

A Hello Robot Stretch navigates a maze from its onboard camera, trained in SAPIEN and deployed in MuJoCo, two simulators that render the scene and simulate its physics differently. Pretrained ImageNet features split SAPIEN (source) from MuJoCo (target). BIFROST's latents mix the two domains and are organized by goal, the task-relevant variable.

Three t-SNE plots. ResNet-18 ImageNet features form separate clusters for source and target. BIFROST latents mix source and target, and group by goal, with source and target points of the same goal overlapping.
t-SNE on egocentric navigation. (a) ImageNet ResNet-18 image features separate by domain. (b) BIFROST history latents mix the domains. (c) The same latents colored by goal: source and target points of each goal share a cluster. Goal 1 splits into two, one per corridor on either side of a wall.

Citation

@inproceedings{deng2026bifrost,
  title={BIFROST: Bridging Invariant Feature Representation for Observation-space Sim2Real Transfer},
  author={Deng, Yunfu and Hanna, Josiah P},
  booktitle={2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year={2026},
  organization={IEEE}
}