Every simulator is, to some degree, an abstraction of the real system: it models part of the real state and leaves the rest out. Most sim2real methods, including domain randomization, system identification and learned dynamics corrections, treat the gap as a mismatch in parameters over a state space the simulator and the robot share. This paper studies the extreme end of the abstraction spectrum: an abstract simulator reduced to the most basic information about the task, such as a 2D point mass with position and velocity standing in for a NAO humanoid with a 30-dimensional state. Such simulators are cheap to build and fast to run, but they leave out most of the dynamics the policy will face on the robot.
Abstraction makes the transfer problem partially observable. The simulator's state is the real state passed through a many-to-one map, so real states with different dynamics share one abstract state, and the abstract state is no longer Markov. The key to transfer is to condition both the simulator correction and the policy on the history of abstract states and actions. ASTRA (Augmented Simulation with self-predicTive abstRAction) learns such a history representation from a small amount of real-world data with self-predictive losses, uses it to ground the abstract simulator, and transfers policies from a point-mass simulator to a physical NAO.
Write \(s^t\) for the state of the real (target) robot, \(a\) for an action, and \(\phi\) for the known abstraction that maps a real state to the simulator's abstract state, so that \(\bar s_i=\phi(s^t_i)\) is the abstract state of the real robot at step \(i\); for the NAO in the maze, its planar position and velocity.
Paired data. Collect trajectories on the real robot. For each real transition \((s^t_i,a_i,s^t_{i+1})\), reset the abstract simulator to \(\bar s_i\) and execute the same action \(a_i\). The simulator's next state \(s^{\mathrm{sim}}_{i+1}\) is its prediction of the real outcome \(\bar s_{i+1}\), and the difference between the two is what grounding has to correct.
History encoder. A GRU \(\psi^s\) reads the abstract history \(h_i=(\bar s_0,a_0,\dots,\bar s_i)\) and outputs a latent state \(z_i=\psi^s(h_i)\). A single abstract state cannot distinguish a biped in a stable stance from one that has become unsteady after a sharp turn; the history can.
Latent dynamics model. \(P_{\mathrm{lat}}(\cdot\mid z,a)\) predicts a Gaussian over the next latent. Training it with the encoder pushes the latent toward being Markov: the next latent depends on the past only through the current one.
Reward predictor. \(R_{\mathrm{pred}}(z,a)\) predicts the real reward. It keeps the latent tied to what the task depends on; in the paper's ablation, removing its loss costs more than removing the latent dynamics loss.
State corrector. \(f_{\mathrm{abs}}(z,a,s^{\mathrm{sim}})\) takes the simulator's prediction and outputs a corrected next abstract state, trained to match the real outcome. This is the grounding signal. The three heads sit on the GRU encoder and are trained jointly with it, one loss per head: \(\mathcal L_{\mathrm{trans}}\) for the latent dynamics model, \(\mathcal L_{\mathrm{rew}}\) for the reward predictor and \(\mathcal L_{\mathrm{abs}}\) for the state corrector.
Policy training. During RL, every simulator step is followed by a correction: the policy acts on \(z\), the abstract simulator steps, \(f_{\mathrm{abs}}\) corrects the result, the simulator is set to the corrected state, and that state extends the history that \(\psi^s\) reads. The reward comes from \(R_{\mathrm{pred}}\). Any RL algorithm can be used; the paper uses PPO and SAC.
Deployment. The policy takes \(z\) as input, so on the robot a target encoder \(\psi^t\) maps the real-state history \((s^t_0,a_0,\dots,s^t_i)\) into the same latent space. It is trained, with the other components frozen, so that its latent transitions match those of \(\psi^s\) under maximum mean discrepancy, and the policy runs on its output.
Write the real robot as the target MDP \(\mathcal M_t=(\mathcal S^t,\mathcal A,P_t,r,\gamma)\) and the abstract simulator as \(\mathcal M_s=(\mathcal S^s,\mathcal A,P_s,r_s,\gamma)\): the two share the action space \(\mathcal A\) and the discount \(\gamma\in[0,1)\), and each has its own state space \(\mathcal S\), transition kernel \(P\) and reward, with \(\phi:\mathcal S^t\to\mathcal S^s\). The goal is a policy trained by RL in \(\mathcal M_s\) that maximizes the expected discounted return in \(\mathcal M_t\).
Abstraction is harmless when \(\phi\) is model-irrelevant: states it merges have the same reward and the same distribution over the next abstract state under every action, and the abstract states then form an MDP. A point mass standing in for a biped is far from model-irrelevant. The dropped joint state changes how position and velocity evolve, and the abstract states of a real trajectory are not Markov:
$$\Pr(\bar s_{i+1}\mid \bar s_i,a_i)\neq\Pr(\bar s_{i+1}\mid \bar s_i,a_i,\bar s_{i-1},a_{i-1}).$$
Standard grounding fits the simulator's \(P_s(\bar s_{i+1}\mid \bar s_i,a_i)\) to real data, which is then ill-defined: real states with different dynamics share \(\bar s_i\), so there is no single next-state distribution to fit.
History repairs this. A latent \(z_i=\psi^s(h_i)\) is an approximate information state (AIS) if it predicts the real reward \(r_i\) and its own next value, with \(\mathbb E\) the expectation over the real robot's behavior:
$$\textcolor{#8A4B1F}{\mathbb E[r_i\mid h_i,a_i]\approx R_{\mathrm{pred}}(z_i,a_i)},$$
$$\textcolor{#2F5D8A}{\Pr(z_{i+1}\mid h_i,a_i)\approx P_{\mathrm{lat}}(z_{i+1}\mid z_i,a_i)}.$$
This is model-irrelevance restated over histories: histories that share a latent share the reward and the distribution of the next latent. ASTRA turns the two conditions into losses and adds the grounding loss, summing over the steps \(i\) of the paired data:
$$\textcolor{#8A4B1F}{\mathcal L_{\mathrm{rew}}}=\sum_i\big(R_{\mathrm{pred}}(z_i,a_i)-r_i\big)^2,$$
$$\textcolor{#2F5D8A}{\mathcal L_{\mathrm{trans}}}=-\sum_i\log\mathcal N\big(z_{i+1}\mid\mu_i,\sigma_i^2\big),$$
$$\textcolor{#3F6B3A}{\mathcal L_{\mathrm{abs}}}=\sum_i\big\|f_{\mathrm{abs}}(z_i,a_i,s^{\mathrm{sim}}_{i+1})-\bar s_{i+1}\big\|^2,$$
where \(\mathcal N(\mu_i,\sigma_i^2)\) is the Gaussian that \(P_{\mathrm{lat}}\) outputs for \((z_i,a_i)\). The full objective is the weighted sum \(\lambda_1\mathcal L_{\mathrm{trans}}+\lambda_2\mathcal L_{\mathrm{rew}}+\lambda_3\mathcal L_{\mathrm{abs}}\).
In both real-robot tasks, a NAO runs a policy trained in a 2D abstract simulator. In maze navigation the simulator is a point mass with velocity control; in ball kicking it is a point agent with a limited field of view and simplified ball physics. ASTRA reaches 73% success on navigation, against 17–53% for all baselines, and 56% on ball kicking, against 5–40%.
@article{deng2026abstract,
title={Abstract sim2real through approximate information states},
author={Deng, Yunfu and Li, Yuhao and Hanna, Josiah P},
journal={IEEE Robotics and Automation Letters},
year={2026},
publisher={IEEE}
}