ORIGAMI: Object Representation Inferred Geometrically for Articulated Manipulation

Yunfu Deng1, Daniel Nikovski2
1University of Wisconsin–Madison    2Mitsubishi Electric Research Laboratories
IEEE/RSJ International Conference on Intelligent Robots and Systems, 2026
ORIGAMI overview: from an RGB-D demonstration, keypoint tracking, joint estimation and topology inference produce a compact joint-position state for a reinforcement learning policy, deployed on a real robot.

Geometric representation as policy input

What a policy is given as input decides how much of its sample budget goes to control and how much to perception. Raw pixels put everything on the policy; a learned encoder moves the burden into pretraining or a joint objective, and what it keeps is whatever that objective happens to reward. For physical objects there is a third choice. A rigid body is a geometric invariant, and an articulated object is a few rigid bodies coupled by low-dimensional joints; those constraints hold regardless of appearance, in simulation and on the real thing. A representation built from them is low-dimensional by construction, physically meaningful in every coordinate, and, once the keypoints are tracked, independent of how the object looks. ORIGAMI's claim is that this is the better policy input: constructed from the object's kinematics rather than fitted to pixels, it lets a standard RL algorithm approach the sample efficiency of ground-truth state and transfer to hardware without adaptation.

Links from rigidity

Two points on the same rigid link keep a constant distance in every frame; a pair that straddles a joint does not:

$$\mathcal C(i)=\mathcal C(j)\;\Longleftrightarrow\;\|\mathbf p_{i,t}-\mathbf p_{j,t}\|=\mathrm{const}\quad\forall\,t.$$

From one RGB-D demonstration, keypoints are sampled inside a SAM mask, tracked with CoTracker3 and lifted to 3D with the depth map. For every pair, the variability of its distance over the demonstration, measured by the median absolute deviation \(m_{ij}\), becomes an affinity \(W_{ij}=\exp(-m_{ij}^2/\gamma^2)\). Spectral clustering on that graph gives a first partition, with the number of clusters read off the largest gap in the Laplacian spectrum; clusters that turn out to move as one rigid body under the joint fit below are merged, and what remains are the \(L\) links.

frame t frame t′ ABC ABC same link: AB never changes across the joint: AC changes
The one constraint the segmentation rests on. Grouping keypoints by how little their pairwise distances vary separates the links.
Two matrices over tracked keypoints of the folding ruler: the pairwise distance range, with dark blocks for rigid pairs, and the affinity matrix used for spectral clustering, with three clear blocks.
The same constraint measured on the folding ruler. Left: how much each pair's distance varies over the demonstration; dark blocks are rigid. Right: the affinity fed to spectral clustering, which recovers the three links.

Joints from relative motion

Given the links, each joint is read off the relative rigid motion between two of them. For link \(B\), the transform that aligns its keypoints at frame \(t\) to a reference frame \(t_0\) is the closed-form least-squares solution of Kabsch and Umeyama. With the cross-covariance over link \(B\)'s keypoints, \(\mathbf H=\sum_{i\in\mathcal I_B}(\mathbf p_{i,t_0}-\bar{\mathbf p}^{B}_{t_0})(\mathbf p_{i,t}-\bar{\mathbf p}^{B}_{t})^{\top}=\mathbf U\boldsymbol\Sigma\mathbf V^{\top}\),

$$\mathbf R_t^{B}=\mathbf U\,\mathrm{diag}\bigl(1,1,\det(\mathbf U\mathbf V^{\top})\bigr)\,\mathbf V^{\top},\qquad \mathbf t_t^{B}=\bar{\mathbf p}_{t_0}^{B}-\mathbf R_t^{B}\,\bar{\mathbf p}_{t}^{B},$$

and the motion of \(B\) relative to \(A\) is \(\mathbf R_t^{AB}=(\mathbf R_t^{A})^{-1}\mathbf R_t^{B}\), \(\mathbf t_t^{AB}=(\mathbf R_t^{A})^{-1}(\mathbf t_t^{B}-\mathbf t_t^{A})\), each solve wrapped in RANSAC against mis-segmented points near the boundary.

A one-degree-of-freedom joint leaves a signature in this sequence. A revolute joint accumulates rotation, \(\Theta=\sum_t|\theta_t-\theta_{t-1}|\); a prismatic joint accumulates translation, \(\Delta=\sum_t\|\mathbf t^{AB}_t-\mathbf t^{AB}_{t-1}\|\), with almost no rotation. Thresholding the ratio \(\Theta/\Delta\) classifies the type. For a revolute joint the axis direction \(\mathbf a_e\) is the dominant singular vector of the per-frame rotation axes, and the axis position \(\mathbf o_e\) is the fixed point of every observed rotation, the least-squares solution of

$$(\mathbf R_t^{AB}-\mathbf I)\,\mathbf o_e=-\mathbf t_t^{AB},\qquad t=1,\dots,T.$$

For a prismatic joint, \(\mathbf a_e\) is the dominant singular vector of the translations \(\{\mathbf t_t^{AB}\}\). Which links are actually connected follows from how well each pair fits a single such joint: adjacent pairs leave a small residual, non-adjacent pairs compound several joints and do not, and the minimum spanning tree over the residuals is the kinematic tree.

revoluteprismatic θt oe, ae ⊙ RAB turns about a fixed point; tAB = (I − RAB) oe ae RAB = I; tAB slides along ae
What the relative motion between two links looks like for the two joint types. Rotation without translation of a fixed point identifies a revolute joint and its axis; translation along one direction without rotation identifies a prismatic joint. The residual of each pair against these one-joint models decides which links are adjacent.

A state RL can use

At deployment the same machinery runs on every frame: keypoints are assigned to links, \(\mathbf R_t^{AB}\) is recovered by Kabsch–Umeyama for every joint, and the joint position is its projection onto the discovered axis,

$$q_e(t)=\operatorname{atan2}\bigl((\mathbf a_e\times\mathbf v)\cdot\mathbf v',\;\mathbf v\cdot\mathbf v'\bigr),\qquad \mathbf v'=\mathbf R_t^{AB}\mathbf v,$$

for a revolute joint, with \(\mathbf v\perp\mathbf a_e\) fixed at discovery time, and \(q_e(t)=\mathbf t_t^{AB}\cdot\mathbf a_e\) for a prismatic one. The vector \(\hat{\mathbf q}\in\mathbb R^{L-1}\) replaces the image, and a goal-conditioned policy sees the displacement \(\Delta\hat{\mathbf q}=\hat{\mathbf q}_{\mathrm{goal}}-\hat{\mathbf q}_{\mathrm{curr}}\). The estimator's errors come mostly from fixed properties of the estimator, the discovered axes, the segmentation and tracker bias, rather than from the configuration, so they cancel to first order:

$$\Delta\hat{\mathbf q}=\Delta\mathbf q^{*}+\boldsymbol\epsilon(\mathbf q^{*}_{\mathrm{goal}})-\boldsymbol\epsilon(\mathbf q^{*}_{\mathrm{curr}})\approx\Delta\mathbf q^{*}.$$

The raw estimates go to the policy as they are, with no calibration of their bias and no learned correction. With this input a standard RL algorithm approaches the sample efficiency of ground-truth state, and because keypoint geometry looks the same in rendered and real images, the policy transfers to a real arm zero-shot.

Citation

@inproceedings{deng2026origami,
  title={ORIGAMI: Object Representation Inferred Geometrically for Articulated Manipulation},
  author={Deng, Yunfu and Nikovski, Daniel},
  booktitle={2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year={2026},
  organization={IEEE}
}