> > > MSP: Learning Manipulation-Sufficient Representations via Outcome Bottlenecks

Learning Manipulation-Sufficient Representations via Outcome Bottlenecks

Md Selim Sarowar, Sungho Kim

Corresponding author. Yeungnam University.

Under review · IEEE International Conference on Robotics and Automation (ICRA'27)

Two learned modules. No reconstruction, no dynamics rollout, no reinforcement learning.

The MSP framework: a hidden world state produces RGB-D observations from networked cameras; a belief encoder returns a posterior over a 64-dimensional code; an outcome head maps code and action to success, margin and slip; an inference block marginalises, certifies by split conformal, acts risk-aversely, senses a new view and adapts after a probe
Perception estimates a belief over a code, not a pose. The world state x = (S, T, φ) collects shape, pose and physics, and only the first two are visible in an image at all. A per-view backbone pools into a posterior qθ(z | o) trained by a variational information bottleneck: minimise the rate R while paying a weighted distortion on the outcomes that decide what happens. Every deployment behaviour, certified abstention, risk-averse action, active sensing and post-probe adaptation, is an inference procedure over the same two frozen modules. Nothing in the objective asks for a reconstruction or a pose, and nothing retrains at deployment.
0.000
Analytic grasp score on curved objects, below chance
0.000
MSP on the same objects
0.000
Empirical coverage at a 0.90 target, n = 16,000
0.00 nats
Rate the code costs on the link, per scene

Abstract

Networked manipulation endpoints couple perception to actuation across compute- and bandwidth-limited links, yet commonly exchange dense geometric states optimized for fidelity rather than action outcomes. We learn a stochastic representation with a policy-free, action-conditioned outcome bottleneck: the marginal outcome log-loss supplies the distortion and a KL term regularizes the rate. The construction is motivated by the minimal statistic that preserves the outcome distribution of every admissible action, while the implemented finite model is evaluated as a rate-regularized mixture predictor. One encoder and one outcome head then support grasp selection, singleton conformal filtering, active viewpoint selection and latent test-time adaptation: no reconstruction, no dynamics rollout, no reinforcement learning. A finite-probe theorem identifies the local level-set tangent space with the null space of an outcome Jacobian; the synthetic oracle verifies it, and on scanned objects an analytic surrogate agrees with measured simulator invariances to within 0.56° of principal angle. Across 11,979 simulated grasps on 13 scanned LIBERO groceries, a reconstructed-geometry wrench score reaches 0.542 AUC against lift success and falls below chance on curved objects, while our representation reaches 0.876. At 25% commitment, executed-grasp success is 0.503 against 0.984. The entropy-coded payload is 10 bytes per scene and runs at 16 ms per CPU decision. On unseen objects within-scene AUC falls to 0.569, and a full-feedback update raises empirical mean pairwise coverage from 0.728 to 0.883.

Four Stages, Two Modules

Everything at deployment is inference over the same two frozen modules. Pick a stage.

What the Theory Buys

Four claims, and the limit of each one stated next to it.

The statement

Under the Markov chain ZOXY, the data processing inequality bounds what any code can carry. The minimal sufficient statistic is the information-bottleneck optimum in the limit of a large sufficiency budget, so the code that survives compression is exactly the one that changes outcomes.

What follows practically

There is no reconstruction loss and no pose loss anywhere in the objective. The rate is the number of nats the belief costs to transmit, which on the network edge is the quantity that actually bills.

The budget is not uniform across outcome channels. Success, margin and slip have log-likelihoods on incommensurate scales, and putting equal weight on them starves the one channel the decision rule reads. The ablation section shows what that costs: the model stops ranking grasps entirely.

The statement

States are identifiable only up to the kernel of the outcome Jacobian. The predicted subspace is compared against one measured independently, by perturbing the state and recording which directions leave real outcomes unchanged. Where the theorem has content the prediction is essentially exact, with principal angles under one degree.

Where it says nothing

For 6 of 13 objects the Jacobian has full rank and the kernel is trivial: nothing is continuously unidentifiable and the theorem makes no claim at all. That is a real limit on its reach, and the object table below marks it per object rather than reporting an average that hides it.

The statement

Split conformal prediction over a held-out calibration fold gives marginal coverage without any distributional assumption. Actions whose prediction set is the singleton {1} form the certified set; an empty set is an abstention.

The estimand trap

Coverage is Pr(succ in the set) over all test points. Certified precision is Pr(succ = 1 | the action was certified). They are different quantities and reporting the second while claiming the first is the standard way a conformal result gets overclaimed. Both are on this page.

Second views help

Acquiring one more viewpoint cuts ambiguity by 13.8% on average. Choosing which viewpoint, selected on one Monte-Carlo estimate of the information gain and scored on an independent one, adds 2.9 points for 16.7% in total, over 512 scenes and 8 candidate views.

The number we refuse to quote

Selecting and scoring on the same estimate reports 130.8%, of which 114.1 points are the winner's curse. We report the honest estimate and print the inflated one beside it so the size of the selection bias in this kind of experiment is visible.

Most of the benefit is in taking a second view at all rather than in choosing it well, and we say so rather than attributing the whole reduction to the policy.

The Proxy Fails Where Its Geometry Is Wrong

Ferrari-Canny epsilon is computed on an oriented bounding box, roughly the geometry a pose-and-shape pipeline recovers. The outcome is a rigid-body lift against the object's true mesh.

The natural control group. Six of the thirteen groceries have a single-box collision hull. Butter, cookies and cream cheese are boxes, so the bounding-box reconstruction is exact for them and there is no hallucinated surface. The remaining seven are cans, bottles and cartons. The thesis predicts the analytic proxy works where its geometry is right and fails where it is not, and it does. On a parametric box this experiment is impossible, because reconstruction and truth are the same object and the gap is identically zero. That is why real meshes are not a nicety.
Bar chart: on box-shaped objects Ferrari-Canny reaches AUC 0.574 and MSP 0.797; on curved objects Ferrari-Canny falls to 0.473, below the chance line, while MSP rises to 0.891
AUC against the real lift outcome, split by whether the reconstruction is exact. On curved objects the analytic proxy scores below chance: it is anti-correlated with the truth, systematically preferring the grasps that fail. 7,194 grasps on box-shaped objects and 4,785 on curved ones.
11,979 executable grasps on 13 scanned LIBERO groceries, success rate 0.597
Predictor Pooled AUC Within-scene AUC
Ferrari-Canny ε on the reconstruction0.5420.548
MSP, belief trained on outcomes Ours0.8760.639
Chance0.5000.500
Oracle: an MLP on the true state0.685

Read the right column. Within-scene AUC ranks the candidate grasps of a single settled pose, so it cannot be won by recognising the object. The pooled column can be: per-object base rates on this corpus span 0.043 to 0.999, and a model that has learned nothing but object identity still scores about 0.72 pooled. The oracle bound matters too: no perception system can beat 0.685 within-scene here, so MSP's 0.639 is against a ceiling, not against 1.0.

When It Commits, Does the Object Come Up?

The system ranks a scene's candidate grasps, executes its favourite, and acts only when confident enough. Sweeping that confidence threshold sweeps the act rate. Drag it.

Act rate0.25
MSP
0.984
of executed grasps lifted the object
Ferrari-Canny
0.503
of executed grasps lifted the object
Random pick
0.623
the control that does no ranking at all

1,997 scenes with at least one executable grasp. A useful score rises as it grows more selective; the analytic proxy's falls, and only reaches the random-pick control by committing to nearly every scene. It is not mis-calibrated, which a monotone transform would fix. It is uninformative, and on the objects whose geometry it gets wrong it is worse than that.

The Certificate Holds on Images

Split conformal calibration on a held-out fold, evaluated over 16,000 test points on real RGB-D. Pick the risk level.

Target coverage
0.90
Empirical coverage
0.8973
Certified precision
0.7435
Abstention rate
0.7430
Certified action fraction
0.2347
A high abstention rate is not a bug. From a photograph you cannot see friction, mass or the centre of mass, so on most scenes no action can be certified at the requested confidence and the system declines rather than guessing. That is the framework being honest about what an image contains. Bringing the number down is exactly what active touch is for, and the tables above show what a second viewpoint already buys.
Two panels: empirical coverage tracks the nominal target along the diagonal at 0.80, 0.90 and 0.95; and the cost of the certificate, where both certified precision and abstention rate rise with the target
Coverage tracks its target, and the certificate has a price. Tightening the risk level raises certified precision and raises the abstention rate with it: the system buys confidence by acting less often, which is the trade a deployed manipulator should be allowed to choose.

What the Link Has to Carry

The sufficiency budget β multiplies the relevance term, so a larger β purchases sufficiency at the price of rate. Hover a point for its coverage and abstention.

The purchased outcome is success: Dsucc falls monotonically as the rate rises, from 0.656 at 0.004 nats to 0.412 at 3.29 nats. The unweighted total over all three channels need not, because it is dominated by the margin and slip likelihoods the budget deliberately sacrifices, and reporting only that total would make a working budget look broken. Coverage holds near 0.90 across the whole sweep, so the certificate is not being paid for out of the rate.

Why an edge system cares

0.53 nats per scene is what the belief costs to send

The rate is not a regulariser we chose for convenience: it is the mutual information between the observation and the code, which on a networked manipulator is the payload. A pipeline that ships a reconstructed mesh or a dense pose distribution ships orders of magnitude more, and the evidence above says the extra bits do not change the decision. Measured as a payload rather than a code length, the belief quantised to two bits per dimension and entropy-coded is 10 bytes per scene — measured, not an fp32 storage figure — which is 14,750× smaller than the 147.5 kB of one RGB-D frame. Ten bytes cost nothing that matters: within-scene AUC 0.636 against 0.638 at fp32, coverage 0.906 against 0.904. One bit per dimension is the cliff, at 0.584.

Because the statistic is formed before a grasp is chosen, one payload scores every candidate. A retrained Dex-Net grasp-quality network forms its representation after the grasp is applied, so it must send one payload per candidate. Sweeping the same quantiser over both traces a rate–quality frontier that crosses near 600 bits per scene: below it ours ranks better on eight times fewer bits; above it the baseline is the stronger ranker.

Loss behaves the way a certificate should. As packet loss rises from 0 to 20 % the system acts on 0.628 down to 0.507 of scenes while success given action stays flat between 0.852 and 0.860. A degrading link costs throughput, not reliability.

The cheap end of the frontier

At 0.004 nats the certificate still covers 0.896

The lowest-rate point on the sweep holds its coverage target and abstains slightly less than the deployed point. Rate and calibration are close to independent here, which is the property that makes the interface deployable on a constrained link rather than only in a lab.

Thirteen Objects, One Table

Base success rates span 0.043 to 0.999, which is the entire reason a pooled AUC cannot be trusted here. Filter by geometry; click a header to sort.

13
Objects in view
0.000
Ferrari-Canny, mean over rankable objects
0.000
MSP, mean over rankable objects

Object Base rate Ferrari-Canny MSP rank J dim ker J Max angle °

The two AUC columns are within-scene, averaged over the scenes where the ranking question is well posed, and the three figures above them are macro-means over the objects in view, which is a different aggregation from the corpus-level 0.548 and 0.639 quoted earlier: a scene needs at least three executable grasps and both a success and a failure. An object lifted successfully 99.9% of the time supplies almost none, so it is marked not rankable rather than given an AUC over two or three scenes that would read as a catastrophic failure. The last three columns are the identifiability diagnostic: where the kernel of the outcome Jacobian is non-trivial the predicted unidentifiable subspace matches the independently measured one to under a degree, and where the Jacobian has full rank the theorem makes no claim and the cell is a dash.

Per-object chart: grey bars give base success rates from 0.04 to 0.999, with Ferrari-Canny and MSP within-scene AUC markers above them; three objects are labelled no rankable scenes
Base rates span the whole range, which is why a pooled AUC cannot be trusted. Objects with too few rankable scenes get their base-rate bar and no marker, which is the honest thing to draw: the question cannot be asked of them.

What Each Piece Is Worth

Scored on the within-scene axis, because a pooled AUC cannot referee these: a model that has learned nothing but object identity still scores about 0.72 pooled, so every ablation would look harmless.

The dotted line is chance. Response is the standard deviation of the score across a scene's candidate grasps: when it collapses towards zero the model has stopped attending to the action and is emitting one number per scene.

The budget is load-bearing

A uniform budget destroys the ranker, and a pooled AUC would hide it

Placing an equal multiplier on outcome channels whose log-likelihoods live on incommensurate scales starves success, the one outcome the decision rule reads. Within-scene AUC falls to 0.511 and the action response collapses to 0.002: the model has stopped ranking grasps entirely. Its pooled AUC is still 0.710, which is why the axis matters.

Perception is doing the work

Replacing the belief with a constant costs everything

With the code ablated the head can only exploit an action prior, and it lands at 0.506 within-scene, chance. This is the control for the reviewer concern that the head could ignore the belief and memorise action priors, and it does not survive.

Evidence Boundaries

What the measurements do not cover, stated rather than left for a reader to discover.

The identifiability theorem is often vacuous

For 6 of 13 objects the outcome Jacobian has full rank and its kernel is trivial, so nothing is continuously unidentifiable and the theorem says nothing at all. Where it does have content the prediction matches an independently measured subspace to under a degree, but the reach is genuinely limited and the per-object table shows it case by case.

Ranking transfers to unseen objects, but less than half of it

Held out object by object, within-scene AUC drops from the in-distribution 0.639 to 0.569 on the seven objects whose base rates make the ranking question well posed. It stays above chance, so the representation does carry grasp-ranking structure to geometry it has never seen, but the claim that it transfers intact is not supported at this scope.

A held-out object breaks the fixed quantile until feedback arrives

Calibrated on twelve objects and deployed on a thirteenth, the fixed quantile undercovers badly: mean pairwise coverage 0.728, worst 0.533 on ketchup. An anchored full-feedback update lifts it to 0.883 and compresses dispersion from 0.110 to 0.032, worst case 0.780. This is a clustered empirical diagnostic under simulated shift, not a selective-feedback validity claim.

Three objects cannot be scored at all

Milk, orange juice and bbq sauce are lifted so reliably that almost no scene contains both a success and a failure, so the within-scene ranking question is not well posed for them. They are marked, not imputed.

The certificate is marginal, not conditional

Coverage holds over the test distribution as a whole. It is not a per-object or per-scene guarantee, and certified precision, the number an operator actually feels, is a different estimand that we report separately rather than in place of it.

Abstention is high at the deployed point

0.743 at α = 0.1. The framework declines on most scenes because a photograph does not contain friction or mass. Active touch is the intended remedy and the equations for it exist, but the touch loop is not wired into the deployed system.

Active perception is evaluated in simulation over 8 views

512 scenes, 8 candidate viewpoints, ambiguity reduction rather than end-to-end success. The honest estimate is 16.7% and most of that, 13.8 points, comes from taking a second view at all rather than from choosing it.

One corpus, one gripper, no real robot

Thirteen scanned LIBERO groceries under a rigid-body lift criterion. Real meshes rather than parametric primitives is what makes the reconstruction-gap experiment possible at all, but nothing here measures sim-to-real transfer of the belief or the certificate.

Bar chart of within-scene AUC: Ferrari-Canny on the reconstruction at 0.548 and MSP on outcomes at 0.639, with a chance line at 0.5
The headline contrast on the axis that cannot be gamed. Drawn within-scene, never pooled: MSP once scored 0.716 pooled while being unable to rank two grasps of the same object, which is the measurement error this axis exists to prevent.

BibTeX

@inproceedings{sarowar2027msp,
title     = {Learning Manipulation-Sufficient Representations via Outcome Bottlenecks},
author    = {Md Selim Sarowar and Sungho Kim},
booktitle = {IEEE International Conference on Robotics and Automation (ICRA)},
year      = {2027},
note      = {Under review},
url       = {https://physical-agi.github.io/MSP-FRAMEWORK/}
}

Placeholder entry. It will be replaced with the conference entry once the paper is accepted.