Files
Mujoco_WASM/src/engine/engine_inverse.c
T
Alessio Quaglino ea230a950c Implicit flex elasticity in the CG constraint solver via an effective metric
This CL replaces the post-hoc implicit flex correction (`flexInterp_cgsolve`) with a **linearly-implicit effective metric** `M̃ = M + (h² + h·damping)·K` carried by the CG constraint solver itself. Contact/friction forces and implicit flex elasticity are now computed against one consistent metric, instead of the solver seeing `M` and a post-solve correction changing `qacc` behind its back.

Gate (unchanged semantics): `solver="CG"` + implicit/implicitfast integrator + pyramidal cones + flex stiffness present. Newton and PGS are untouched. `solver="CG"` remains the user-facing contract — the factorization is an implementation detail of the preconditioner.

### What's in the metric

- **mjData `efm_*`** (arena, efc-like lifetime/skip semantics; built in `mj_fwdPosition`, value-refreshed in `mj_fwdVelocity`): the per-step stiffness CSR `efm_B_*`, its reverse-Cholesky factor `efm_dofid` + `efm_L_*` (nested-dissection ordered, separators-first for the reverse factorization), and the smooth-force shift `efm_c = h·K·qvel`.
- **`mjd_flexStiff_assemble`** now assembles stretch (Gauss–Newton), standard dim-2 bending, and — via the cached corotated stiffness `d->flexelem_krot` — interp stiffness (all node bodies on simple sliders: point Jacobian is I₃, `flex_centered` not required; fixed nodes drop like pins) into one dof-level CSR. `mjd_effMulAdd`/`mjd_effSolve` apply the metric, with matrix-free operator fallbacks where assembly does not apply.
- **mjModel `efm0_*`** (`nefm0dof`/`nefm0L`): the constant part of the metric factor — currently the dim-2 bending factor, computed once in `mj_setConst` — so bending-only models pay zero per-step factorization cost. Naming mirrors mjData's `efm_*` with the standard `0`-suffix (reference/constant) idiom, and is deliberately not bending-specific: future constant contributors extend it without renames.
- The solver consumes the metric through pre-shifted `qfrc_smooth` and the metric products `Ma`/`Mv`/`Mgrad`; `qacc_smooth` becomes the unconstrained minimizer of the implicit dynamics, which makes the no-constraint shortcut and the warmstart choice consistent by construction.
- **`mj_inverse` adds `B·qacc − c`**, making inverse dynamics discrete-consistent with the gated forward dynamics — exact, since the gated path has no qDeriv term (new test `ForwardTest.GatedFlexInverseConsistency`).

### Performance

All numbers: ms/step over the same 2000-step window, models as shipped on each side (old code with the old model settings vs this CL with the new ones).

The new solver path activates on exactly two shipped models — the ponchos, the only flex models that need an implicit integrator (poncho on Euler degenerates to >200 ms/step). For them, this CL trades speed for consistency: the implicit bending solve now runs inside every solver iteration, where the contact solve can see the stiffness, instead of once after the solve. Solver iterations drop because the curvature is visible, but each iteration pays for the implicit solve:

| model | before | after | solver iters/step |
|---|---|---|---|
| poncho | 2.47 | 3.30 (1.33×) | 16.8 → 11.8 |
| poncho_edgeequality | 1.96 | 2.72 (1.39×) | 13.2 → 10.0 |

What that price buys: contact forces consistent with the implicit elasticity (previously the post-hoc correction changed `qacc` after the constraint solve), discrete-consistent inverse dynamics, and the removal of the post-hoc special case from the integration path. Raising poncho's timestep from 2 to 5 ms leaves its per-step cost nearly flat, so the consistency price can be recovered by taking fewer steps where accuracy allows.

Every other flex model was measured stable on Euler at its shipped timestep and switches to it (these models predate the post-hoc integrator; implicit was never load-bearing for them). They end up equal or faster than before: bunny_multicell 0.47 → 0.40, trampoline 0.28 → 0.25, plate 1.02 → 0.99, pancake 0.34 → 0.33.

Finally, the per-step factorization makes configurations practical that the old code could only integrate explicitly: implicit stretch elasticity (`elastic2d="stretch"`/`"both"`, dim-3 solids) and factorized interp stiffness. No before/after exists for these — stock has no implicit treatment of stretch at all.

### Behavior changes

- With the post-hoc correction deleted, interp/bending models running `solver="Newton"` (or elliptic cones, or islands) now integrate flex elasticity **explicitly** (previously: post-hoc implicit). Affects e.g. `gripper_trilinear` (stable, and faster, but different semantics). Follow-up options: Newton-side metric support, or a documented fallback.
- With the gate on, `mj_forward` outputs are timestep-dependent for gated models (they answer the linearly-implicit discrete problem); `qacc_smooth` and `mj_inverse` change accordingly. Non-gated models are bit-identical (full suite green throughout).

### Validation

- 1737/1737 tests, including new: `FlexStretchDerivatives` (FD-validated GN operator), `FlexStiffAssemble`/`FlexStiffAssembleInterp` (CSR ≡ operators), `GatedFlexInverseConsistency` (fails pre-change), equivalence tests vs the old post-hoc treatment (bending matches to 2e-11).
- Fingerprint discipline throughout: bending-only models bit-exact across every refactor; permutation/kernel changes verified iteration-identical.

### Known follow-ups (not in this CL)

3×3-block sparse Cholesky kernel (the numeric factorization is index-bound; projected ~3× on the factor); mjModel persistence of the factor's symbolic pattern (rest-pose ND makes sizes compile-time); the general effective-metric mode (all solvers, all PSD-safe force classes, behind an enable flag).

PiperOrigin-RevId: 948561856
Change-Id: I8b8e32ebd0428042af71647d0470d10773bf6daf
2026-07-15 14:57:42 -07:00

343 lines
9.5 KiB
C

// Copyright 2021 DeepMind Technologies Limited
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
#include "engine/engine_inverse.h"
#include <stddef.h>
#include <mujoco/mjdata.h>
#include <mujoco/mjmacro.h>
#include <mujoco/mjmodel.h>
#include <mujoco/mjsan.h> // IWYU pragma: keep
#include "engine/engine_collision_driver.h"
#include "engine/engine_core_constraint.h"
#include "engine/engine_core_smooth.h"
#include "engine/engine_core_util.h"
#include "engine/engine_derivative.h"
#include "engine/engine_memory.h"
#include "engine/engine_macro.h"
#include "engine/engine_forward.h"
#include "engine/engine_sensor.h"
#include "engine/engine_support.h"
#include "engine/engine_util_blas.h"
#include "engine/engine_util_errmem.h"
#include "engine/engine_util_misc.h"
#include "engine/engine_util_sparse.h"
// position-dependent computations
void mj_invPosition(const mjModel* m, mjData* d) {
TM_START1;
TM_START;
// clear flag for lazy evaluation
d->flg_energypos = 0;
mj_kinematics(m, d);
mj_comPos(m, d);
mj_camlight(m, d);
mj_flex(m, d);
mj_tendon(m, d);
TM_END(mjTIMER_POS_KINEMATICS);
mj_makeM(m, d); // timed internally (POS_INERTIA)
mj_factorM(m, d); // timed internally (POS_INERTIA)
mj_collision(m, d); // timed internally (POS_COLLISION)
TM_RESTART;
mj_makeConstraint(m, d);
TM_END(mjTIMER_POS_MAKE);
// compute exact diagonal if enabled
if (mjENABLED(mjENBL_DIAGEXACT)) {
TM_RESTART;
mj_projectConstraint(m, d);
TM_END(mjTIMER_POS_PROJECT);
}
TM_RESTART;
mj_transmission(m, d);
TM_ADD(mjTIMER_POS_KINEMATICS);
// implicit effective metric: multiply-only build (no factorization) for the inverse
mjd_effBuild(m, d, mj_flexCG(m), /*flg_factor=*/0);
TM_END1(mjTIMER_POSITION);
}
// velocity-dependent computations
void mj_invVelocity(const mjModel* m, mjData* d) {
mj_fwdVelocity(m, d);
}
// convert discrete-time qacc to continuous-time qacc
static void mj_discreteAcc(const mjModel* m, mjData* d) {
int nv = m->nv, nC = m->nC, nD = m->nD, dof_damping;
mjtNum *qacc = d->qacc;
mj_markStack(d);
mjtNum* qfrc = mjSTACKALLOC(d, nv, mjtNum);
// use selected integrator
switch ((mjtIntegrator) m->opt.integrator) {
case mjINT_RK4:
// not supported by RK4
mjERROR("discrete inverse dynamics is not supported by RK4 integrator");
return;
case mjINT_EULER:
// check for dof damping if disable flag is not set
dof_damping = 0;
if (!mjDISABLED(mjDSBL_EULERDAMP)) {
for (int i=0; i < nv; i++) {
if (m->dof_damping[i] > 0 ||
!mju_isZero(m->dof_dampingpoly + mjNPOLY*i, mjNPOLY) ||
m->jnt_actuatorid[m->dof_jntid[i]] != -1) {
dof_damping = 1;
break;
}
}
}
// if disabled or no dof damping, nothing to do
if (!dof_damping) {
mj_freeStack(d);
return;
}
// set qfrc = (M + h*diag(B)) * qacc
mj_mulM(m, d, qfrc, qacc);
for (int i=0; i < nv; i++) {
mjtNum v = d->qvel[i];
mjtNum poly[mjNPOLY];
mju_copy(poly, m->dof_dampingpoly + mjNPOLY*i, mjNPOLY);
mjtNum damping = m->dof_damping[i]
+ mj_actuatorDamping(m, mjOBJ_JOINT, m->dof_jntid[i], poly);
mjtNum damp_deriv = mjd_xPolyForce(damping, poly, v, mjNPOLY, 1);
qfrc[i] += m->opt.timestep * damp_deriv * d->qacc[i];
}
break;
case mjINT_IMPLICIT:
// compute qDeriv
mjd_smooth_vel(m, d, /* flg_bias = */ 1);
// gather qLU <- M (lower to full)
mju_gatherMasked(d->qLU, d->M, m->mapM2D, nD);
// set qLU = M - dt*qDeriv
mju_addToScl(d->qLU, d->qDeriv, -m->opt.timestep, m->nD);
// set qfrc = qLU * qacc
mju_mulMatVecSparse(qfrc, d->qLU, qacc, nv,
m->D_rownnz, m->D_rowadr, m->D_colind, /*rowsuper=*/NULL);
break;
case mjINT_IMPLICITFAST:
// compute analytical derivative qDeriv; skip rne derivative
mjd_smooth_vel(m, d, /* flg_bias = */ 0);
// save mass matrix
mjtNum* Msave = mjSTACKALLOC(d, m->nC, mjtNum);
mju_copy(Msave, d->M, m->nC);
// modified mass matrix: gather qH <- qDeriv (full to lower)
mju_gather(d->qH, d->qDeriv, m->mapD2M, nC);
// set qH = M - dt*qDeriv
mju_addScl(d->qH, d->M, d->qH, -m->opt.timestep, nC);
// set qfrc = (M - dt*qDeriv) * qacc
mju_mulSymVecSparse(qfrc, d->qH, qacc, m->nv, m->M_rownnz, m->M_rowadr, m->M_colind);
// standalone free bodies: overwrite block rows with the unsymmetric local product,
// including the bias (gyroscopic) derivative, mirroring mj_implicitSkip
for (int j=0; j < m->njnt; j++) {
mjtNum A[36];
if (!mjd_freeMhat(m, d, j, m->opt.timestep, A)) {
continue;
}
int adr = m->jnt_dofadr[j];
mju_mulMatVec(qfrc+adr, A, qacc+adr, 6, 6);
}
break;
}
// solve for qacc: qfrc = M * qacc
mj_solveM(m, d, qacc, qfrc, 1);
mj_freeStack(d);
// refresh the effective-metric velocity shift
mjd_effShift(m, d);
}
// inverse constraint solver
void mj_invConstraint(const mjModel* m, mjData* d) {
TM_START;
int nefc = d->nefc;
// no constraints: clear, return
if (!nefc) {
mju_zero(d->qfrc_constraint, m->nv);
TM_END(mjTIMER_CONSTRAINT);
return;
}
mj_markStack(d);
mjtNum* jar = mjSTACKALLOC(d, nefc, mjtNum);
// compute jar = Jac*qacc - aref
mj_mulJacVec(m, d, jar, d->qacc);
mju_subFrom(jar, d->efc_aref, nefc);
// call update function
mj_constraintUpdate(m, d, jar, NULL, 0);
mj_freeStack(d);
TM_END(mjTIMER_CONSTRAINT);
}
// inverse dynamics with skip; skipstage is mjtStage
void mj_inverseSkip(const mjModel* m, mjData* d,
int skipstage, int skipsensor) {
TM_START;
mj_markStack(d);
mjtNum* qacc;
int nv = m->nv;
// position-dependent
if (skipstage < mjSTAGE_POS) {
mj_invPosition(m, d);
if (!skipsensor) {
mj_sensorPos(m, d);
}
if (mjENABLED(mjENBL_ENERGY) && !d->flg_energypos) {
mj_energyPos(m, d);
}
}
// velocity-dependent
if (skipstage < mjSTAGE_VEL) {
mj_invVelocity(m, d);
if (!skipsensor) {
mj_sensorVel(m, d);
}
if (mjENABLED(mjENBL_ENERGY) && !d->flg_energyvel) {
mj_energyVel(m, d);
}
}
if (mjENABLED(mjENBL_INVDISCRETE)) {
// save current qacc
qacc = mjSTACKALLOC(d, nv, mjtNum);
mju_copy(qacc, d->qacc, nv);
// modify qacc in-place
mj_discreteAcc(m, d);
}
// acceleration-dependent
mj_invConstraint(m, d);
// sum of bias forces in qfrc_inverse = centripetal + Coriolis + tendon bias
mj_rne(m, d, 0, d->qfrc_inverse);
mj_tendonBias(m, d, d->qfrc_inverse);
if (!skipsensor) {
d->flg_rnepost = 0; // clear flag for lazy evaluation
mj_sensorAcc(m, d);
}
// compute Ma = M*qacc
mjtNum* Ma = mjSTACKALLOC(d, nv, mjtNum);
mj_mulM(m, d, Ma, d->qacc);
// implicit effective metric (built in mj_invPosition): the forward dynamics solved
// (M+K)*qacc = qfrc + c + J'*f, so the discrete-consistent inverse adds K*qacc - c
if (d->efm_active) {
mjd_effMulAdd(m, d, Ma, d->qacc);
mju_subFrom(Ma, d->efm_c, nv);
}
// qfrc_inverse += Ma - qfrc_passive - qfrc_constraint
for (int i=0; i < nv; i++) {
d->qfrc_inverse[i] += Ma[i] - d->qfrc_passive[i] - d->qfrc_constraint[i];
}
if (mjENABLED(mjENBL_INVDISCRETE)) {
// restore qacc
mju_copy(d->qacc, qacc, nv);
}
mj_freeStack(d);
TM_END(mjTIMER_INVERSE);
}
// inverse dynamics
void mj_inverse(const mjModel* m, mjData* d) {
mj_inverseSkip(m, d, mjSTAGE_NONE, 0);
}
// compare forward and inverse dynamics, without changing results of forward
// fwdinv[0] = norm(qfrc_constraint(forward) - qfrc_constraint(inverse))
// fwdinv[1] = norm(qfrc_applied(forward) - qfrc_inverse)
void mj_compareFwdInv(const mjModel* m, mjData* d) {
int nv = m->nv, nefc = d->nefc;
mjtNum *qforce, *dif, *save_qfrc_constraint, *save_efc_force;
// clear result, return if no constraints
d->solver_fwdinv[0] = d->solver_fwdinv[1] = 0;
if (!nefc) {
return;
}
// allocate
mj_markStack(d);
qforce = mjSTACKALLOC(d, nv, mjtNum);
dif = mjSTACKALLOC(d, nv, mjtNum);
save_qfrc_constraint = mjSTACKALLOC(d, nv, mjtNum);
save_efc_force = mjSTACKALLOC(d, nefc, mjtNum);
// qforce = qfrc_applied + J'*xfrc_applied + qfrc_actuator
// should equal result of inverse dynamics
mju_add(qforce, d->qfrc_applied, d->qfrc_actuator, nv);
mj_xfrcAccumulate(m, d, qforce);
// save forward dynamics results that are about to be modified
mju_copy(save_qfrc_constraint, d->qfrc_constraint, nv);
mju_copy(save_efc_force, d->efc_force, nefc);
// run inverse dynamics, do not update position and velocity,
mj_inverseSkip(m, d, mjSTAGE_VEL, 1); // 1: do not recompute sensors and energy
// compute statistics
mju_sub(dif, save_qfrc_constraint, d->qfrc_constraint, nv);
d->solver_fwdinv[0] = mju_norm(dif, nv);
mju_sub(dif, qforce, d->qfrc_inverse, nv);
d->solver_fwdinv[1] = mju_norm(dif, nv);
// restore forward dynamics results
mju_copy(d->qfrc_constraint, save_qfrc_constraint, nv);
mju_copy(d->efc_force, save_efc_force, nefc);
mj_freeStack(d);
}