Optimize mj_flex performance.
Two performance optimizations for interpolated flex objects: 1. Hoist loop-invariant stride calculations in `mju_cellLookup` out of nested loops. This reduces multiplications from 16 to 6 (linear) and 54 to 12 (quadratic) per vertex. 2. Extract and consolidate optimized 3D interpolation logic into a new reusable utility `mju_evalBasisArray` in `engine_util_misc.c`. This function uses nested loops and precomputed 1D shape functions to avoid expensive dynamic `phi` calls and branching, and leverages stack-buffered outputs to eliminate compiler pointer-aliasing barriers. We propagate this optimization to both kinematics (`mju_interpolate3D`) and constraint setup (`engine_core_constraint.c`). Together these changes yield a ~50% overall speedup in `mj_fwdKinematics` for interpolated flexes with ~10k vertices in the collision meshes and ~10 nodes in the deformation grid. PiperOrigin-RevId: 922198093 Change-Id: I9c804e65cd532e2f9ac6e9a71d01c30cd64186fa
This commit is contained in:
committed by
Copybara-Service
parent
af4be63cd8
commit
f6c85287bf
@@ -89,6 +89,9 @@ MJAPI void mju_defGradient(mjtNum res[9], const mjtNum p[3], const mjtNum* dof,
|
||||
// evaluate the basis function at x for the i-th node
|
||||
MJAPI mjtNum mju_evalBasis(const mjtNum x[3], int i, int order);
|
||||
|
||||
// evaluate the basis functions at x for all nodes in the cell
|
||||
MJAPI void mju_evalBasisArray(mjtNum* basis, const mjtNum x[3], int order);
|
||||
|
||||
// map global parametric coord to cell-local coord and build node indices
|
||||
MJAPI int mju_cellLookup(const mjtNum coord[3], const int cellnum[3], int order, mjtNum local[3],
|
||||
int* nodeindices);
|
||||
|
||||
Reference in New Issue
Block a user