Refactor and speed up mj_crb and add a benchmark.
This is a no-op "prefactor" of `mj_crb` to reduce the number of lines changed in the upcoming sleeping CL. The main change here is that `mj_crb` now avoids accessing the model and data pointer repeatedly, but instead has the input and output pointers explicitly declared as local variables. A benchmark test found a healthy **14% perf bump** due to two changes: - Adding `restrict` to `mju_mulInertVec` - The function-local pointers. Adding `restrict` to the local pointers had no effect. Note that `mj_crb` is not a particularly expensive function so these speed bumps are not significant per se, but rather indicative of possible future gains with these techniques. ``` Benchmark Time(ns) CPU(ns) Iterations --------------------------------------------------------------- ORGINAL BASELINE BM_CRB_BASELINE_mean 4325 4348 993200 230.052k items/s BASELINE + RESTRICT BM_CRB_BASELINE_mean 4057 4090 1186150 244.573k items/s LOCAL POINTERS + RESTRICT BM_CRB_mean 3772 3800 1001750 263.209k items/s ``` PiperOrigin-RevId: 815798225 Change-Id: Iffdf57b544e8e10562807c617f73ca4ccd1c614f
This commit is contained in:
committed by
Copybara-Service
parent
b75af81c75
commit
ab9575d8a6
@@ -431,7 +431,7 @@ void mju_inertCom(mjtNum res[10], const mjtNum inert[3], const mjtNum mat[9],
|
||||
|
||||
|
||||
// multiply 6D vector (rotation, translation) by 6D inertia matrix
|
||||
void mju_mulInertVec(mjtNum res[6], const mjtNum i[10], const mjtNum v[6]) {
|
||||
void mju_mulInertVec(mjtNum* restrict res, const mjtNum i[10], const mjtNum v[6]) {
|
||||
res[0] = i[0]*v[0] + i[3]*v[1] + i[4]*v[2] - i[8]*v[4] + i[7]*v[5];
|
||||
res[1] = i[3]*v[0] + i[1]*v[1] + i[5]*v[2] + i[8]*v[3] - i[6]*v[5];
|
||||
res[2] = i[4]*v[0] + i[5]*v[1] + i[2]*v[2] - i[7]*v[3] + i[6]*v[4];
|
||||
|
||||
Reference in New Issue
Block a user