Skip to content

Repository files navigation

GRiD

CI docs license python agent-ready All Contributors

A GPU-accelerated library for robot dynamics, kinematics, and collisions, with analytical derivatives and Hessians for supported numerical operations.

The GRiD package ecosystem: a user's URDF goes through URDFParser to the code generator (built on GLASS) and RBDReference, producing CUDA C++ with NumPy, JAX, and PyTorch wrappers; benchmarks and tests, backed by pytest-gpu-proof and external oracles, produce validated outputs and performance benchmarks.

GRiD turns a URDF into optimized, per-robot CUDA C++ for rigid-body dynamics, kinematics, their analytical first- and second-order derivatives and a trajectory-optimization plant layer, then hands you that code three ways — a numpy handle, a jax.jit-able FFI surface, or torch.autograd-aware ops — from one content-addressed .so cache. One CUDA block per problem, batched, bit-deterministic and thread-count invariant; the same model and API rebuild for each target architecture (one artifact per sm_XX), with runtime memory adaptation from embedded Jetson class devices to desktop GPUs. Website: https://a2r-lab.github.io/GRiD/.

GRiD builds on our URDFParser, RBDReference, and GLASS packages (URDF parsing, Pinocchio-validated reference dynamics, and GPU linear algebra), together with its own bundled code generator. Using its scripts, users can easily generate and test optimized rigid body dynamics CUDA C++ code for their URDF files.

Ongoing development and the upcoming rerelease live in A2R-Lab/GRiD. The original ICRA 2022 paper describes the implementation preserved in the archival robot-acceleration/GRiD repository, not the full feature set or performance of the upcoming release. See the project website for the overview. Collision routines use the generated CUDA interface; numerical Python interface coverage is documented separately.

I want to…

Task Start here
Call GRiD from Python (numpy/JAX/torch) grid_rbd.load_robot("robot.urdf", backend=...) — Python wrappers docs · agent guide
Generate CUDA for a new robot grid-generate config/robot_assets/iiwa14.urdf — see Quick Start below
Fit a humanoid build in RAM fast robot setup (algorithm_list=, enable_mujoco_kernels=False)
Add an algorithm adding an algorithm
Run tests / fix a red receipt CI job CUDA validation + test/run_gpu_proof.sh --help
Benchmark benchmarks
Debug a CUDA-vs-numpy mismatch docs/agent_debugging_guide.md — the bug-class bible
Get MuJoCo/mjx-convention I/O handle.mujoco.<method>(...) — values AND derivatives/second-order
Everything else How do I…? on the docs site

Start-here track: examples/README.md routes the four usage tracks — the examples/notebooks/ Python-bindings tour (01-quickstart → 07-inline-cuda), the runnable bindings/examples/ scripts, codegen scripts, and hand-written-CUDA walkthroughs.

This package contains submodules make sure to run git submodule update --init --recursive after cloning!

Quick Start

Install (creates a local venv and registers the grid-generate CLI):

bash install/base_install.sh
source .venv/bin/activate

Generate CUDA code for your robot (ten ready-to-use URDFs ship in config/robot_assets/ — iiwa14, go2, fr3, g1, h1_2, …):

# Via the installed CLI (works on a clean base install):
grid-generate config/robot_assets/iiwa14.urdf                # arm, fixed base
grid-generate config/robot_assets/go2.urdf -f                # quadruped, floating base
grid-generate path/to/robot.urdf [-t EE_JOINT_NAME] [-n NAMESPACE] [-f] [--algorithm-list LIST] [-o OUT.cuh]

# Or via a hardcoded zero-config example (these two pull their URDFs from the
# robot_descriptions package — a DEV dependency; install install/requirements-dev.txt first):
python examples/codegen/generate_iiwa14.py       # iiwa14 fixed base
python examples/codegen/generate_go2_floating.py # Go2 floating base

Validate and debug:

# Print CPU reference values for all algorithms:
python examples/codegen/print_reference_values.py path/to/robot.urdf

# Compile and run the CUDA print kernel (requires nvcc):
python examples/codegen/print_grid.py path/to/robot.urdf

Write your own CUDA kernel against the generated header:

# Step-by-step walkthrough + compiling/validated example kernels:
#   examples/cuda/README.md   (and examples/cuda/wrapper_types.md)
bash examples/cuda/build_and_validate.sh   # generate → nvcc → run → validate

Requires a C++17-capable host compiler (e.g., g++ ≥ 7, clang++ ≥ 5). The benchmark and codegen runtime compile with -std=c++17 — needed for inline variables in the bench common header. With the [torch] extra the per-robot .so follows torch's ATen requirement (-std=c++20 from torch 2.14; needs CUDA 12+ and g++ ≥ 10).

Usage

  • grid-generate PATH_TO_URDF — generate grid.cuh; add -d for full debug mode, -f for floating base, -t JOINT_NAME to target a specific end-effector joint
  • python examples/codegen/print_reference_values.py PATH_TO_URDF — print CPU reference values for all algorithms to validate CUDA output
  • python examples/codegen/print_grid.py PATH_TO_URDF — compile and run the CUDA print kernel against the generated header

Floating-Base Conventions

Floating-base parsing and the Python reference path now accept a public floating-base convention flag. The default is Pinocchio-compatible:

  • floating_base_convention="pinocchio": q = [x, y, z, qx, qy, qz, qw], v = [vx, vy, vz, wx, wy, wz]
  • floating_base_convention="legacy": q = [x, y, z, qw, qx, qy, qz], v = [wx, wy, wz, vx, vy, vz]

GRiD normalizes both public conventions into one shared internal floating-base representation, so the code generator and RBDReference stay consistent under the hood while callers can choose the input/output ordering they need.

Developer Testing

Contributor-facing test workflows (floating-convention regression suite, CUDA equivalence env overrides, shared-memory targets) moved to CONTRIBUTING.md; the receipt/verification policy lives in the CUDA validation guide.

Current Support

GRiD currently fully supports any robot model consisting of revolute, prismatic, and fixed joints that does not have closed kinematic loops. Arbitrary/skew joint axes (a non-cardinal <axis>) are also supported via a dense 6-vector motion subspace — currently for inverse_dynamics and crba only (cardinal-axis robots stay byte-identical; other algorithms and the helical/planar/spherical joint types are later stages).

GRiD implements the full modern rigid-body-dynamics stack: RNEA / CRBA / ABA / Minv / forward dynamics; analytical first-order gradients (ID + FD, incl. external-force gradients); the second-order derivatives (IDSVA-SO both frames with a codegen-time dispatcher, FDSVA-SO); the kinematics family (EE pose / Jacobian / Hessian, general-frame frame_jacobian/J̇/OSC inertia, runtime multi-EE targets); integrators + integrator gradients; the centroidal family (CoM, CCRBA, dccrba, CMM time-variation, Coriolis matrix, energy/ID regressors); the inertial-parameter (π) family (the joint-torque regressor Y with its analytic gradient ∂Y/∂(q,v) and the FD parameter gradient ∂q̈/∂π); contact-frame wrench mapping (contact_fext / register_robot(contact_frames=...)); runtime tool/payload welding (attach_tool/tool_fext); runtime multi-target positions; the collision family (two-tier config_free); a trajectory-optimization grid_plant cost/step layer; and runtime-mutable inertia/transform/joint-dynamics tables. The complete per-algorithm catalog with citations and per-feature detail lives in the CUDA support status page.

RBDReference additionally provides numpy reference oracles — validated against Pinocchio — for generalized gravity, nonlinear effects, kinetic/potential/mechanical energy, the Coriolis matrix, the centroidal quantities (CoM, CoM Jacobian, CCRBA, centroidal momentum) and their derivatives (the analytic dccrba ∂A/∂q tensor — replacing the prior finite-difference oracle — and cmm_time_variation Ȧ), the inverse-dynamics and kinetic/potential-energy regressors, the general-frame Jacobian / J̇ / OSC inertia described above, and the plant/cost/barrier layer above.

Dual-surface equivalence. Every algorithm exists on two surfaces that are tested for numerical agreement: the RBDReference numpy implementation (the oracle, checked against Pinocchio) and the generated CUDA C++ kernels (checked against that same numpy reference). This keeps the GPU codegen honest against an independent, Pinocchio-validated baseline.

Mimic-joint support: per-robot gating is now essentially eliminated. Non-gradient algorithms (RNEA, forward dynamics, ABA, CRBA, …) work for robots with mimic joints, and every gradient emits a correct mimic-reduced result on both the fixed and floating base: inverse_dynamics_gradient/forward_dynamics_gradient, end_effector_pose_gradient/end_effector_pose_hessian, the second-order idsva_so/fdsva_so, the external-force gradients (f_ext_gradient), and the integrator gradients. The centroidal family — com, ccrba, energy, and the centroidal derivatives dccrba/cmm_time_variation — now also runs on mimic robots (the per-body Jacobian and per-unit motion columns carry the mimic multiplier α, validated against the mimic-aware reference). dccrba/cmm_time_variation additionally run on big floating-base robots (e.g. g1/h1_2-floating) via the sweep-pool spill path. No algorithm raises NotImplementedError for mimic robots anymore.

Additional algorithms and features are in development. If you have a particular algorithm or feature in mind please let us know by posting a GitHub issue. We'd also love your collaboration in implementing the Python reference implementation of any algorithm you'd like implemented!

Repo map

Directory Owns Entry doc
grid_codegen/ the code-generation engine: emits grid.cuh AND the checked-in generated binding regions, all driven by the abi_specs.py table codegen architecture
bindings/ the grid-rbd Python package (numpy/jax/torch handles over a cached per-robot .so) bindings/README.md · agent guide
external/ the peer-product submodules: GLASS (GPU linear algebra), RBDReference (Pinocchio-validated numpy oracle), URDFParser each submodule's README
examples/ the start-here track: notebooks/ (Python tour), codegen/, cuda/ examples/README.md
test/ pytest suites + the split-suite/receipt machinery (run_split_suite.py, run_gpu_proof.sh, compile_sched.py) CUDA validation
config/ ten sample URDFs (robot_assets/) + tuned per-GPU launch configs (launch_configs/) + autotune_robot.sh config/robot_assets/URDF_SOURCES.md
docs/ the Sphinx site (source/) + agent_debugging_guide.md (the bug-class bible) docs site
install/ install scripts (base_install.sh, developer_install.sh) + requirements files installation guide

C++ API

For each algorithm GRiD emits four layers: *_inner (core math on shared-mem inputs), *_device (allocates scratch + calls _inner), *_kernel (global entry point with batched timestep loop), and the host wrapper (CPU launcher with H↔D copies). See the codegen architecture docs for the rationale and concrete signatures.

Python API (grid-rbd)

For Python users the grid-rbd package (in bindings/) wraps the per-robot codegen behind a register-then-run UX with numpy, jax, and torch backends. It ships as part of the single repo distribution — a pip install -e . (what install/base_install.sh runs) installs the codegen toolkit and the grid_rbd wrapper together. The base install is minimal; pick a backend extra for the surface you want:

pip install -e "."          # base: numpy backend only
pip install -e ".[jax]"     # + JAX FFI surface
pip install -e ".[torch]"   # + torch backend (CUDA wheel matching your GPU arch)
pip install -e ".[all]"     # jax + torch

See the install matrix in bindings/README.md for what each extra unlocks (and the torch CUDA-wheel note).

import grid_rbd

# numpy (default), jax, or torch; urdf_string= also accepted instead of urdf_path
handle = grid_rbd.register_robot("iiwa14", urdf_path="iiwa.urdf", backend="torch")

qdd = handle.forward_dynamics(q, qd, u)   # autograd-aware torch.Tensor
qdd.sum().backward()                      # gradients flow to q, qd, u

The torch backend exposes autograd-aware inverse_dynamics / forward_dynamics / aba / integrator (analytic backward passes) plus CUDA-Graphs capture, and the handle also surfaces the grid_plant cost/barrier methods. inverse_dynamics (alias rnea) / forward_dynamics (alias fd) take an optional qdd= (the autograd gradient is qdd-aware, returning the correct ∂τ/∂(q,q̇) including the ∂(M·q̈)/∂q term), and all three backends expose the value ops coriolis_matrix, kinetic_energy_regressor, potential_energy_regressor, dccrba, and cmm_time_variation (forward-only on jax/torch). The π-regressor family (inverse_dynamics_regressor, the differentiable inverse_dynamics_wrt_params/forward_dynamics_wrt_params, and the forward_dynamics_parameter_gradient ∂q̈/∂π), runtime tool welding (attach_tool/tool_fext, via enable_tool=True), and multi-contact wrench mapping (contact_fext, via register_robot(contact_frames=[...])) are bound as well. For true fp64 compute build with register_robot(..., dtype="float64") (its own cache entry); allow_fp64=True is only the numpy handle's fp64-in/fp64-out convenience cast on an fp32 build (ignored when dtype="float64"). See bindings/README.md and the Python wrappers docs.

Citing GRiD

To cite GRiD in your research, please use the following bibtex for our paper "GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients":

@inproceedings{plancher2022grid,
  title={GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients}, 
  author={Brian Plancher and Sabrina M. Neuman and Radhika Ghosal and Scott Kuindersma and Vijay Janapa Reddi},
  booktitle={IEEE International Conference on Robotics and Automation (ICRA)}, 
  year={2022}, 
  month={May}
}

Performance

Release measurements from the 27 September 2026 run on one NVIDIA RTX 5090 with an Intel Core Ultra 9 285K cover RNEA, its analytical gradient (∇RNEA), and its analytical Hessian (∇²RNEA) on iiwa14 (fixed base, 7 velocities), go2 (floating base, 18), and G1 (floating base, 35) at batch sizes 16–1024. The release measurements page gives the method, every timing boundary, and the caveats; the benchmark harness reproduces the collection.

Core-operation speedups against seven baseline modes, with timing boundaries and fp64 exceptions labeled.

Ratios are baseline time divided by GRiD time; above 1× favors GRiD. Each column names its timing boundary: GRiD host calls including copies against the CPU libraries, and GRiD compute-only calls against the GPU libraries' resident calls. * marks cells where the evaluated baseline path required fp64 and ~ a side whose run means span more than 1.5×. Colors are clipped at 100×.

Clustered GRiD, Pinocchio and MuJoCo timing bars on three robots; Hessians compare GRiD with Pinocchio's standard API only.

Microseconds per complete batch on a log axis. GRiD's bar splits into its CUDA compute-only call, the GPU–CPU I/O increment, and the JAX wrapper increment. These are differences of measured call times, not isolated measurements of each component.

Call wall times for CUDA Device, C++ Host, NumPy, PyTorch, and JAX, in that order, for RNEA, its gradient and its Hessian on three robots; Python bars are solid to the allocate-once call and hatched up to the default call.

Call wall times through each API boundary: native CUDA, the C++ host call, NumPy, PyTorch, and JAX. Pick the boundary your application uses. For the Python surfaces the solid bar is the call with its buffers allocated once and reused (measured 2 October 2026), and the hatched cap reaches the default call, which allocates its output every time. With reused buffers, NumPy and PyTorch land within a few percent of the C++ host call on large outputs.

Installation

The Quick Start above covers the common-case install. For CUDA Toolkit setup, developer dependencies (Pinocchio, robot_descriptions, benchmarks), and Docker, see the full installation guide.

Troubleshooting

Bench harness nvcc hangs on floating-base kernels (sm_8x)

On Ampere (sm_86 / CUDA 12.6) the bench harness can wedge nvcc / ptxas at 100 % CPU when compiling heavy floating-base GRiD harnesses. Pass --ptxas-opt-level 2 to test/benchmarks/run_multi_version.py — it forwards -Xptxas -O2 to floating-base compiles only. Blackwell (sm_120) does not hit this. Typical user code that includes grid.cuh and calls the batch host wrappers (e.g. grid::forward_dynamics<T>(...)) does not trigger the hang — it's specific to the timing-bench template surface.

Contributing

Contributions welcome — see CONTRIBUTING.md for the workflow (and CLAUDE.md for the repo conventions AI agents and humans both follow).

Contributors

Brian Plancher
Brian Plancher

Zachary Pestrikov
Zachary Pestrikov

Kwamena A
Kwamena A

Danelle Tuchman
Danelle Tuchman

Ann Li
Ann Li

=
=

EmreAdabag
EmreAdabag

Cael Yasutake
Cael Yasutake

Naren Loganathan
Naren Loganathan

emilyburnett2003
emilyburnett2003

pruyontrarakk
pruyontrarakk

Kimiya Shahamat
Kimiya Shahamat

Kimiya Shahamat
Kimiya Shahamat

About

A GPU accelerated library for computing rigid body dynamics with analytical gradients

Resources

Code of conduct

Contributing

Security policy

Stars

17 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages