In productionProadvancedService recovery

A restart loop erases the useful log

Bound restarts and preserve the first failure evidence.

Pairs with systemd and startup order in Production ROS Operations. Read the concept first, then diagnose it here.

Unlock this lab with Pro

Operator report

A misconfigured node restarts 400 times, floods logs, and exhausts CPU.

Ubuntu 24.04, ROS 2 Jazzy, Cyclone DDS, containers, systemd, OpenTelemetry fleet sandbox

System boundary

Trace only the relevant path.

process failuresystemd restart policyhost resources

How this lab works

You diagnose it — no command list.

Open the repair workspace and gather your own evidence in a real terminal. No diagnostic commands are handed to you — finding the fault is the exercise. Stuck? Progressive hints unlock inside the workspace.

Verification

Your repair must pass more than the visible symptom.

  • Retries are bounded
  • First failure remains queryable
  • Hidden regression behavior
  • Root cause explanation