In productionProadvancedIncident command

Recovery causes a second outage

Contain an overloaded discovery service without restarting the whole fleet simultaneously.

Pairs with Fleet incident response in Production ROS Operations. Read the concept first, then diagnose it here.

Unlock this lab with Pro

Operator report

Operators restart every robot, creating a discovery storm that overloads the network again.

Ubuntu 24.04, ROS 2 Jazzy, Cyclone DDS, containers, systemd, OpenTelemetry fleet sandbox

System boundary

Trace only the relevant path.

operator recovery actionfleet restart wavenetwork capacity

How this lab works

You diagnose it — no command list.

Open the repair workspace and gather your own evidence in a real terminal. No diagnostic commands are handed to you — finding the fault is the exercise. Stuck? Progressive hints unlock inside the workspace.

Verification

Your repair must pass more than the visible symptom.

  • Staged recovery remains below capacity
  • Abort threshold stops regression
  • Hidden regression behavior
  • Root cause explanation