In productionProadvancedIncident command
Recovery causes a second outage
Contain an overloaded discovery service without restarting the whole fleet simultaneously.
Pairs with Fleet incident response in Production ROS Operations. Read the concept first, then diagnose it here.
Unlock this lab with ProOperator report
Operators restart every robot, creating a discovery storm that overloads the network again.
Ubuntu 24.04, ROS 2 Jazzy, Cyclone DDS, containers, systemd, OpenTelemetry fleet sandbox
System boundary
Trace only the relevant path.
operator recovery actionfleet restart wavenetwork capacity
How this lab works
You diagnose it — no command list.
Open the repair workspace and gather your own evidence in a real terminal. No diagnostic commands are handed to you — finding the fault is the exercise. Stuck? Progressive hints unlock inside the workspace.
Verification
Your repair must pass more than the visible symptom.
- Staged recovery remains below capacity
- Abort threshold stops regression
- Hidden regression behavior
- Root cause explanation