The dashboard went blank
Most edge incidents start the same way. Grafana goes empty. The GitOps agent stops reporting. Someone concludes the NVIDIA Jetson is down, the model crashed, or K3s ate itself. A site visit later, the camera pipeline is still producing boxes. The node never crashed. The path to it did.
Cloud GPU clusters train teams to treat silence as death. If Prometheus cannot scrape, the replica is gone. If the API server cannot schedule, the workload is gone. That heuristic is cheap in a datacenter. It is expensive in a factory, a store, or a vehicle, where the first thing to fail is almost never TensorRT or GPU memory. It is reachability.
Edge failure modes are network failure modes first. Until the platform can distinguish a partition from a crash, every blank panel becomes an outage, and every outage becomes a truck roll.
