Comment by theoli

9 hours ago

You need some sort of external monitor for a missed pulse style alert. Personally I do something similar by collecting a "last success" metric with Prometheus and alerting on it with Grafana if the value is too far in the past. The local system cannot reliably alert if your job does not fail into the alert path or the system is just down.

I can see something like this being a great intermediate option to a full monitoring and alerting stack.