There Is No Haystack
I was reminded by a recent incident that it’s often just as important to be intentionally made aware that things are working, as it is to know when they’re not. And it's a lot easier to spot the needle if you reduce the haystack to a few blades of grass.
Every morning for the past few years, I've been reviewing and deleting 18 separate emails reporting the status of 18 separate backup jobs.
Below is my response after I was asked to investigate a failed operation someone noticed this morning.
Good morning!
Actually, at 12:01am this morning we received a notification via email that the server named "Justin" was unable to create a backup of server "Justin" (itself).
The 9:01am CDT email you saw is a status-summary from one of our two backup servers; the summary from the other server, "Case," showed all its' jobs completed successfully.
The alert from Justin indicated a file-access conflict - one process trying to read a file while another process is trying to write to it.
Every few days an access-conflict like this prevents the backup servers from backing themselves up.
The backup servers don’t really change much, and in that context this error/failure is benign unless it happens multiple days in a row.
We could try rescheduling the backups to mitigate the conflict, but I don’t think it’s worth rocking that boat for this error related to this particular server.
There are nine servers being backed up, including the two backup servers.
Until yesterday, successful or not, each server triggered 2 status-alert emails per day: local backup status, and off-site copy status.
Yesterday I reconfigured daily notifications so we're sent summaries from each server, and individual alerts only from failed operations.
This morning someone noticed an error that’s been happening, on and off, for a good, long while; and about which they've been receiving status alerts amidst those 18 daily emails.
And that’s why I switched to digest emails.