We have one of our primary mx servers not responding. It seems to be hung at a hardware level. We have removed it from our DNS MX record list and are working with the vendor to resolve the issue. This could have caused some temporary delays while attempting deliveries to this server.
May 26, 23:38 UTC
The server issue has been resolved by our hosting provider. Everything should be back to normal.
May 27, 10:28 UTC
The MX server with hardware failure was still rejecting a certain portion of messages. This caused delays mostly coming from Microsoft servers. We are currently investigating and have removed the server from our MX pool again.
May 27, 15:07 UTC
After a thorough investigation the root cause of the rejected messages was determined to be a simple disk space issue for the message queue. During the hardware failure the disk was filled with errors, but enough space remained to process some messages while rejecting others depending on the volume of inbound messages.
To compound the issue the internal alerting had been erroneously paused for an excessive duration while the initial hardware replacement was done. Once completed the internal alerting should have been reactivated, but wasn't. This would have alerted us immediately to the problem. We are reviewing the procedures concerning the internal alerting.