Sporadic Email Failures on EU Production Instance

Incident Report for Viedoc

Postmortem

Sporadic Email Failures on EU Production Instance

Description

During the two windows below, email alerts triggered by actions in Viedoc were sent intermittently. Most were delivered as expected. A subset were delayed, and some were not delivered at all.

  • Monday, August 31, 11:02–17:25 CEST
  • Thursday, September 3, 13:31 CEST – Friday, September 4, 22:07 CEST

Email alerts in Viedoc are not sent by the application at the moment a user saves data. They are produced by background processing that runs immediately afterwards and decides, based on your study's configuration, which alerts need to be sent. During these two windows that background processing did not always complete, so the alerts it would have produced were delayed or never generated.

Data entered and saved by users during these periods was stored as expected. The interruption affected work handled by particular processing servers rather than the platform as a whole, which is why alerts were intermittent rather than stopping entirely.

Cause

Viedoc's background processing runs across several servers working in parallel. Each server needs a working connection to a shared internal service that Viedoc uses to coordinate this work safely - for example, to ensure two servers never process the same subject at the same moment.‌ During these windows, individual servers lost that connection and did not re-establish it. An affected server continued to accept work but could not complete it. Other servers running at the same time were unaffected, which is why the impact was intermittent and why the platform continued to appear healthy overall.

Corrective action

  • The affected servers were identified and replaced. Normal processing, including alert generation, resumed immediately afterwards.
  • We confirmed that subsequent processing completed correctly, and that no further email failures occurred.

Preventative action

Increased monitoring

We have increased the monitoring and added alerting that detects this specific failure on any individual server much faster, preventing this to fail silently.

Underlying cause - ongoing investigation

‌The investigation into why these connections failed remains open, and we are treating it as separate from the mitigations above.

Posted Sep 17, 2026 - 08:17 UTC

Resolved

This incident has been resolved.
Posted Sep 11, 2026 - 08:25 UTC

Monitoring

We have not observed any new email failures since our previous report. We are now monitoring the situation and will post a further update once we can confirm the issue is fully resolved.
Posted Sep 10, 2026 - 13:51 UTC

Investigating

Email alerts triggered by EDC actions in Viedoc have been failing intermittently on the EU production instance during the following time windows:

Monday, August 31, 11:02–17:25 CEST
Thursday, September 3, 13:31 CEST – Friday, September 4, 22:07 CEST

Most emails are being sent as expected, but a subset are delayed or, in some cases, not delivered at all. No other instances are affected. We are actively investigating and will post updates as more information becomes available.
Posted Sep 08, 2026 - 14:07 UTC
This incident affected: Viedoc 4 - Europe (Background processes (import/export/revisions/alerts/archive/disposal)).