Skip to content

Incident detail

Resolved: KeeperPAM Connection Errors in US Data Center

Resolved incidentMajor

Timeline window

to

Outage alerts

Get alerted the next time Keeper Security breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are investigating errors in establishing KeeperPAM connections in the US Data Center.

  2. Identified

    We have identified the cause of the KeeperPAM connection errors which are related to a networking issue within the AWS ECS environment. We are actively troubleshooting with the AWS support team and will update the ticket as soon as there is a status update.

  3. Monitoring

    A fix has been implemented and we are monitoring the KeeperPAM connection stability. We will update this case with additional details soon.

  4. Resolved

    The issue has been resolved. See postmortem page with additional details.

  5. Resolved

    At 9:05 AM PST, alerts were triggered for issues affecting KeeperPAM connections managed through Keeper’s ECS deployments in the US-EAST region. There were no recent changes to the application or environment.

    Investigation identified a low-level concurrency bug in the Keeper EPM service that caused request failures under high simultaneous load. These failures led to instability in the ECS services supporting KeeperPAM connections.

    As a temporary mitigation, we blocked the error condition, restoring KeeperPAM connectivity by approximately 1:00 PM PST.

    The engineering team then developed and deployed an updated Keeper Router version to address the underlying issue and prevent EPM agents from triggering server errors. The fix was fully validated by 3:00 PM PST, at which point all KeeperPAM services were stable and operating normally.