Skip to content

Incident detail

Data ingestion is delayed on Traceable US production

Resolved incidentMajor

Timeline window

to

Get alerted the next time Harness breaks

Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update on the official status source, oldest to newest, exactly as it appeared there.

  1. Investigating

    We are currently investigating this issue.

  2. Investigating

    We are continuing to investigate this issue.

  3. Investigating

    We are continuing to investigate this issue.

  4. Identified

    The issue has been identified and a fix is being implemented.

  5. Monitoring

    A fix has been implemented and we are monitoring the results.

  6. Monitoring

    We are continuing to monitor for any further issues.

  7. Resolved

    This incident has been resolved.

  8. Resolved

    Summary

    On 19 August 2026 between 12:35 and 17:29 UTC, the Harness Application Security service experienced a significant disruption affecting both the customer-facing console and the data ingestion pipeline in the SaaS Production and US1 regions.

    Root Cause

    The internal configuration service that supplies runtime settings to nearly every other component became overloaded and entered a repeated restart cycle. Because so many services depend on it, the effects were broad: console pages such as protection policies, posture views, activity logs, API inventory, and custom policy failed to load or timed out, and downstream processing stalled while waiting for configuration it could not obtain.

    # Customer impact

    | Dimension | Detail |

    | --- | --- |

    | Console (UI) impact | Multiple pages failed to load or timed out, including protection policies, posture event pages and posture views inside dashboards and insight pages, activity log queries, API inventory screens, custom policy, and sensitive-data views and widgets. |

    | Ingestion impact | Security telemetry processing degraded severely and, in some paths, stopped entirely. Consumer lag grew across normalisation, grouping, anomaly detection, generation, and related processing stages. |

    | Data loss | A subset of telemetry ingested during the disruption was permanently dropped. |

    Mitigation

    Several intermediate mitigations additional CPU and memory, relaxed health-check thresholds, a database restart, and a larger connection pool ameliorated the issue. Disabling the new feature in both affected regions restored throughput sharply and durably. The incident was resolved at 17:29 UTC.

    # Preventive actions

    The following actions are committed and tracked internally to completion. The feature that triggered this incident remains disabled and will not be re-enabled until the work below is complete and validated.

    | Action |

    | --- |

    | |

    | OPtimize the code by tuning parameters such as cache eviction and retention , evaluate cursor-based pagination for bulk rule retrieval as rule counts grow |

    | Add a purpose-built database index for the service-scoping access pattern |

    | Remediate pipeline recovery semantics so consumers replay safely after position-marker loss instead of skipping backlog |

    | Mandate staged rollout for configuration overrides that alter downstream request patterns: low-volume cluster, then mid-volume, then high-volume |

    | Add backpressure and concurrency protection to the configuration service: circuit breaking, bounded queues, and timeout isolation |

    | Enhance observability by Instrumenting more detailed metrics |

Keep exploring

More from Harness

Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.