Skip to content

Incident detail

Custom Dashboards are failing intermittently in prod3

Resolved incidentMinor1 affected service

Timeline window

to

Outage alerts

Get alerted the next time Harness breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are currently investigating this issue.

  2. Monitoring

    A fix has been implemented and we are monitoring the results.

  3. Monitoring

    We are continuing to monitor for any further issues.

  4. Resolved

    This incident has been resolved.

  5. Resolved

    # Summary

    Dashboards service in Prod3 experienced intermittent failures, causing dashboards to return errors and become unavailable to some customers.

    ## Root Cause Analysis

    The incident was caused by slow Looker `search_dashboards` API calls degrading from 7 seconds to over 30 seconds, which saturated the worker thread pool and prevented the health endpoint from responding to Kubernetes liveness and readiness probes.

    ## Mitigation Steps Taken

    ### Immediate Mitigations

    1. Increased liveness probe timeout from 15s to 30s (failure threshold kept at 3, allowing up to 90s tolerance)

    2. Doubled thread count , increasing total concurrency

    3. Increased pod replica count to maintain higher minimum availability and reduce risk of all pods becoming simultaneously unavailable

    ## Actions

    Harness will work on the following Action items to prevent recurrence.

    1. Optimize Looker SDK timeout: Tune timeout to Looker SDK `search_dashboards` calls to prevent indefinite thread occupation

    2. Optimize Looker queries : Investigate Looker-side query performance degradation to understand why `search_dashboards` latency increased