{
  "meta": {
    "version": "v1",
    "pricing": {
      "public": {
        "label": "Public",
        "description": "Read-only API access for lightweight status checks and public integrations."
      },
      "premium": {
        "label": "Premium",
        "description": "API keys with higher hourly quotas, plus Slack, Teams, Discord, webhook, and email outage alerts across your vendor stack."
      }
    },
    "generatedAt": "2026-09-13T16:05:25.874Z"
  },
  "data": {
    "id": "incident_statuspage_harness_dwb2r7637mp5",
    "slug": "harness-prod3-environment-is-down-2026-07-26",
    "title": "The Prod3 & Prod1 environment is experiencing intermittent outages. We are currently investigating the issue.",
    "summary": "The Prod3 & Prod1 environment is experiencing intermittent outages. We are currently investigating the issue.",
    "status": "resolved",
    "stale": false,
    "severity": "major",
    "startedAt": "2026-07-27T06:42:31.598+00:00",
    "updatedAt": "2026-08-07T18:16:09.379+00:00",
    "resolvedAt": "2026-07-27T13:33:56.03+00:00",
    "provider": {
      "slug": "harness",
      "name": "Harness"
    },
    "affectedServices": [
      {
        "slug": "harness-cicd",
        "name": "CI/CD"
      },
      {
        "slug": "harness-feature-flags",
        "name": "Feature flags & FME"
      },
      {
        "slug": "harness-platform",
        "name": "Platform & dashboards"
      }
    ],
    "links": {
      "html": "/incidents/harness-prod3-environment-is-down-2026-07-26",
      "api": "/api/v1/incidents/harness-prod3-environment-is-down-2026-07-26",
      "providerHtml": "/providers/harness",
      "alerts": "https://outagedeck.com/account?stack=harness&utm_source=api&utm_medium=response&utm_campaign=api_alerts&utm_content=incident"
    },
    "impactSummary": "Harness reported a major event for the affected tracked services.",
    "source": {
      "id": "source_harness_status",
      "kind": "official_api",
      "name": "Harness Status",
      "checkedAt": "2026-09-13T16:00:09.159+00:00",
      "stale": false,
      "officialUrl": "https://status.harness.io",
      "statusPageUrl": "https://status.harness.io"
    },
    "updates": [
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_yztg3nlx3v9n",
        "status": "investigating",
        "body": "We are currently investigating this issue.",
        "createdAt": "2026-07-27T06:42:31.787+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_1ljnl6lkbyft",
        "status": "identified",
        "body": "The issue has been identified and a fix is being implemented.",
        "createdAt": "2026-07-27T06:45:36.501+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_s6s680d9ysxs",
        "status": "monitoring",
        "body": "A fix has been implemented and we are monitoring the results.",
        "createdAt": "2026-07-27T06:49:17.069+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_xr2bnxtzgpks",
        "status": "investigating",
        "body": "We are currently investigating this issue.",
        "createdAt": "2026-07-27T07:52:01.328+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_kcc2m0v7l4qp",
        "status": "investigating",
        "body": "We are continuing to investigate this issue.",
        "createdAt": "2026-07-27T08:13:40.048+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_2dynfq8t0fxr",
        "status": "identified",
        "body": "The issue has been identified and a fix is being implemented.",
        "createdAt": "2026-07-27T09:04:11.85+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_nb631655g50n",
        "status": "identified",
        "body": "We are continuing to work on a fix for this issue.",
        "createdAt": "2026-07-27T09:05:35.379+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_d3dhp1nspnrm",
        "status": "monitoring",
        "body": "A fix has been implemented and we are monitoring the results.",
        "createdAt": "2026-07-27T09:07:24.674+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_rj223dtz6ly6",
        "status": "monitoring",
        "body": "We are continuing to monitor for any further issues.",
        "createdAt": "2026-07-27T09:07:50.382+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_tvqc1w5sgw0y",
        "status": "monitoring",
        "body": "We are continuing to monitor for any further issues.",
        "createdAt": "2026-07-27T09:09:54.925+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_hf127nyzgb18",
        "status": "resolved",
        "body": "This incident has been resolved.",
        "createdAt": "2026-07-27T13:33:56.03+00:00"
      },
      {
        "id": "update_statuspage_harness_dwb2r7637mp5_s6pl9w1qnlp1",
        "status": "resolved",
        "body": "# Summary\n\nDuring a recent production deployment, a defect in our internal deployment tooling caused two critical services   to run with incorrect, non-production configuration values in our production environment This led to a related set of four distinct symptoms: incorrect configuration behavior, intermittent login/access failures, a filestore access issue affecting one customer environment, and delayed pipeline status updates in the UI.\n\n‌\n\nWe have identified and are implementing a permanent fix for the underlying configuration defect, and have already put in place resource and capacity changes that resolve the UI delay symptom.\n\n‌\n\nAt no point during this incident were pipeline executions themselves lost, corrupted, or left in a stuck state. Where execution behavior was affected, it was limited to delays in status visibility, not in the underlying processing.\n\n# Incident Details\n\n## Incorrect Production Configuration Values Applied\n\nOur engineering team confirmed a defect in the Service Manager deployment pipeline that caused certain production services to be deployed using configuration values intended for a different environment, rather than the correct production configuration.\n\n‌\n\n**Root Cause**\n\n‌\n\nThe service responsible for fetching configuration overrides during deployment queries an internal API that returns a maximum of 1,000 results per request. The total number of services in the environment recently grew beyond that limit. As a result, any service beyond the first 1,000 returned was not included in the response, and the deployment pipeline silently fell back to default configuration values for those services. This is a confirmed pagination defect in the deployment tooling, not an issue with the configuration values themselves.\n\n‌\n\n**Resolution**\n\n‌\n\nEngineering has confirmed the mechanism and is implementing a permanent fix to remove this limit-related gap in the deployment pipeline.\n\n## Intermittent Login / Access Failures\n\nDuring the Service Manager deployment referenced above, some users experienced intermittent login or access failures. Under normal operation, previously running instances should continue serving traffic without interruption while a new deployment is in progress. In this incident, that fallback behavior did not occur as expected, contributing to access failures during the deployment window.\n\n## Filestore Access Issue\n\nA filestore access issue was identified that was specific to the Prod-3 environment and affected a single customer's environment.\n\n**Root Cause**\n\nThis is related to an IAM / storage-bucket permission configuration on Service Manager, potentially triggered by rollback activity. \n\n## Delayed Pipeline Execution Status Updates in UI\n\nSome users observed that the pipeline execution graph in the UI was slow to refresh and did not reflect the latest status promptly. Importantly, this was a visibility delay only: there was no impact to actual pipeline executions, and no executions were stuck or failed as a result of this issue.\n\n**Root Cause**\n\nThe pipeline execution graph relies on a message stream \\(the orchestration log\\) to receive status updates. During the incident window, consumer processing of this stream fell behind \\(high consumer lag\\), which delayed how quickly status updates reached the UI. This was caused by the fact that the underlying database was in the middle of a planned scaling operation at the same time, and a traffic spike during that window further exacerbated the delay. Users experienced this as apparent pipeline slowness, even though the underlying executions were running normally.\n\n**Resolution**\n\nWe have increased resource capacity for the affected components to maintain more than 50% spare headroom going forward, reducing sensitivity to similar load spikes. This change has been implemented and is currently being validated as part of longer-term hardening for this part of the platform.\n\n# Impact Summary\n\n* Service Manager and License Manager ran with incorrect configuration values in the Prod-1 and Prod-3 environments.\n* Some users experienced intermittent login or access failures during the affected deployment window.\n* One customer environment in Prod-3 experienced a filestore access issue.\n* Users across affected environments saw delayed pipeline execution status updates in the UI; underlying pipeline executions continued to run correctly and were not lost, stuck, or corrupted.\n\n# Preventive Actions\n\nThe following corrective and preventive actions have been identified.\n\n‌\n\n| **Corrective / Preventive Action** |\n| --- |\n| Correct the pagination limit in the configuration-lookup service so that all services are returned and evaluated, regardless of total count. |\n| Add safeguards so that a service which cannot retrieve its configuration fails safely \\(e.g. alerts and blocks the deployment\\) rather than silently falling back to non-production defaults. |\n| Increase Postgres and messaging-pipeline resource headroom \\(target: greater than 50% spare capacity\\) to reduce sensitivity to concurrent load and scaling events. |\n\n‌\n\n_We recognize the impact this incident had across multiple areas of the platform and appreciate your patience as we work through a complete resolution._",
        "createdAt": "2026-08-07T18:15:34.165+00:00"
      }
    ],
    "access": {
      "plan": "public",
      "keyed": false
    }
  }
}