CXone Knowledge Management - – Monitoring complete. Status = All Services Running Normally

Incident Report for CXone Expert US

Postmortem

Impact Start Time (UTC) 07/09/2026 07:18 AM UTC

Impact End Time (UTC) 07/09/2026 08:18 AM UTC

Incident Summary

On 07/09/2026, some NiCE CXone Mpower customers experienced slow page load times, intermittent access issues, or temporary unavailability of the CXone Mpower Expert knowledge portal. The issue stemmed from a scheduled production release that inadvertently introduced a configuration change, reducing the resource allocation for the autoscaling service on the affected production cluster. The impact was resolved after the change was rolled back, redirecting traffic to the previous stable version and restoring service.

Root Cause

The issue stemmed from a scheduled production release that inadvertently introduced a configuration change, reducing the resource allocation for the autoscaling service on the affected production cluster. During the investigation, engineers determined that a previous resource increase unexpectedly failed to successfully merge into the codebase. As a result, the deployment reapplied the code-defined configuration, inadvertently overwriting the existing production setting and reverting the autoscaler to its previous resource allocation. This caused the autoscaler to encounter out-of-memory (OOM) conditions and was unable to scale cluster capacity as required, leading to degraded platform performance and temporary site unavailability for some customers. The release successfully passed standard staging validation and was deployed to another production environment without issue. The condition only manifested in the affected production cluster under higher traffic volume. While incidents of this nature are rare during routine production releases, the event highlighted opportunities to further strengthen deployment controls and configuration validation to reduce the risk of similar occurrences in the future

Corrective Actions

Detection: Internal monitoring and automated alerts identified the service failure, which was subsequently confirmed through customer reports of service degradation and temporary inaccessibility to some CXone Mpower Expert sites. Corrective Actions

Remediation: The impact was resolved after the change was rolled back, redirecting traffic to the previous stable version and restoring service. This reduced cluster resource demand, allowing the autoscaling service to recover and resume normal operations. Completed on 07/09/2026.

Prevention: The Engineering team restored the autoscaler's resource allocation in the affected production environment to its intended increased values, ensuring sufficient capacity to support cluster demand and reducing the risk of similar service impacts in the future. Completed on 07/10/2026.

The release pipeline will be enhanced to introduce additional validation, and readiness checks earlier in the deployment process ("shift-left" testing). This change will ensure that critical workloads, infrastructure capacity, and service health are fully validated before the traffic is routed to the new release version. An update will be provided by End of Day MT on 07/24/2026.

NiCE resiliency teams remain focused on strengthening system reliability, stability, platform-wide resilience, and post-deployment validation, monitoring, and performance. These efforts are aligned under the Operational Resiliency Excellence initiative to drive sustained progress and measurable outcomes.

Incident Timeline (UTC)

07/09/2026 07:18 AM (UTC) - Internal monitoring and deployment validation detected failures in the production environment. Engineering teams immediately began investigating the issue and initiated recovery actions.

07/09/2026 07:23 AM (UTC) - The first customer case was opened, and Tech Support (TS) engineers began their initial validation and troubleshooting investigation.

07/09/2026 07:56 AM (UTC) - While the remediation effort was ongoing, TS engineers notified the Network Operations Center (NOC) engineers about the reported customer impact; a major incident was proposed and confirmed.

07/09/2026 08:18 AM (UTC) - The affected release was successfully rolled back, restoring normal service. Following successful validation, the major incident was marked as resolved.

Posted Jul 16, 2026 - 17:20 UTC

Resolved

CXone Knowledge Management - Service Disruption Resolved - All Services Running Normally. The CXone Mpower Expert Engineering team has deployed a fix and monitored the deployment to make sure sites are stable. The issue is now resolved at this time. Event duration 38 minutes
Posted Jul 09, 2026 - 08:37 UTC

Monitoring

CXone Knowledge Management - Fix Deployed - All Services Running Normally. The CXone Knowledge Management Engineering team has deployed a fix and all services are running normally. We are currently monitoring sites for deployment stability. Event duration 31 minutes
Posted Jul 09, 2026 - 08:29 UTC

Identified

CXone Knowledge Management Service Degradation: Sites unavailable. The issue has been identified and a fix is being worked on for deployment.
Posted Jul 09, 2026 - 08:17 UTC

Investigating

CXone Mpower Expert Service Degradation: Sites unavailable. The CXone Mpower Expert Engineering
team is investigating reports of site unavailability.
Posted Jul 09, 2026 - 07:58 UTC
This incident affected: Application (General Service), Search, Generative Search, In-Product Contextual Help, Email Services, MindTouch Success Center, Analytics, and Geoblocking for Russia.