Service Disruption Affecting Some Customers

Incident Report for Brillium

Postmortem

Post-Incident Review and Service Improvement Summary

Incident Date: September 2026
Status: Resolved
Audience: Business Users, Customers, Partners, and Stakeholders
Reference: Brillium Incident Notification and Brillium Scheduled Maintenance Notice

Executive Summary

During a period of unusually high system activity, a subset of Brillium customers experienced intermittent access and performance issues while using the platform. While the majority of users remained unaffected, some customers encountered difficulties completing actions within the application.

The issue was isolated to a legacy component responsible for managing user sessions. Under heavier-than-normal traffic conditions, this component experienced contention that prevented some user sessions from being processed efficiently, resulting in delayed responses or temporary interruptions.

No customer data was lost, altered, exposed, or compromised at any time during this event. The incident affected system availability only and did not impact the confidentiality, integrity, or security of customer information.

Brillium's Operations and Engineering teams responded immediately, identified the root cause, implemented a corrective solution, and conducted extensive validation before releasing the fix. Following successful testing, the solution was deployed incrementally to a limited group of customers and monitored closely before broader rollout.

The platform is operating normally.

What Happened?

Brillium utilizes technology that maintains a user's active session while they work within the platform. This session management capability ensures that users remain authenticated and can move seamlessly between application functions.

During periods of elevated platform activity, a legacy component involved in this process became a bottleneck. Specifically, the system's session coordination mechanism was unable to efficiently handle the volume of simultaneous session activity, causing some session requests to wait longer than expected.

As demand increased, this created intermittent service degradation for a subset of users.

It is important to note that this issue was related solely to session management and application availability. It was not caused by a cybersecurity incident, unauthorized access attempt, software defect affecting customer data, or infrastructure failure.

Customer Impact

Some customers may have experienced:

  • Intermittent application responsiveness
  • Delayed page loads
  • Temporary session interruptions
  • Difficulty accessing specific areas of the platform during peak activity periods

Customers not utilizing the affected application paths generally experienced normal service.

What Was Not Impacted

The following remained protected throughout the event:

  • Customer assessment data
  • Candidate data
  • User account information
  • Assessment content
  • Reports and analytics
  • Integrations and stored records
  • Security controls and access management

There is no evidence of data loss, data corruption, unauthorized access, or privacy impact.

Root Cause

The root cause was identified as a limitation within a legacy session-management architecture.

The affected component relied upon a centralized mechanism for tracking active user sessions. Under high transaction volumes, this mechanism could become constrained, reducing the system's ability to process concurrent session updates efficiently.

As platform demand increased, contention within that process created temporary delays that affected some customer experiences.

The issue was related to system scalability rather than security, data integrity, or infrastructure capacity.

Immediate Response

Upon detection of the issue, Brillium initiated its incident response process, which included:

  1. Immediate monitoring and investigation.
  2. Identification and isolation of the affected component.
  3. Development of a corrective workaround.
  4. Extensive internal testing and quality validation.
  5. Controlled deployment to a limited set of customer environments.
  6. Continuous monitoring to verify successful resolution.
  7. Execution of scheduled maintenance activities to support the corrective actions and long-term stability improvements.

The workaround successfully mitigated the issue and restored normal operation.

Validation Process

Before broad deployment, Brillium followed a controlled release methodology to minimize customer risk.

The corrective solution was:

  • Tested in non-production environments.
  • Subjected to high-load validation scenarios.
  • Released to a limited group of customers.
  • Monitored for stability and performance.
  • Verified through operational metrics and customer observations.

The results confirmed that the identified issue was resolved.

Long-Term Improvements

While the immediate issue has been addressed, Brillium is implementing a strategic modernization initiative designed to eliminate this category of issue entirely.

Planned Platform Enhancement

Brillium will replace the remaining legacy session-management architecture with a modern stateless architecture.

A stateless design eliminates dependency on centralized session coordination mechanisms and provides significant operational benefits, including:

  • Greater scalability under heavy usage
  • Improved platform resilience
  • Faster recovery during traffic spikes
  • Reduced risk of session-related bottlenecks
  • Improved operational efficiency
  • Enhanced long-term reliability

This modernization effort aligns with Brillium's broader platform scalability and cloud modernization roadmap.

SOC 2 Compliance Considerations

Brillium maintains its commitment to SOC 2 security, availability, and operational control requirements.

During this incident:

✅ Incident management procedures were followed.

✅ Monitoring and alerting controls functioned as designed.

✅ The issue was investigated, documented, and remediated.

✅ Corrective actions were validated prior to broader deployment.

✅ No customer data confidentiality concerns were identified.

✅ No unauthorized access occurred.

✅ No customer records were lost, altered, or corrupted.

✅ Post-incident corrective and preventive actions have been established and tracked.

This post-incident review is part of Brillium's continuous improvement process and commitment to transparency with customers and partners.

Frequently Asked Questions

Was customer data compromised?

No. Customer data remained secure throughout the event. The issue affected application availability only and did not impact data confidentiality, integrity, or security.

Did assessments or results get lost?

No. Assessment content, candidate information, and reporting data were preserved and remained intact.

Has the problem been fixed?

Yes. A corrective solution has been implemented, tested, validated, and deployed. Continuous monitoring confirms normal platform operation.

Could this happen again?

The immediate issue has been resolved. In addition, Brillium is undertaking a strategic architectural upgrade to remove the legacy session-management dependency entirely, significantly reducing the likelihood of similar issues in the future.

What is Brillium doing differently going forward?

Brillium is accelerating modernization efforts focused on:

  • Stateless application architecture
  • Increased scalability
  • Improved resilience during peak usage
  • Enhanced operational monitoring
  • Continuous performance testing under high-load conditions

Closing Statement

Brillium recognizes the importance of platform reliability to our customers' business operations and assessment programs. We sincerely apologize for the disruption experienced by affected users.

Our teams responded quickly, identified the root cause, implemented a validated solution, and have initiated long-term architectural improvements designed to further strengthen platform performance and scalability. We remain committed to transparency, operational excellence, and the secure delivery of our services.

Brillium Operations & Engineering Team
September 2026

Posted Sep 30, 2026 - 11:24 EDT

Resolved

All customers that experienced issues have been remediated and tested.
Posted Sep 29, 2026 - 20:12 EDT

Update

The issue has been identified and the team is applying remediation at this time. We expect all instance issues for affected customers to be resolved shortly.
Posted Sep 29, 2026 - 14:17 EDT

Update

The team has rulled out some possible causes for the small number of customers experiencing access issues. They are continuing the investigation and will provide additional updates shortly.
Posted Sep 29, 2026 - 13:28 EDT

Update

The team is still working to resolve issue experienced by a few customer instances.
Posted Sep 29, 2026 - 12:55 EDT

Update

The team is still working to resolve issue experienced by a few customer instances.
Posted Sep 29, 2026 - 11:58 EDT

Identified

We are aware that a small number of customers are currently experiencing service disruptions and may be unable to access the platform. Our team has identified the issue and is aggressively working to restore service as quickly as possible.

We understand the impact this may have on your operations and are treating this as our highest priority. We will continue to provide updates as more information becomes available.

Thank you for your patience and understanding while we work to resolve this issue.
Posted Sep 29, 2026 - 10:58 EDT
This incident affected: API & Integrations (API, Zapier Integration) and Assessment Builder (Assessment Authoring, Assessment Delivery).