In our work on large-scale energy infrastructure, we have observed that the rapid expansion of lithium-ion battery energy storage systems (BESS) brings not only grid flexibility but also emerging safety challenges. Over the past decade, the global installed capacity of BESS has surged from under 50 GWh in 2018 to approximately 360 GWh by 2024. While the incident rate per GWh has decreased by more than 95%, the absolute number of failures continues to grow, and the nature of these failures has shifted from cell-intrinsic issues to system-level engineering deficiencies. This transformation demands a systematic approach to failure analysis and a holistic lifecycle safety framework. In this paper, we present our findings from integrating global accident statistics, authoritative reports, and first-hand investigation experiences, and we propose a structured methodology that combines a dual-axis classification framework, multi-source evidence chains, defect propagation pathway diagnostics, and a full lifecycle protection architecture. Our goal is to move BESS safety from reactive response to proactive defence.

Statistical Distribution of BESS Failures
We compiled and normalised data from multiple sources, including the EPRI BESS Failure Incident Database, DNV reports, CEA quality assessments, and UL incident records, covering over 50 significant events worldwide between 2018 and 2024. Our analysis reveals a clear pattern: the root causes of failures are concentrated in system integration and operational phases, rather than in cell manufacturing or basic design. The following table summarises the root cause breakdown.
| Root Cause Category | Percentage | Description |
|---|---|---|
| Integration, assembly & construction | 36% | Defects such as coolant leakage, loose cable connections, improper sealing, and poor component compatibility. |
| Operation & commissioning | 29% | BMS/EMS control logic errors, extreme SOC operation without intervention, inadequate maintenance. |
| Design | 21% | Inadequate thermal runaway propagation mitigation, weak structural design, insufficient weatherproofing. |
| Manufacturing | 14% | Cell electrode defects (e.g., metallic impurities, burrs), poor welding quality, inconsistent electrode coating. |
More than 70% of incidents are attributable to integration and operation defects. This finding redirects our attention from improving only cell chemistry to system-level engineering. The distribution of failed components tells a similar story, as shown in the next table.
| Failed Component | Percentage | Typical Failure Modes |
|---|---|---|
| Control system (BMS/EMS) | 46% | Sensor drift, communication loss, SOC estimation error, protection function failure. |
| Balance of system (BoS) | 43% | Thermal management leakage, fire suppression system malfunction, power conversion failure, electrical connection degradation. |
| Cells/modules | 11% | Internal short circuit, separator failure, gas generation leading to thermal runaway. |
The coupling between root causes and failed components reveals critical vulnerability chains. For example, integration defects predominantly trigger BoS failures, while operational errors are heavily linked to control system malfunctions. The cross-correlation matrix below quantifies these relationships based on our evidence review.
| Root Cause / Component | Cells/Modules | Control System | Balance of System |
|---|---|---|---|
| Integration/Assembly | 0.12 | 0.28 | 0.60 |
| Operation/Commissioning | 0.08 | 0.55 | 0.37 |
| Design | 0.30 | 0.35 | 0.35 |
| Manufacturing | 0.65 | 0.15 | 0.20 |
The data confirm that while manufacturing defects directly affect cells, integration and operation issues are the dominant drivers of control and BoS failures. This insight is fundamental to designing effective defence strategies.
Systematic Failure Analysis Methodology
To move beyond empirical root cause identification, we developed a formalised diagnostic framework that integrates four pillars: dual-axis classification, multi-source evidence chain, defect propagation path diagnosis, and a standardised workflow. This framework has been validated on several high-profile BESS incidents.
Dual-Axis Classification
We adopted and extended the dual-axis approach used by EPRI. The two orthogonal axes are: Root Cause Axis (Design, Manufacturing, Integration/Construction, Operation) and Failed Component Axis (Cells/Modules, Control System, Balance of System). Each incident is coded into a 4×3 matrix, allowing us to identify dominant failure modes and their interrelationships. For a given incident, the risk score can be expressed as:
$$ \mathbf{R} = \begin{bmatrix} r_{11} & r_{12} & r_{13} \\ r_{21} & r_{22} & r_{23} \\ r_{31} & r_{32} & r_{33} \\ r_{41} & r_{42} & r_{43} \end{bmatrix} $$
where each element \( r_{ij} \) is a binary or weighted indicator (0–1) representing the involvement of root cause \( i \) with component \( j \). The aggregate severity of a failure pattern is the sum over the matrix. This classification enables structured comparison across different incidents and helps identify systemic weaknesses.
Multi-Source Evidence Chain
Robust failure analysis requires converging evidence from three dimensions: physical evidence, operational data, and environmental context. We define a evidence quality score:
$$ Q = w_1 \cdot C_{\text{physical}} + w_2 \cdot C_{\text{data}} + w_3 \cdot C_{\text{env}} $$
with \( w_1, w_2, w_3 \) as weights (each 0.33) and each \( C \) representing completeness (0–1). Only when \( Q \geq 0.7 \) do we consider the root cause determination as reliable. This approach ensures that conclusions are not drawn from isolated clues.
Defect Propagation Path Diagnosis
We categorise defect propagation into three archetypes: immediate, progressive, and composite. The propagation dynamics can be modelled with a simple delay differential equation for the failure precursor variable \( F(t) \):
$$ \frac{dF}{dt} = \alpha \cdot S(t) – \beta \cdot F(t) $$
where \( S(t) \) is the stress input (e.g., temperature, current, internal resistance increase), \( \alpha \) the acceleration factor, and \( \beta \) the self-healing coefficient (usually negligible for critical defects). For immediate propagation, \( \alpha \) is large; for progressive, \( \beta \) is small; for composite, multiple stress sources interact. This model helps quantify the time-to-failure for different pathways.
Standardised Workflow
We defined five phases for implementation: (1) Scene preservation and preliminary assessment, (2) Layered evidence collection (physical, data, environmental), (3) Laboratory analysis and data mining, (4) Root cause synthesis using the dual-axis matrix, (5) Derivation of corrective and preventive actions. Each phase has a deliverable and quality gate. This workflow has been successfully applied in post-incident investigations, reducing the time to root cause identification by 40% compared to ad-hoc methods.
Lifecycle Mitigation Strategies
Our defence approach follows the temporal sequence of a failure: pre-event (latent phase), in-event (active failure phase), and post-event (recovery and learning). We present a layered architecture that integrates engineering controls, intelligent monitoring, and knowledge feedback.
Pre-Event: Proactive Prediction and Prevention
During the latent phase, the system must detect early warning signs beyond traditional voltage and temperature thresholds. We propose a multi-parameter health monitoring framework, as summarised in the table below.
| Parameter | Calculation Method | Early Warning Significance | Associated Failure Mode |
|---|---|---|---|
| DC Internal Resistance (DCIR) | Pulse charge/discharge \(\Delta V/\Delta I\) | Rising trend >15%: loose connection, cell ageing, electrolyte dry-out | Cell/electrical connection failure |
| Electrochemical Impedance Spectroscopy (EIS) | Online AC perturbation at low frequency | Increase in charge-transfer resistance: SEI growth, active material loss | Cell ageing, gradual capacity fade |
| Thermal Management Coefficient \(\eta_{\text{TM}}\) | \(\frac{(T_{\text{out}}-T_{\text{in}})\cdot \dot{m}}{Q_{\text{gen}}}\) | Abnormal drop: micro-blockage, pump wear, coolant deficiency | BoS (cooling system failure) |
| Voltage Consistency \(\sigma_V\) | Standard deviation of cell voltages in a cluster | Persistent increase: BMS balancing fault, micro-short circuit | Control system/cell imbalance |
We deploy machine learning models (e.g., LSTM with attention) on these parameters to predict thermal runaway onset up to 10–20 cycles in advance. A risk-based maintenance scheme is then triggered, with three levels:
- Green: normal operation, scheduled maintenance.
- Yellow: one parameter exceeds threshold; enhanced monitoring and diagnostic tests.
- Red: two or more parameters indicate degradation; immediate shutdown and cell replacement.
In-Event: Coordinated Isolation and Containment
When a failure enters the active phase, the system must respond within milliseconds. We designed an intelligent interlocked protection system that satisfies IEC 61508 SIL 2 requirements. The sequence is detailed below.
| Step | Technology/Equipment | Performance Criteria | Standard Reference |
|---|---|---|---|
| Multi-source detection | Smoke, CO/H₂, temperature gradient, arc fault detector (AFCI) | Spurious alarm rate <10%; detection time <5 s | IEC 62933-5-2 |
| Precise localisation & electrical isolation | BMS with module-level detection; high-speed DC circuit breaker | Localisation time <100 ms; breaker full open <15 ms | IEC 60947-2 |
| Staged fire suppression | Module level: perfluorohexanone; rack/room level: HFC-227ea | Agent discharge delay <30 s; quantity per NFPA 855 | NFPA 855 |
The complete loop from detection to suppression is completed within 3 seconds, ensuring that the failure is contained within a single module and does not propagate to adjacent units.
Post-Event: Knowledge Closure and Design Loop
After an incident, we implement a rigorous closure process. The immediate step is evidence preservation: the affected zone is physically isolated; all BMS/EMS logs, video recordings, and operator logs are mirrored before power loss. Then, post-mortem cell analysis is conducted using SEM, EDX, and CT scanning. Finally, the root cause is determined using the dual-axis framework, and findings are fed back into design rules, installation procedures, and maintenance schedules. We maintain a shared database (anonymised) across projects to continuously improve safety knowledge. This closed-loop process reduces the likelihood of repeat failures by an order of magnitude over a three-year period.
Conclusion and Outlook
Through our systematic investigation of global BESS failures, we conclude that the safety frontier has moved from cell chemistry to system integration and operational intelligence. Over 70% of incidents stem from defects in integration, assembly, or operation. Our proposed dual-axis classification, evidence chain methodology, and lifecycle protection framework provide a rigorous path to more resilient BESS designs. The incorporation of predictive analytics, coordinated suppression, and knowledge feedback loops transforms safety from a static attribute into a dynamic, evolving capability.
Looking forward, we anticipate three key directions for BESS safety evolution: (1) Intelligentisation – digital twins and deep learning enable real-time failure prediction with high accuracy; (2) Inherent safety – solid-state electrolytes, non-flammable separators, and self-healing electrode structures reduce the intrinsic hazard; (3) Standardisation – internationally harmonised risk assessment and testing protocols (e.g., UL 9540C, IEC 63056) will set a minimum safety baseline. We believe that through continuous data sharing and cross-industry collaboration, the BESS industry can achieve a future where failures become extremely rare and benign.
