Software FMEA per SAE J1025: a dedicated framework for software risk

Hardware wears out, software does not, and it still fails. SAE J1025 is the first FMEA standard to give software risk its own type of analysis: the Software FMEA (SwFMEA). This article shows how it assesses Occurrence in a structured way, where the failure modes come from, and how it supports the evidence for ISO 26262 and MIL-STD-882E.

Comparison of hardware and software logic in the FMEA per SAE J1025

In vehicles, control units and medical devices, most of the functional scope now sits in software. Yet many teams still assess this software with failure modes such as fracture, wear and failure rate that were actually meant for sheet metal and shafts. Software does not wear out. It carries latent defects that only fire under certain inputs and states, and it therefore needs its own risk logic.

After this article you will know why SAE J1025 treats the Software FMEA as its own type, how Occurrence is assessed for software in a structured way, and how the analysis supports the evidence for ISO 26262 and MIL-STD-882E.

Why software needs its own type of FMEA


Hardware fails because parts change over time. The failure follows a failure rate, mixes random and systematic causes, and is detected through inspection and measurement. Time-based ageing dominates the picture. The classic DFMEA is built for that world.
Software does not wear out. What sits inside it are latent defects, present from day one and triggered only by certain inputs and states. The AIAG-VDA handbook absorbs software into the DFMEA. SAE J1025 takes the other route and carves it out as its own type of analysis, so that state, sequence, timing and data become the leading categories. How far apart the two ways of thinking are is shown by the comparison in Figure 1.

Figure 1: Comparison: hardware logic (DFMEA) with failure rate, wear and fatigue versus software logic (SwFMEA) with state, sequence, timing and data

Occurrence becomes a structured estimate


Figure 2: Formula: Occurrence = Existence × Manifestation × Prevention × Detection / 4 with a short explanation of the four factors

The core of SAE J1025 lies in the assessment of Occurrence (O). In a classic FMEA a team often says "that is a 4" and means a gut feeling. For software, SAE J1025 instead defines Occurrence as a composite of four separately assessable factors:Occurrence = Existence × Manifestation × Prevention × Detection / 4

Existence asks whether the defect plausibly exists in the code. Manifestation asks under what fraction of inputs and states it shows. Prevention stands for the preventive controls from coding standards, reviews and static analysis. Detection stands for the detection controls from unit, integration and system testing. The result sits on a scale of 1 to 5 and is combined with Severity (S) in an [SO] matrix of Severity and Occurrence. No Risk Priority Number (RPN), no Action Priority (AP): criticality becomes provable factor by factor.

The six-step process: same shape, different content


Anyone who knows the FMEA per AIAG-VDA will find their way around the SwFMEA quickly.SAE J1025 runs through six steps:
  1. Planning
  2. Preparation and scope
  3. Technical risk analysis
  4. Criticality analysis
  5. Risk reduction actions
  6. Updates and version control

The shape resembles the DFMEA, the content is different.

The analysis draws on the design flows of state, sequence, timing and data, together with the software architecture and its interfaces, the Software Operating Environment (SOE) and the CWE. What comes out are failure modes per design-flow element, an [SO] criticality with high, medium and low bands, recommended design, test and process actions, and an evidence trail for ISO 26262 and MIL-STD-882E. For many projects this evidence trail is the real gain, because the SwFMEA supplies the proof that a safety case will demand later anyway.

CWE as the source of failure modes: the bridge to safety and security


SAE J1025 explicitly references the Common Weakness Enumeration (CWE) from MITRE as a source of software failure modes. The failure modes therefore do not have to be invented. The CWE is a maintained, community-curated catalogue of software weaknesses. The team no longer works on the question of which faults could exist at all, but focuses on relevance, severity and detection.

With this, the SwFMEA connects two worlds that run separately in many organisations. On the functional-safety side it supplies the evidence layer for the Level of Rigor per MIL-STD-882E and for the ASIL decomposition per ISO 26262. On the weakness side, the CWE brings the security case and the safety case into the same assessment.

Three patterns that break a software FMEA


Occurrence from gut feeling. If Occurrence is estimated as a single number, the rating stays unproven. The countermeasure lies in the decomposition into Existence, Manifestation, Prevention and Detection.

Reusing mechanical failure modes for code. Fracture, wear or short circuit do not fit software functions. Software failure modes come from the CWE and from the analysis of the design flows in state, sequence, timing and data.

Doing the SwFMEA on the side. The hardware DFMEA team cannot take on the software analysis in passing. Software needs its own cross-functional team of software architects, test engineers and a safety expert.

Example


A team assesses an input validation per SAE J1025. Existence is high, because the affected path demonstrably accepts unchecked values. Manifestation is low, because only a narrow value range triggers the fault. Prevention is in the middle, because a coding standard applies but a static analysis run is missing. Detection depends on the test depth that covers the path. From these four justified individual judgements comes an Occurrence value that an auditor can follow, instead of having to accept a single number from gut feeling.

A second example shows where the failure modes come from. A counter that overflows at an unusually large input value is not a mechanical failure mode but a known weakness from the CWE (integer overflow). The team selects it from the catalogue and relates it to its own function, instead of inventing a new failure picture.

Result


Software fails differently from hardware. Treating it with a dedicated FMEA per SAE J1025 finds risks that a DFMEA alone misses, and at the same time delivers the evidence for functional safety.

Do you want to embed software risk into your FMEA in a method-safe way? Talk to us about a SwFMEA pilot analysis or the right DC module: www.dietz-consultants.com

Author: Winfried Dietz, CEO of Dietz Consultants GmbH
Winfried Dietz is CEO of Dietz Consultants GmbH and has supported development organisations worldwide for more than 30 years in adopting and maturing the FMEA methodology. He is a trainer, author and speaker for FMEA per AIAG-VDA. Find more from the FMEA Quick Tip series on LinkedIn
.