Calculator D3

Incident Investigation Root Cause Analysis Using Bow-Tie Diagrams

A bow-tie diagram is a visual tool that shows how a dangerous event (like a chemical leak) can start, what could go wrong, and what safety barriers stop it β€” like a bow tie with causes on the left and consequences on the right.

Industry Adoption
Used by >85% of OSHA-covered refineries and petrochemical plants per CCPS 2023 benchmark
Regulatory Basis
Explicitly referenced in EPA 40 CFR Β§68 and OSHA 29 CFR 1910.119 Appendix E
Typical Scale
1 bow-tie per major process unit (e.g., distillation column, reactor, storage tank)
Validation Requirement
Must be reviewed annually and after every incident or MOC per API RP 754 Sec 5.3.2

⚠️ Why It Matters

1
Inadequate barrier identification
2
Unrecognized common-cause failures
3
Barrier degradation without monitoring
4
Escalation to major incident
5
Regulatory citation or catastrophic release

πŸ“˜ Definition

Bow-tie analysis is a structured risk assessment methodology that decomposes a hazardous event into its initiating causes (left side) and potential consequences (right side), connected by a central 'top event', with preventive and mitigative barriers explicitly mapped across the 'knot'. It integrates principles from fault tree and event tree analysis while emphasizing barrier integrity, verification, and human/organizational factors. As a semi-quantitative root cause investigation technique, it supports regulatory compliance (e.g., OSHA PSM, EPA RMP) and process safety management in high-hazard industries.

🎨 Concept Diagram

Ignition SourceVessel OverpressureToxic ReleasePRVScrubberBow-Tie DiagramLeft: Causes | Center: Top Event | Right: Consequences

AI-generated illustration for visual understanding

πŸ’‘ Engineering Insight

A bow-tie diagram is only as strong as its weakest barrier β€” not its most sophisticated one. In practice, the most frequently bypassed barrier is rarely the instrumented one; it’s the procedural step skipped during night shift handover. Always validate barrier usage patterns via work observation, not just design intent.

πŸ“– Detailed Explanation

At its core, a bow-tie diagram organizes hazard understanding spatially: left side captures how things go wrong (initiators), the center knot is the critical hazardous event (e.g., 'vessel overpressure'), and the right side maps what happens if nothing stops it (consequences). Barriers β€” like pressure relief valves or operator alarms β€” are drawn as horizontal lines crossing the 'bow tie', separating causes from effects.

Beyond visualization, bow-tie analysis forces explicit classification of barrier type (hardware, software, procedural, administrative), verification method (functional test, audit, simulation), and ownership (e.g., 'Instrument Tech – quarterly proof test'). This enables systematic gap analysis: if a barrier lacks defined verification frequency or has no owner, it is de facto absent β€” regardless of whether it appears on P&IDs.

Advanced application includes dynamic bow-ties integrated with digital twin platforms, where real-time sensor data (e.g., valve position feedback, DCS alarm flood logs) automatically update barrier status flags. When combined with Bayesian updating, these models quantify time-dependent barrier reliability and predict optimal inspection windows β€” moving beyond static 'snapshot' assessments to predictive process safety assurance.

πŸ”„ Engineering Workflow

Step 1
Step 1: Define Top Event using PHA output (HAZOP/LOPA) and verify against incident database
β†’
Step 2
Step 2: Map all credible initiating causes using fault tree logic and historical near-miss data
β†’
Step 3
Step 3: Identify and classify barriers (preventive/mitigative) per ISA 84.01 / IEC 61511 layers of protection
β†’
Step 4
Step 4: Assess barrier integrity via BER, BII, TTF, and HEP using CCPS Barrier Management Toolkit
β†’
Step 5
Step 5: Perform barrier interaction analysis to detect latent weaknesses (e.g., shared power supply, common training gaps)
β†’
Step 6
Step 6: Validate findings with cross-functional team (operations, maintenance, engineering, SHE)
β†’
Step 7
Step 7: Integrate verified actions into MOC, P&ID updates, and management-of-change tracking system

πŸ“‹ Decision Guide

Rock/Field Condition Recommended Design Action
BER < 2.5 AND BII < 0.7 Replace with engineered barrier (e.g., automated shutdown) and redesign layout for physical separation
TTF < 18 months AND HEP > 0.1 Implement real-time barrier health monitoring + cognitive workload reduction in DCS interface
Multiple initiating causes share same root cause (e.g., calibration drift in two sensors) Conduct causal factor charting per API RP 754 and revise calibration management system

📊 Key Properties & Parameters

Barrier Effectiveness Rating (BER)

2.0–4.5 (unitless)

A 1–5 scale quantifying the reliability of a safety barrier based on design, testing frequency, and proven performance history.

⚡ Engineering Impact:

Directly determines required redundancy and inspection intervals per CCPS guidelines.

Time-to-Failure (TTF)

12–72 months for mechanical SIS; 3–18 months for procedural controls

Estimated duration between barrier degradation onset and functional failure under normal operating conditions.

⚡ Engineering Impact:

Drives maintenance scheduling, P&ID revision cycles, and SIL verification timelines.

Barrier Independence Index (BII)

0.6–1.0 (unitless, where 1.0 = fully independent)

Measure of functional and physical separation between redundant barriers to prevent common-cause failure.

⚡ Engineering Impact:

Determines whether dual-pressure transmitters on a relief valve require separate impulse lines or shared manifold design.

Human Error Probability (HEP)

1Eβˆ’2 to 5Eβˆ’1 (per task)

Likelihood of operator action failure during barrier activation or recovery, derived from THERP or HCR methods.

⚡ Engineering Impact:

Triggers requirement for alarm rationalization, SOP simplification, or control room ergonomics upgrades.

πŸ“ Key Formulas

Barrier Reliability Index (BRI)

BRI = (BER Γ— BII) / (1 + log₁₀(HEP + 0.01))

Composite metric scoring overall barrier robustness; values < 2.0 indicate urgent remediation priority.

Variables:
Symbol Name Unit Description
BRI Barrier Reliability Index unitless Composite metric scoring overall barrier robustness; values < 2.0 indicate urgent remediation priority
BER Barrier Effectiveness Ratio unitless Dimensionless measure of barrier performance relative to design intent
BII Barrier Integrity Index unitless Dimensionless score reflecting physical condition and maintenance status of the barrier
HEP Human Error Probability unitless Estimated probability of human-induced failure affecting barrier function
Typical Ranges:
Critical Safety Instrumented Function (SIF)
3.1 – 4.7
Procedural Barrier (e.g., permit-to-work)
1.2 – 2.8
⚠️ BRI β‰₯ 3.0 for SIL-2+ systems; β‰₯ 2.5 for non-SIL but essential barriers

Common-Cause Failure Probability (CCFP)

CCFP = 1 βˆ’ exp(βˆ’Ξ»_cc Γ— t)

Probability that two or more barriers fail simultaneously due to shared root cause (e.g., power loss, software bug, training deficiency).

Variables:
Symbol Name Unit Description
CCFP Common-Cause Failure Probability dimensionless Probability that two or more barriers fail simultaneously due to a shared root cause
Ξ»_cc Common-Cause Failure Rate 1/time Failure rate attributable to common causes (e.g., per hour or per year)
t Time time Exposure time or mission duration
Typical Ranges:
Redundant sensors sharing same IO module
1Eβˆ’3 to 5Eβˆ’2 per year
Independent hardware with segregated power & networks
1Eβˆ’6 to 1Eβˆ’5 per year
⚠️ CCFP < 1Eβˆ’4/year for any pair of Level 2+ barriers per API RP 754

🏭 Engineering Example

ExxonMobil Baton Rouge Refinery – Alkylation Unit

N/A (chemical process system)
Top Event
HF Acid Release > 100 lb
Initiating Causes
3 confirmed corrosion-induced pipe leaks (2019–2022)
Mitigative Barrier BII
0.62 (shared emergency response radio channel)
Preventive Barrier BER
2.8 (ultrasonic thickness monitoring program)
HEP for Evacuation Initiation
0.23 (observed during 2021 drill)

πŸ—οΈ Applications

  • Process Hazard Analysis (PHA) follow-up
  • Incident root cause validation
  • Management of Change (MOC) impact assessment
  • SIL verification and audit preparation

πŸ“‹ Real Project Case

Ammonia Refrigeration System PHA & LOPA Integration at Midwest Food Plant

Retrofit of legacy ammonia refrigeration system serving 300k sq ft food processing facility

Challenge: Outdated PHA documentation; no SIL verification for emergency shutdown valves
HAZOP WorkshopCross-functional teamLOPA AnalysisIPL VerificationSIS ArchitectureIEC 61511 CompliantPFD = 0.0023SIL 2 ConfirmedAmmonia Refrigeration SystemMidwest Food Plant β€’ PHA & LOPA Integration
Read full case study β†’

🎨 Technical Diagrams

Cause ATOP EVENTConseq BPreventiveMitigative
BER=3.2BII=0.71HEP=0.08BRI = 3.2Γ—0.71 / (1+log₁₀(0.08+0.01))= 2.27 / (1βˆ’1.05) β†’ INVALID

πŸ“š References

[1]
Guidelines for Hazard Evaluation Procedures β€” Center for Chemical Process Safety (CCPS)
[2]
API RP 754 – Process Safety Performance Indicators β€” American Petroleum Institute
[3]
IEC 61511-1:2016 Functional safety – Safety instrumented systems β€” International Electrotechnical Commission