If you write code that sits close to a plant's control or shutdown logic (PLC programs, firmware in a field transmitter, DCS or safety system configuration, or the tooling around them), you will eventually see a requirement like "this function must be SIL 2." The term gets used loosely, and the loose usage causes real engineering mistakes.
This guide explains what a safety integrity level actually is, how it is determined and verified, and what it means for the people who build and maintain software in oil and gas, chemical, and other process industries.
What a safety integrity level is
A safety integrity level (SIL) is one of four discrete levels, SIL 1 to SIL 4, that express how reliably a safety function must perform when it is called upon. A higher level means a lower tolerated probability of the function failing to act.
Two standards define the framework:
IEC 61508 is the generic functional safety standard for electrical, electronic and programmable electronic systems.
IEC 61511 applies the same ideas to the process industry for the people who specify, design, integrate and operate safety systems. In the US it is adopted as ANSI/ISA-84. It addresses SIL 1 to 3.
The detail most often missed: a SIL is assigned to a safety instrumented function (SIF), not to a product. A SIF is the complete chain, for example "close shutdown valve XV-101 within 5 seconds if vessel pressure exceeds the high-high setpoint." That chain includes the sensor, the logic solver, the final element, and the software running in between. Individual devices can be described as SIL capable based on certification and failure data, but only the whole function achieves a SIL.
What the SIL numbers mean
For functions that are rarely demanded (the "low demand" mode, roughly no more than once a year), the target is the average probability of failure on demand (PFDavg). For continuous or high-demand functions, it is the probability of dangerous failure per hour (PFH).
SIL PFDavg (low demand) Risk reduction factor PFH (high demand / continuous)
1 ≥ 10⁻² to < 10⁻¹ 10 to 100 ≥ 10⁻⁶ to < 10⁻⁵ per hour
2 ≥ 10⁻³ to < 10⁻² 100 to 1,000 ≥ 10⁻⁷ to < 10⁻⁶ per hour
3 ≥ 10⁻⁴ to < 10⁻³ 1,000 to 10,000 ≥ 10⁻⁸ to < 10⁻⁷ per hour
4 ≥ 10⁻⁵ to < 10⁻⁴ 10,000 to 100,000 ≥ 10⁻⁹ to < 10⁻⁸ per hour
How the required SIL is determined
The required level is not chosen from a menu. It comes out of a risk assessment, typically in this order:
Identify hazards. A HAZOP study walks through the process deviation by deviation (more pressure, no flow, wrong composition) and records credible causes and consequences.
Set a tolerable risk target for each hazardous scenario, based on the owner's criteria and applicable regulation.
Credit the other protection layers. The basic process control system, alarms with operator response, relief valves, and physical containment each reduce risk, but only if they are genuinely independent of the initiating event.
Calculate the gap. In a layer of protection analysis (LOPA), you multiply the initiating event frequency by the failure probabilities of the independent layers and compare the result with the tolerable frequency. The remaining gap is the risk reduction the SIF must deliver, which maps to a SIL target. Risk graphs and risk matrices are used for simpler or earlier screening.
The independence rule in step 3 matters to software people. If the control loop that causes the upset is the same one you are crediting as a protection layer, it cannot be counted.
How the achieved SIL is verified
Once a SIF is designed, you check that it meets its target. The calculation depends on:
the dangerous undetected failure rate of each component (λDU),
the proof test interval and how much of the failure modes the proof test actually reveals,
the architecture (1oo1, 1oo2, 2oo3 and so on),
diagnostic coverage, and
a common cause factor for redundant channels.
Architectural constraints (hardware fault tolerance and safe failure fraction) then limit the SIL you can claim regardless of how good the probability number looks.
Here is the simplest possible case, a single-channel (1oo1) function, in Python. The numbers are illustrative, not real device data:
def pfd_avg_1oo1(lambda_du_per_hr: float, proof_test_interval_hr: float) -> float:
"""Simplified PFDavg for a 1oo1 function: lambda_DU * TI / 2.
Ignores diagnostics, repair time, and proof test imperfection."""
return lambda_du_per_hr * proof_test_interval_hr / 2
def sil_band_low_demand(pfd: float) -> str:
bands = [(1e-5, 1e-4, "SIL 4"), (1e-4, 1e-3, "SIL 3"),
(1e-3, 1e-2, "SIL 2"), (1e-2, 1e-1, "SIL 1")]
for low, high, label in bands:
if low <= pfd < high:
return label
return "outside SIL 1-4 bands"
pfd = pfd_avg_1oo1(lambda_du_per_hr=1e-6, proof_test_interval_hr=8760) # annual test
print(f"PFDavg = {pfd:.2e} -> {sil_band_low_demand(pfd)}")
PFDavg = 4.38e-03 -> SIL 2
Real verification uses vendor safety data, the full set of formulas from the standards (or a validated tool), and documented assumptions about demand rate and proof testing. If you want the consultancy view of how that workflow is run on projects, this reference page on Safety Integrity Level assessment and verification is one place to look.
What the number does not cover
The probability calculation covers random hardware failures. It says nothing about systematic failures: a wrong setpoint in the specification, a logic error, a misunderstood requirement, a bad modification. Those are controlled by process, not by arithmetic. That is why the standards are organised around a safety lifecycle: hazard and risk assessment, a safety requirements specification (SRS), design, verification, installation and validation, operation and proof testing, management of change, and eventually decommissioning.
What this means if you write the code
Know which standard your code falls under. Application logic on a certified safety logic solver is generally governed by IEC 61511, which favours limited-variability languages such as ladder logic and function block diagrams. Device-level embedded firmware is generally governed by IEC 61508-3, where the required development techniques scale with the SIL.
Trace everything to the SRS. Each safety function should have a written requirement, and each test should trace back to one.
Keep safety logic simple and separate. Avoid mixing unrelated functionality into the safety application, and avoid sharing components with the basic control system in ways that undermine independence.
Treat changes as safety events. A "small" trip setpoint tweak or a temporary bypass goes through management of change and re-validation, because the verification assumptions may no longer hold.
Do not label code or products as "SIL 3" on their own. Claims belong to the function and the lifecycle evidence behind it.
Common pitfalls
Treating SIL as a product label. Buying a "SIL 3 valve" does not give you a SIL 3 function.
Over-specifying "to be safe." A higher SIL than the risk requires adds cost, complexity and a heavier proof testing burden, and more complexity can introduce more systematic faults.
Ignoring operational assumptions. If the calculation assumes an annual proof test with a given coverage and the plant tests every three years, the claimed SIL no longer holds.
Crediting non-independent layers in the LOPA.
Skipping validation of the whole function against the original requirement, not just component checks.
Key takeaways
A safety integrity level is a target for a safety function, not a rating for a device or a piece of software.
The required SIL comes from a documented risk assessment (HAZOP, then LOPA or a risk graph), and the achieved SIL is verified by calculation plus architectural constraints.
Probability maths covers random hardware faults; lifecycle discipline covers systematic faults, which is where most software risk sits.
If you touch safety logic, learn the relevant standard, work from the SRS, and treat every change as a safety-relevant event.
Conclusion
SIL is easier to work with once you see it as a chain of evidence: a hazard, a risk target, a function designed to close the gap, and calculations plus lifecycle records showing it does. Engineers who build control and safety software do not need to become risk assessors, but understanding that chain helps you ask better questions of the specification and write code that holds up under audit.
Top comments (0)