Chapter 21 / 36

Diagnose with alarms

Capture useful fault evidence and separate acknowledgement from reset.

An alarm should earn the operator's attention

A screen fills with messages: motor stopped, belt stopped, carton absent, downstream not ready, cycle incomplete. One air-pressure failure caused all five. The messages are technically related to the machine, but they do not help the operator find the first useful action.

An alarm should identify an abnormal condition that requires an operator response. A routine state change can be an event. “Waiting for next carton” can be status. When everything is called an alarm, priority stops communicating anything.

Write each alarm as a short operational contract: trigger, consequence, required response, priority, delay, clearing condition and reset behaviour. If nobody can name a useful response, reconsider whether the condition belongs in the alarm list or in a diagnostic log.

Qualification comes before latching

A pressure switch flickers for one scan during a normal valve transition. An immediate alarm may be too sensitive. A five-second delay may be too slow. Choose the qualification from the physical behaviour and the time available for the required response.

For an analog threshold, hysteresis can prevent chatter. An illustrative cooling alarm becomes active above 75 degrees Celsius and returns inactive below 72. Between those values it keeps its previous condition. Those numbers are teaching choices, not general temperature limits.

Delay and hysteresis solve different problems. Delay rejects short excursions; hysteresis separates activation and clearing thresholds. Filtering the measured signal adds yet another delay. Account for their combined effect instead of adding each independently until the screen looks quiet.

Preserve invalid-measurement handling. A bad sensor must not be silently treated as “temperature below the alarm threshold.” Otherwise the temperature alarm disappears precisely when you lose the ability to assess temperature.

Active, acknowledged and cleared are different facts

The process condition can be active or inactive. The event can be acknowledged or unacknowledged. A remembered trip can remain latched after the process recovers. These dimensions explain why a single Boolean called Alarm is often insufficient.

Consider low air pressure on a clamping station. Pressure drops, the process holds and an alarm appears. The operator acknowledges it: the message has been seen, but pressure is still low. Pressure returns: the initiating condition clears, but the station may still need a deliberate recovery check. Reset then clears the eligible process latch. Start remains a separate request.

That separation prevents a common surprise: pressing Acknowledge accidentally restarts equipment. It also keeps the event history truthful. An alarm that is no longer active may still need acknowledgement because the operator has not yet seen it.

Trace a deliberately simple alarm

This cyclic ST excerpt demonstrates one latching policy. QualifiedCondition already includes the chosen validity, delay and hysteresis rules. AckPulse and ResetPulse are one-cycle requests. Variables are persistent across scans, not automatically retained across power loss. A production alarm system also needs identifiers, timestamps, permissions and restart rules.

NewOccurrence := QualifiedCondition AND NOT PreviousCondition;

IF NewOccurrence THEN
    AlarmLatched := TRUE;
    AlarmAcknowledged := FALSE;
END_IF;

IF AckPulse AND AlarmLatched AND NOT NewOccurrence THEN
    AlarmAcknowledged := TRUE;
END_IF;

IF ResetPulse
    AND AlarmLatched
    AND AlarmAcknowledged
    AND NOT QualifiedCondition THEN
    AlarmLatched := FALSE;
    AlarmAcknowledged := FALSE;
END_IF;

PreviousCondition := QualifiedCondition;

A new occurrence wins over an acknowledgement in the same scan. That is a deliberate policy: a coincident button press must not quietly acknowledge an event that just arrived. Reset only works after the condition has cleared and acknowledgement has occurred. Different plants may select a different documented policy, but hidden scan-order accidents should never choose it.

If the condition clears and returns while the alarm remains latched, it becomes unacknowledged again. That is useful when a recurring problem needs renewed attention. Record occurrences separately from the latch if frequency matters; one latched bit cannot tell you how often the pressure failed.

First-out captures evidence, not certainty

A first-out record preserves the earliest detected initiating condition for a trip group. Suppose a drive fault is detected, then conveyor speed falls, then a transfer times out. Keeping the drive fault helps diagnosis after all three conditions are visible.

Several conditions can arrive in the same PLC scan. In that case, program order or an explicit priority list may decide which is recorded first. Label that limitation honestly. A scan-based first-out is not a high-resolution proof of physical causality. For faster events, device timestamps or dedicated event recording may be needed, with synchronised clocks and known uncertainty.

Store the relevant context too: operating state, batch identifier, command, key feedback and timestamp source. “Fault 83 happened” is weaker evidence than “transfer timed out while sending carton 418; upstream sensor stayed occupied and downstream acceptance was true.”

Reduce floods at the source

If the master air supply fails, twenty cylinders may all report missing position. Decide which downstream alarms remain useful and which should be suppressed by a documented dependency. Suppression must be visible and controlled; deleting messages simply because they are inconvenient can hide an independent fault.

State-dependent alarming helps distinguish normal absence from abnormal absence. Low flow may be expected while a pump is stopped and abnormal after a running pump's startup allowance. Changes to alarm behaviour should be reviewed with operations because they change what attention is requested. ISA's discussion of alarm management includes state-dependent treatment for batch and discrete processes. Read the alarm-management context.

Priority should reflect consequence and response time, not how annoying a message is to the programmer. A frequently recurring minor alarm does not become urgent through repetition. Its frequency is evidence that the design or equipment needs attention.

Try it

A fill station reports “Target weight not reached” immediately when filling begins. The programmer adds a 30-second delay. The normal fill takes eight seconds, but a disconnected weighing instrument holds its last value at zero. Explain the two separate problems and propose better triggers.

Work through the answer

Target not reached is normal while filling is in progress. It becomes an abnormal timeout only if the operation exceeds its validated allowance. Tie that monitor to the filling state and define what completion evidence stops it.

The disconnected instrument is a measurement-validity failure. Detect it through channel diagnostics and freshness rules, then follow the process's invalid-measurement response. Do not wait for a generic fill timeout before admitting that the stop measurement is unavailable. The alarm text should identify the failed weighing signal and the appropriate recovery procedure.

Finally test both faults independently: a healthy scale with no material flow, and an invalid scale while material is available. They can produce the same zero value but require different evidence and potentially different responses. Good alarms preserve that distinction.

Now make the decision yourself

Use the chapter’s model on a fresh question, then compare your reasoning with the worked decision.

How this becomes a program

The cause disappeared before anyone saw it. What evidence remains?

Pressure drops for one scan during filling. The alarm must preserve the event, the phase and the relevant reading for diagnosis.

First causeLatched eventRecovery conditionstime →

The low-pressure condition disappears, but its record stays latched. Conditions may later permit recovery; they do not clear the record or command motion by themselves.

Your first artifact

Decide what gets captured on the first occurrence and which action clears that record. Distinguish ‘seen by operator’ from ‘cause gone’.

Open the worked decision
IF PressureLow AND NOT FaultLatched THEN
  FaultState := State;
  FaultPressure := Pressure_bar;
  FaultLatched := TRUE;
END_IF;

Why this line belongs here

The first-cause guard prevents later symptoms from overwriting the useful evidence. Acknowledge can change notification state without erasing the record or requesting motion.

Change the task

After a pressure fault, arrival also times out. Decide whether to preserve one cause, record both events, or show a causal group; state your retention rule.