Chapter 28 / 36

Debug with a hypothesis

Use traces and controlled experiments to find causes rather than guess at fixes.

Debugging chain from physical result through the output boundary, final command, permissions, remembered state and request acceptance
Follow evidence backward from the symptom. Do not jump from a stopped belt to rewriting the sequence. First find the boundary where observed behavior differs from the intended contract.

If Q_Belt is FALSE, inspect the terms that own that decision. If Q_Belt is TRUE but the physical drive input is absent, move outward toward mapping, communication and electrical evidence. If the drive runs but the carton stays still, inspect the mechanical process and arrival supervision. Each observation narrows the search to a different layer. A useful hypothesis names the expected difference before you change anything.

Stop asking which line looks wrong

The model transfer conveyor stops halfway through a tray movement. You notice a timer preset, a sensor filter, and a sequence transition. All three look suspicious. Changing all three may make the animation finish, but you will not know what failed or whether it can fail again.

Debugging begins with a narrower question: what is the first observable difference between intended and actual behavior? If the motor request disappears at 1.24 seconds, investigate why it disappeared. Do not start by increasing a four-second arrival timeout. That timeout has not yet expired.

A professional does not need to guess correctly on the first attempt. They need a method that makes each attempt informative.

Preserve the failure before improving it

Record the software revision, parameters, initial state, input sequence, and observed result. Save the smallest trace that contains the event and enough preceding history to explain it. If the failure is intermittent, record occurrence frequency and what changed between successful and failed cycles.

In the browser model, reset to a known initial condition and replay the same disturbances. On an actual machine, reproducing a fault may be unsafe or disruptive; investigation must follow the site's authorized conditions. Sometimes the correct next step is analysis of existing records rather than another physical trial.

Write the symptom without a diagnosis: “Motor request became false while Transfer was active and no arrival had been accepted.” Avoid “The timer is broken.” The first sentence preserves possibilities. The second encourages you to ignore contrary evidence.

Follow the chain backward

A useful cause chain for a stopped conveyor is:

Physical conveyor stopped
  ← drive stopped producing motion
  ← drive command or permission changed
  ← final motor request changed
  ← sequence request or production permission changed
  ← sampled signal, state, parameter, or diagnostic changed

Start at the earliest level you can observe reliably. A green HMI icon is not proof of a drive command if the display updates slowly or uses a different tag. Compare the exact variables and timestamps.

Suppose the trace shows TransferState = Moving, SequenceRunRequest = TRUE, and MotorRun = FALSE. That directs attention to output arbitration or permissions, not the transition out of Moving. If MotorRun = TRUE throughout the stop, ordinary sequencing may be innocent; the command path, drive status, feedback, or plant model needs investigation.

Build a hypothesis table

For the model fault, collect this trace:

TimeMovingSequence requestMotor requestDownstream ready
1.22 sTrueTrueTrueTrue
1.23 sTrueTrueTrueTrue
1.24 sTrueTrueFalseFalse
1.25 sTrueTrueTrueTrue

Now state competing explanations and predictions:

HypothesisPredictionDistinguishing observation
Arrival timeout expiredTimer Q becomes trueQ remains false, so reject
Downstream readiness gates motion continuouslyMotor follows a ready glitchMatches trace; inspect interface contract
Another routine overwrites MotorRunLast writer changes valueCross-reference and task trace

Do not stop at correlation. Downstream readiness and motor request may share a third cause. Inspect the actual assignment and task ownership to test causation.

Find the design mistake behind the symptom

The code contains:

MotorRun := SequenceRunRequest AND DownstreamReady;

This may be correct for one interface and wrong for another. In our model, downstream readiness means “I can accept a new tray.” After accepting the transfer, the downstream station withdraws readiness to prevent a second reservation. The local conveyor mistakenly interprets that withdrawal as “stop the accepted transfer.”

Adding a sensor filter would hide the issue only while the withdrawal remained short. Increasing the timeout would allow a longer wrong behavior. The root correction is an interface contract: acceptance reserves the destination for this transaction, and motion follows that accepted transaction until completion or an explicit abort condition.

Update the requirement and interface diagram before editing the expression. A debug session often discovers an incomplete design, not a misspelled variable.

One change, one prediction

Change the model to latch an accepted transfer under a transaction identifier. Do not simultaneously adjust speed, timeout, and sensor filtering. Predict the result: readiness withdrawal after acceptance will not interrupt the accepted movement; readiness false before acceptance will still block it.

Run both tests. Then run explicit abort and communication-loss tests. Removing a continuously evaluated ready condition must not accidentally ignore a different stop requirement. The corrected contract should distinguish readiness, acceptance, completion, and loss of permission.

When a result contradicts your prediction, revise the hypothesis. Do not keep layering conditions until the visible symptom disappears. That produces programs nobody can explain and regressions nobody can confidently isolate.

Timing deserves its own evidence

Many PLC faults involve events that are individually correct but ordered badly. A completion bit may be raised and cleared between HMI polls. A pulse may be shorter than the receiving task interval. Two tasks may read different generations of shared data. A timer may be conditionally skipped, retaining internal state longer than expected.

Log event identifiers or counters when a short pulse is hard to observe. Record both the sender's transition and the receiver's acknowledgement. Use the controller's appropriate tracing tools on hardware, with attention to acquisition rate, task loading, and timestamp meaning. Increasing trace frequency does not automatically create perfect information.

For a browser exercise, deliberately vary the communication delay independently of the control step. You should be able to explain which contract still holds when the screens update slowly. A robust handshake does not require two displays to blink together.

Leave a useful repair record

A repair note should explain symptom, cause, correction, affected requirements, and verification. “Fixed conveyor bug” is almost useless six months later. “Reserved-transfer state replaces continuous Ready gating; verified readiness withdrawal after acceptance, pre-acceptance blocking, abort, and stale-communication handling” preserves the reasoning.

If the diagnosis remains uncertain, say so. A temporary containment can be appropriate under the project's authority, but label it and define the evidence needed for a permanent correction. Do not promote a lucky parameter change into a proven root cause.

Try it

An actuator sometimes reports a travel timeout immediately after reset. The timer is called only while State = Moving. Reset changes the state to Idle, and the next start returns it to Moving. Propose a hypothesis, a decisive trace, and a correction to test. Assume the selected timer retains its instance state until called with a resetting input.

Work through the answer

The hypothesis is that leaving Moving stops execution of the timer without resetting its internal state. On the next entry, its old completion state is still present. Trace the timer's invocation, IN, Q, elapsed value, and sequence state across fault, reset, and restart.

Call the timer consistently according to the target library's contract, with IN derived from whether movement is currently being supervised. Ensure it receives false during the reset condition. Then test a full timeout, reset, fresh start, normal completion, and repeated short attempts. Do not merely assign a new preset or clear a display alarm.

Save that regression test with the repair. It becomes part of the maintenance evidence discussed in maintain and secure.

Now make the decision yourself

Use the chapter’s model on a fresh question, then compare your reasoning with the worked decision.

How this becomes a program

What observation would prove your guess wrong?

The belt sometimes stops early. A sensor blink, a readiness loss and a timer error could produce similar symptoms.

HypothesisTimestamped evidenceOne controlled change

Your first artifact

Capture state, input image, output request and the first reason for transition. Write a hypothesis that predicts a specific ordering in that trace.

Open the worked decision
(* Capture on MOVING → INTERRUPTED *)
(* State, AtDrop, DriveReady, TravelWatch.ET *)

Why this line belongs here

Changing three filters and two timers at once destroys your ability to attribute the result. The smallest useful trace preserves the information needed to distinguish the competing explanations.

Change the task

Your hypothesis is ‘the arrival sensor bounces’. Which transition and input history would contradict it?