Chapter 29 / 36

Maintain and secure

Plan changes, access, backups and restoration for the life of the machine.

A program is only useful if the next person can recover it

The machine runs. Its original laptop is gone. A replacement controller has arrived, but the available project requires a library version nobody recorded. The drive backup predates a gearbox change, and the HMI contains recipes missing from the server copy.

This is a controls failure even though the sequence logic is excellent. Maintainability includes the ability to identify, understand, change, test, and restore the complete system. Security protects those abilities as well as normal operation.

In this chapter, build a maintenance package for a small inspection cell. The model has a PLC, HMI, camera interface, motor drive, and engineering workstation. Those are five configuration owners, not one project file.

Inventory the things that can change behavior

Start with a configuration inventory. Include assets whose parameters influence the process even when they contain no code you wrote.

AssetControlled materialRestore dependency
PLCSource, compiled deployment where applicable, hardware configurationFirmware, libraries, engineering tool
HMIApplication, users and role configuration, recipe definitionsRuntime version and licensed features
DriveMotor data, limits, control mode, network settingsDrive type and firmware
CameraJob, calibration references, result mappingCamera software and optical setup
WorkstationApproved tools and configurationInstallation media, access, compatibility

Separate secrets from ordinary project files. Credentials, private keys, and recovery codes need approved protected storage and access processes. Committing them beside Structured Text does not make the system easier to maintain; it exposes the system and makes rotation harder.

Record network addresses and device identity without treating an address as proof of identity. Replacement equipment may inherit an address while carrying different firmware or parameters. Recovery verification needs both identification and behavioral checks.

Version the decision, not just the file

A useful change record connects a reason to a tested result. For example: “Reject confirmation timeout increased after validated actuator replacement; process response budget reviewed; tests RC-04 through RC-08 rerun.” That is more helpful than a file named final_really_final_3.

Use the organization's version-control approach for source and configuration exports. Keep vendor binary files when required, but also keep readable exports or comparison reports where the tool supports them. A reviewer should be able to see whether a change affects a timer preset, a state transition, or an entire I/O mapping.

Tag or otherwise identify released packages. Distinguish a working copy from the version deployed to each machine. “Latest in the repository” and “running in production” are separate facts until verified.

A backup is a hypothesis until restored

Copying a project proves that a copy was made. It does not prove the copy can restore the intended service. Test restoration in an appropriate isolated or approved environment, including dependencies and parameters. Record the result and expected recovery sequence.

For our model cell, simulate loss of the PLC configuration. Restore the controlled package, then verify I/O mapping, communication identity, recipe limits, restart behavior, and product tracking initialization. A successful download followed by the wrong reject timing is not a successful recovery.

CISA advises maintaining and testing recovery capabilities and software backups for operational technology. Its guidance also emphasizes reducing unnecessary exposure and controlling remote access. These are practical reasons to make recovery an exercised process rather than a folder nobody opens. CISA primary OT mitigations.

Do not store the only backup on the same workstation that controls every device. Use the organization's protected backup strategy, including copies resilient to accidental deletion or compromise. Access to backups and the ability to deploy them are themselves powerful privileges.

Give each person the access their task needs

An operator changing an approved recipe, a technician viewing diagnostics, and an engineer deploying software have different responsibilities. Design accounts and roles accordingly. Shared administrator credentials erase accountability and make it difficult to revoke access for one person.

Remote access deserves explicit ownership, approval, authentication, and logging. Avoid exposing controllers directly to the public internet. Use the site's approved architecture and coordinate with OT security personnel. Network segmentation, least privilege, and maintained recovery plans are recurring themes in official CISA guidance. CISA StopRansomware guide.

Security settings can affect availability and maintenance. Plan certificate renewal, account recovery, and access during an incident. A locked-down system with undocumented recovery credentials may become impossible for authorized personnel to restore when time matters most.

Treat an OT change as an engineering change

A patch, firewall rule, firmware update, or replacement switch can alter timing or communication behavior. Review compatibility and impact before production deployment. Test in a representative environment where possible, define a rollback path, obtain the required approvals, and schedule the work with operations.

Do not interpret this as a reason never to update. Permanently postponing maintenance creates its own risk. The task is to manage changes with evidence and ownership, taking both security and process consequences into account.

For the inspection cell, a camera firmware update changes a result field from a one-cycle pulse to a latched value. A superficial connection check passes. The product counter then increments repeatedly because the PLC increments on every true evaluation, assuming the field lasts only one invocation. A rising-edge detector would count that sustained high once; it could instead miss later results if the signal never returned low. The change review should have identified the interface contract and rerun the transaction tests.

Preserve what happened during an incident

When something unusual occurs, follow the site's incident process. Keep relevant logs, times, versions, and observations. Coordinate with the responsible operations and security teams before making changes that destroy evidence or alter equipment state. A communication outage can have ordinary technical causes, malicious causes, or both; the symptom alone does not settle the question.

In the training model, practice writing a factual timeline: first alarm, affected devices, observed commands, last known configuration change, and recovery actions. Avoid assigning blame in the log. Accurate records help both technical diagnosis and organizational learning.

Try it

The only engineer who knows a packaging cell is leaving. You have one afternoon to reduce the handover risk. Choose five concrete deliverables. Explain why a source-code zip and a password in an email are insufficient.

Work through the answer

First, identify the actual deployed versions across PLC, HMI, drives, and peripheral devices. Second, create the controlled recovery package with dependencies and readable configuration exports. Third, demonstrate restoration or document the latest witnessed restore test and remaining gaps. Fourth, transfer access through approved account and secret-management processes. Fifth, provide the operating boundaries, known issues, acceptance tests, and support contacts.

A source zip omits device parameters, recipes, firmware dependencies, and possibly the actual production revision. A password in email neither establishes appropriate roles nor provides controlled long-term access and revocation. The aim is continuity of capability, not possession of a few files.

Add a small change exercise: ask the receiving engineer to locate the reject timeout requirement, identify its implementation, and run its test without coaching. That reveals whether the package communicates intent. With maintainability established, measure and improve shows how to change performance without sacrificing that clarity.

Now make the decision yourself

Use the chapter’s model on a fresh question, then compare your reasoning with the worked decision.

How this becomes a program

Can you restore the working machine, not just its source file?

A replacement controller must match firmware, I/O, drive parameters, recipes and network interfaces as well as code.

PLCPLCVersioned baselineControlled changeVerified restorerequest → evidence → acknowledgement

Your first artifact

Inventory the whole deployable configuration. Define authorized access and a change record. Rehearse a restore using the actual supported tools and approved conditions.

Open the worked decision
(* Restore checklist is an engineering artifact, *)
(* not a PLC instruction or a shared password. *)

Why this line belongs here

A backup is useful only if it can be restored and identified. Network credentials, controller settings and retained process data have different protection and recovery needs.

Change the task

A backup contains last year's recipe limits. Write how a restore detects that mismatch before production starts.