Guide

The ultimate guide to managing your PI System Download now

Incident Management in the PI System: Best Practices

Incident Management in the PI System: Best Practices

A PI data incident can affect dashboards, calculations, reports, and operational decisions. A structured response helps the team restore reliable data without creating additional risk.

Use a simple incident lifecycle: detect, assess, contain, diagnose, recover, verify, and learn.


Define roles before an incident

Do not decide ownership during a high-pressure event.

Common roles include:

  • Incident lead: coordinates work and decisions.

  • PI or data-system specialist: investigates the technical data path.

  • Process or operations specialist: determines operational meaning and priority.

  • Scribe: records timeline, actions, and decisions.

  • Communications owner: provides updates to affected stakeholders when needed.

One person can hold more than one role in a small team.


Assess severity by operational effect

Do not copy an IT severity model without adapting it to operations.

Consider:

  • Safety relevance

  • Production impact

  • Number of affected assets or sites

  • Effect on operator displays

  • Effect on critical calculations or reports

  • Data loss or recoverability

  • Regulatory or contractual exposure

A data gap on an unused test tag is not the same as a gap in a critical production or safety-related signal.


Contain before you optimize

The first objective is to stop the incident from causing more bad data or bad decisions.

Containment can include:

  • Disable or isolate a failed calculation.

  • Restore the previous known configuration.

  • Mark a dashboard or output as unreliable.

  • Stop a faulty writer.

  • Move an interface to a validated recovery path.

Do not backfill or overwrite history until the team understands the source and required audit controls.


Diagnose the data path

Trace the signal in order:

  1. Instrument or source application

  2. PLC or DCS when applicable

  3. Interface or connector

  4. PI Data Archive

  5. AF attribute or analysis

  6. PI Vision or downstream consumer

Review recent configuration changes and timestamps. This often narrows the cause quickly.


Recover data carefully

Recovery can include restarting a failed service, correcting a mapping, reprocessing an analysis, or backfilling missing history.

Verify that the source data is authoritative before backfill. Record the method and time range of any historical correction that has business or regulatory importance.


Verify downstream systems

A successful source repair does not prove that all consumers recovered.

Confirm important:

  • AF analyses

  • PI Vision displays

  • Reports

  • Alerts

  • External data pipelines

Check both current values and affected historical periods.


Record the incident timeline

Capture:

  • Detection time

  • Start time if known

  • Affected systems and data

  • Important changes before the event

  • Actions taken

  • Recovery time

  • Verification steps

  • Follow-up actions

A useful record helps the next investigation and supports governance.


Run a no-blame review

The review should focus on system controls, not individual fault.

Ask:

  • Why was the issue not detected earlier?

  • Which dependency was difficult to trace?

  • Was ownership clear?

  • Could a monitoring rule have detected the failure?

  • Could change review have prevented it?

The best incident process reduces future detection and diagnosis time. It also turns each event into better monitoring, documentation, and change control.