PI System Data Quality: How to Find, Diagnose, and Prevent Bad PI Data
A PI Tag can be online and still provide bad data. The value may be stale, flatlined, delayed, incorrectly scaled, mapped to the wrong source, or behaving in a way that no longer represents the physical process.
These problems can spread because PI data often feeds other systems and workflows, including:
PI Vision displays
PI Analyses
Reliability workflows
Reports and dashboards
Power BI
Machine learning
AI applications
Bad data rarely stays isolated to one sensor or tag. Once other applications depend on it, the impact can spread across calculations, displays, reports, analytics, and AI systems.
PI System data quality is the degree to which operational data in PI is complete, available, timely, valid, and representative of the physical process it describes.
A practical PI System data quality program should evaluate seven areas:
Availability
Completeness
Timeliness
Validity
Behavior
Compression
Context
The objective is not to make every PI Tag perfect. Teams need to identify the data that matters, understand how it should behave, and detect problems before they affect operations or downstream applications.
1. Check data availability
Start with a basic question: Is the PI Tag providing data when expected? A tag can stop updating because of problems anywhere between the sensor and the PI System.
Common causes include:
Sensor failure
Communication loss
Interface failure
Connector problems
Source system problems
Network issues
Configuration changes
The expected update rate matters when deciding whether data is stale. A pressure sensor that normally changes every second should not be evaluated the same way as a laboratory value that updates once per day.
Practical checks
For important tags:
Define the expected update rate.
Detect stale values.
Detect tags that stop updating.
Compare actual update behavior with expected behavior.
Review interface and connector health.
Identify repeated communication failures.
Do not use the same stale-data threshold for every tag. The threshold should reflect how the measurement normally behaves.
2. Identify missing data and gaps
A tag can still be updating today while parts of its history are missing. Those gaps can cause problems when the data is later used for trending, calculations, reliability analysis, machine learning, AI models, or regulatory reporting.
Historical gaps can also be difficult to notice when users only look at the latest value. A tag may appear healthy even though an important operating period is missing from the archive.
Practical checks
Review:
Missing historical periods.
Unexpected gaps between events.
Archive gaps.
Repeated communication interruptions.
Missing periods during important operating events.
Also determine whether missing periods are actually unexpected. A value that updates only when a process is running will behave differently from a continuously sampled sensor.
3. Check whether the data is timely
Availability and timeliness measure different problems. A value can arrive successfully and still arrive too late to be useful.
Delays can be caused by:
Network latency
Interface buffering
Store-and-forward behavior
Source system delays
Timestamp problems
Incorrect clocks
A delayed value may look normal once it reaches PI. The problem appears when a real-time calculation, dashboard, analytics application, or AI system assumes the value is current.
Practical checks
Check:
Difference between event time and arrival time.
Unexpected delays.
Clock synchronization.
Timestamp consistency.
Repeated latency patterns.
Acceptable delay depends on the use case. A five-minute delay may have little effect on a monthly report but can be unacceptable for real-time monitoring.
4. Check whether values are valid
Some PI data problems can be detected with straightforward rules. A value may fall outside a physical limit, use the wrong engineering units, contain an invalid digital state, or have incorrect scaling.
Examples include:
Negative flow where negative flow is impossible.
Temperature outside the physical range of the instrument.
Invalid digital states.
Impossible pressure values.
Incorrect engineering units.
Incorrect scaling.
Basic validity checks are useful because they find obvious errors quickly. They are also relatively easy to implement and explain.
Practical checks
Define:
Expected minimum and maximum values.
Engineering limits.
Allowed states.
Valid units.
Expected data types.
Reasonable rate-of-change limits.
Static thresholds cannot detect every problem. A value can stay inside its expected range and still be wrong.
5. Understand whether the signal is behaving normally
Many PI data quality problems show up as changes in behavior rather than obvious invalid values. The individual values can look reasonable while the signal starts behaving differently from its historical pattern.
Examples include:
A pressure signal suddenly becomes flat.
A temperature signal becomes much noisier.
A level signal stops following related process conditions.
A tag begins changing much more slowly than normal.
A signal develops repeated spikes.
A calculated value changes pattern after an upstream configuration change.
Historical behavior provides the baseline for deciding whether a change is unusual. Without that baseline, a system may either miss real problems or generate excessive false positives.
Practical checks
Compare current behavior with:
Historical ranges.
Normal variability.
Typical rate of change.
Expected update frequency.
Related measurements.
Similar equipment.
Behavioral monitoring can detect problems that simple range checks miss. It is especially useful for measurements that remain within valid limits even when something has changed.
6. Evaluate tag compression
A PI value can pass availability, completeness, and validity checks while still representing the physical process incorrectly. Data fidelity looks at that relationship between the digital value and the real-world condition it is intended to measure.
Compression describes how well a digital value represents the physical process or condition that it is intended to measure.
Consider a temperature sensor where the PI Tag:
Updates every second.
Has no missing data.
Stays within its expected range.
Has a valid engineering unit.
Those checks would make the tag appear healthy. If the sensor has drifted by 15 degrees, however, the digital value no longer represents the actual process temperature accurately.
Other examples include:
Incorrect PLC scaling.
A sensor connected to the wrong source.
A calculation using the wrong input.
Sensor drift.
Incorrect tag mapping.
Timestamp shifts.
An instrument frozen at a plausible value.
These problems are difficult to detect because the values can look normal. Traditional availability and range checks may never generate an alert.
Practical checks
Where possible:
Compare redundant measurements.
Compare related process variables.
Review historical behavior.
Check scaling and engineering units.
Confirm source mappings.
Compare digital values with known operating conditions.
Investigate unexpected differences between related measurements.
Data fidelity often requires process knowledge. Software can identify unusual relationships or behavior, but an engineer may still need to determine whether the measurement is physically reasonable.
7. Add context before deciding that data is bad
Context determines what normal behavior looks like for a measurement. A flat signal, slow update rate, or unusual value pattern may be completely normal for one tag and a serious problem for another.
Consider a few common cases:
A tank level may remain unchanged for several hours because the tank is stable.
A laboratory value may update only once per day.
A valve position may remain at zero for weeks because the valve is normally closed.
A rule that labels every flat signal as bad would generate large numbers of false positives. Data quality rules need to reflect the type of measurement and the operating conditions around it.
Useful context includes:
Measurement type.
Equipment type.
Expected update rate.
Normal operating range.
Historical variability.
Operating state.
Related measurements.
Asset relationships.
The PI Asset Framework can provide part of this context. AF can help identify what a tag represents, which equipment it belongs to, and how similar assets are structured.
Why flatline detection is harder than it looks
Flatline detection is a common PI System data quality check, but simple implementations tend to generate noise. A basic rule such as "alert if the value does not change for 30 minutes" works for some measurements and fails badly for others.
A level measurement may legitimately remain constant, while a laboratory value may change only once per day. A discrete state can stay unchanged for days, and a process sensor can make very small changes that disappear because of compression or rounding.
A better flatline check should consider:
Historical frequency of change.
Typical variability.
Measurement type.
Operating state.
Data collection frequency.
PI compression behavior.
Similar historical periods.
Detecting bad PI data therefore requires a model of expected signal behavior. Static thresholds can still be useful, but they should be applied with enough context to distinguish unusual behavior from normal operation.
What causes bad data in the PI System?
Bad PI data can originate at many points along the data path. Troubleshooting only the PI Tag can miss the real source of the problem.
Common causes include:
Sensor failure.
Instrument drift.
PLC or DCS configuration changes.
Incorrect scaling.
Network communication problems.
Interface or connector failures.
Incorrect PI Tag configuration.
Incorrect source mappings.
Exception configuration.
Compression configuration.
Failed calculations.
Incorrect engineering units.
Timestamp errors.
Source system changes.
Manual configuration errors.
Teams often need to understand where the data came from and what changed upstream. That broader view can reduce the amount of manual investigation required to find the cause.
Data quality and lineage should work together
A data quality alert identifies something that may be wrong. Lineage provides the dependency information needed to investigate where the value came from and what else depends on it.
Consider a calculated PI value that suddenly starts behaving differently. Without lineage, an engineer may need to manually determine:
Which PI Tags feed the calculation.
Which AF Attributes are involved.
Whether another PI Analysis provides an input.
Whether an upstream resource changed.
With lineage, that dependency chain is already available for investigation. The same information can also show the downstream impact of the affected data.
A single bad PI Tag may affect:
Several calculations.
Multiple AF Attributes.
PI Vision displays.
Power BI reports.
Reliability models.
AI applications.
The severity of a data quality issue depends partly on where the data is used. A questionable tag feeding a critical calculation or operational display deserves more attention than an unused tag with the same underlying problem.
Do not monitor every PI Tag the same way
Large PI environments can contain hundreds of thousands or even millions of tags. Applying the same checks and alerting thresholds to every tag creates noise and makes the important problems harder to see.
Start with data that has meaningful operational or downstream impact:
Critical process measurements.
Data used by important PI Analyses.
Data used by control room and operational PI Vision displays.
Data used in reliability workflows.
Data used by reports and analytics.
Data used by machine learning and AI applications.
Lineage and usage information can then help determine which tags have the largest downstream footprint. This gives teams a practical way to focus monitoring where failures are most likely to matter.
PI System Data Quality Checklist
Use this checklist when reviewing important PI data. The checks should be adjusted based on the measurement, equipment, operating state, and expected behavior.
Availability
Are critical tags updating?
Are values becoming stale?
Are interfaces and connectors healthy?
Is the actual update rate consistent with expectations?
Completeness
Are there missing historical periods?
Are there unexpected gaps?
Are important operating periods missing data?
Timeliness
Is data arriving when expected?
Are timestamps correct?
Are there unexpected delays?
Validity
Are values within reasonable ranges?
Are engineering units correct?
Are states valid?
Are scaling and data types correct?
Behavior
Are there unexpected flatlines?
Are there unexpected spikes?
Has normal variability changed?
Has the rate of change changed?
Does current behavior differ significantly from historical behavior?
Data fidelity
Does the digital value represent the physical process?
Could the instrument have drifted?
Is the source mapping correct?
Is scaling correct?
Do related measurements support the value?
Context
Do you know what the measurement represents?
Do you know its expected update rate?
Do you know the equipment it belongs to?
Do you know its normal operating behavior?
Is the tag mapped to an AF Attribute where appropriate?
Impact
Which calculations depend on the tag?
Which PI Vision displays use it?
Which reports depend on it?
Which analytics or AI applications use it?
What does mature PI System data quality look like?
Organizations do not need to solve every data problem at once. A useful maturity path is to progressively understand the data, establish normal behavior, monitor it, and improve the rules based on what the team learns.
A practical sequence is:
Identify → Baseline → Monitor → Trace → Prioritize → Improve continuously
Identify
Determine which PI data supports important operations and applications. Start with the measurements where bad data could affect decisions, operations, reliability, analytics, or AI.
Baseline
Understand normal update frequency, ranges, variability, and behavior. Historical data provides the reference needed to tell the difference between unusual behavior and normal operating conditions.
Monitor
Continuously check for stale data, missing periods, abnormal behavior, and other quality problems. Monitoring should use thresholds and models appropriate for the measurement.
Trace
Understand where questionable data came from and which downstream systems use it. Lineage can reduce investigation time by making those relationships visible before an incident occurs.
Prioritize
Focus investigation on the problems with the highest operational impact. Usage and lineage information can help distinguish a noisy alert from an issue affecting critical workflows.
Improve continuously
Use historical incidents, configuration changes, and resolved false positives to improve the monitoring rules. Data quality monitoring becomes more useful as the system learns what normal and abnormal behavior look like in the environment.
Why PI System data quality matters for AI
AI increases the consequences of poor operational data quality because it can consume incorrect information automatically and at scale. An experienced engineer may recognize that a process value does not make sense, while an AI application may accept the same value unless it has enough context to identify the problem.
AI applications can also depend on hundreds or thousands of measurements. Manually validating every input quickly becomes impractical as those applications expand.
Before PI data is used by AI, teams should understand:
Whether the data is available.
Whether it is complete.
Whether it arrives on time.
Whether the values are valid.
Whether the signals behave normally.
Whether the values faithfully represent the physical process.
Where the data comes from.
What has changed.
AI creates a greater need for reliable operational data because more decisions and workflows can depend on the information automatically. The quality of the AI output still depends on the quality and context of the data it receives.
How Tycho Data supports PI System data quality
Tycho Data helps industrial teams identify and investigate data quality problems across the PI environment. Osprey combines information about data behavior, dependencies, configuration changes, and downstream usage so teams can investigate issues with more context.
Osprey provides visibility into areas such as:
Stale data.
Missing data.
Abnormal signal behavior.
Data quality incidents.
PI dependencies.
Lineage.
Configuration changes.
Downstream usage.
Instead of stopping at "something looks wrong," teams can investigate what is wrong, where it came from, what changed, and what the affected data feeds. This helps reduce the manual work required to understand meaningful operational data problems and resolve them.
The objective is also to avoid creating another noisy alerting system. Data quality monitoring is most useful when it helps teams focus on the issues that have real operational or downstream impact.
Frequently Asked Questions
What is PI System data quality?
PI System data quality describes whether operational data in PI is complete, available, timely, valid, and representative of the physical process it describes. Good data quality also depends on understanding the expected behavior and context of each measurement.
How do you check data quality in the PI System?
Common checks include stale-data detection, missing-data detection, flatline detection, range checks, rate-of-change checks, timestamp checks, calculation monitoring, and comparison with historical behavior. The specific checks and thresholds should reflect how each measurement is expected to behave.
How do you detect stale PI Tags?
A PI Tag can be considered stale when it has not received a new value within its expected update period. The correct threshold depends on the measurement because a tag expected to update every second should not use the same threshold as a laboratory value that updates once per day.
How do you detect flatlined PI Tags?
Flatline detection identifies measurements that continue to report the same or nearly the same value for an unusual period. A useful flatline rule should account for historical variability, expected update frequency, measurement type, operating state, and PI compression behavior.
What causes bad data in the PI System?
Common causes include sensor problems, instrument drift, communication failures, incorrect scaling, PI configuration issues, source mapping errors, failed calculations, timestamp problems, and changes in PLC or DCS configuration. The source of the problem may exist upstream of PI, so diagnosis often requires looking at the full data path.
What is the difference between stale data and flatlined data?
Stale data means that a PI Tag has stopped receiving new events when new events are expected. Flatlined data means new events may still arrive, but the reported value remains unchanged or nearly unchanged for an unusual amount of time.
What is data fidelity?
Data fidelity describes how accurately a digital value represents the real physical process or condition it is intended to measure. A PI Tag can pass basic availability and validity checks while still having poor fidelity because of sensor drift, incorrect scaling, mapping problems, or another source issue.
Why is context important for PI data quality?
Context helps determine what normal behavior should look like for a measurement. A laboratory value, tank level, valve state, and pressure measurement can have very different update frequencies and expected behavior, so applying the same rules to all four can generate false positives.
How does PI System data quality affect AI?
AI applications depend on the operational data provided to them, and stale, missing, incorrect, or misleading PI data can lead to poor results. AI readiness therefore requires monitoring technical data quality while also checking whether the data represents the physical process accurately.
Do you need to monitor every PI Tag?
No. Start with data used by critical operations, calculations, displays, reports, analytics, and AI applications, then prioritize monitoring based on usage and potential business impact.
Is PI System data quality a one-time cleanup project?
No. Sensors fail, configurations change, interfaces change, calculations evolve, and new applications begin using existing data, so data quality requires ongoing monitoring rather than a one-time cleanup.