Webinar

Can you Trust your Process Data in the age of AI? Register Now

Can you Trust your Process Data in the age of AI? Register Now

PI Tag Compression: How It Affects Data Quality, Analytics, and AI

PI compression is easy to treat as a storage setting. For years, that was often how teams thought about it: configure the tags, reduce the number of archived events, and move on.

Historical PI data is now used for much more than trending. It feeds reliability programs, process investigations, analytics, machine learning, and AI. That makes compression worth another look because it affects what the historical record actually contains.

PI Tag compression reduces the number of events stored in the PI Data Archive while preserving the useful shape of a process signal. When the settings are appropriate, PI stores enough information to represent what happened without keeping every sample. When the settings are too aggressive, useful process behavior may be missing from the historical record, sometimes years later when someone needs it.

Why does the PI System use compression?

Industrial sensors generate a lot of data. A temperature measurement collected every second produces 60 values per minute, 3,600 per hour, 86,400 per day, and more than 31 million per year.

Many of those values add very little information. If a temperature moves gradually from 100°F to 101°F, storing every small step between those values may not improve the historical representation of the process. PI compression reduces that volume while preserving the shape of the signal.

AVEVA describes compression as evaluating values through the PI Snapshot subsystem and determining whether their deviation from the expected trajectory requires another value to be archived. In practical terms, the goal is to keep enough information to represent the process at the resolution users need.

Exception and compression do different jobs

Exception and compression are often discussed together, but they happen at different points. With traditional PI Interfaces, exception reporting can reduce the values sent toward the PI Server. Compression happens later and determines which eligible values are written to the PI Data Archive.

A simplified path looks like this:

Sensor → PLC/DCS → PI Interface → Exception Algo → PI Snapshot → Compression Algo → PI Data Archive

The exact path depends on how data enters PI, and newer collection technologies may handle exception behavior differently. The useful distinction is that exception affects what gets sent toward PI, while compression affects what ends up in the historical archive.

What does CompDev do?

CompDev is the compression deviation. It is sometimes interpreted as a simple rule such as "store a value whenever the signal changes by X," but PI compression does more than compare the latest value with the previous one.

The algorithm considers whether values deviate enough from the expected trajectory of the signal to require another archived event. This allows PI to preserve the shape of a changing signal without storing every sample.

The CompDev setting also needs to make sense for the process. Consider a temperature tag with a configured span from 0 to 1,600°F. The process may normally operate within a much narrower range around 1,000°F. A compression setting that looks small compared with the full instrument span can still remove changes that operators care about during normal operation.

The instrument range and the meaningful process range are not always the same.

What does CompMax do?

A stable process can remain within the compression criteria for a long time. CompMax limits how long PI can go without archiving another event under those conditions, which leaves periodic evidence of the tag's state even when the signal changes very little.

CompMax works with CompDev. The two settings solve different parts of the compression problem.

Can PI compression remove important data?

Yes. Compression intentionally means that PI does not archive every value it receives, and that is expected behavior. Problems occur when the values that are not preserved contain process information that someone later needs.

Suppose a pressure signal usually changes slowly, but a short excursion can indicate an equipment problem. If the collection and compression strategy does not preserve enough detail around that event, the archived history can show a simplified version of what happened. The raw samples that were never archived cannot be reconstructed later from the PI Data Archive.

This makes compression relevant to data fidelity, which describes how well digital data represents the physical process. Compression affects how much process detail remains available in the historical record.

What happens when compression is too aggressive?

Aggressive compression reduces historical detail. Depending on the signal, users may have a harder time seeing short process excursions, small but meaningful changes, shifts in process variability, equipment degradation patterns, operating transitions, or relationships between measurements.

Consider a boiler temperature transmitter with a range from 0 to 1,600°F. The process may normally operate near 1,000°F, where a change of only a few degrees is useful information. A setting based only on a percentage of the entire transmitter span may not preserve that level of variation.

This is why the process matters as much as the instrument range.

What happens when compression is too loose?

The opposite problem also exists. Loose compression can preserve far more events than users need, including sensor noise, insignificant fluctuations, and repeated values that add little information.

More archived data does not automatically mean better historical data. If an instrument is accurate to ±0.5°F, preserving every 0.001°F change may add little useful process information. The goal is to retain the resolution needed by the process and its downstream uses.

There is no universal PI compression setting

A reactor temperature, tank level, flow measurement, valve position, laboratory result, and vibration signal behave differently. Their compression settings may need to differ as well.

Useful factors include sensor accuracy, process variability, engineering range, normal operating range, expected rate of change, collection frequency, measurement type, process criticality, and downstream use.

Standards are useful for managing large PI environments. Problems start when one standard is copied across measurements that behave very differently. For each important tag, teams should understand what process behavior needs to remain visible in the historical record.

The Point span is only part of the picture

Suppose a pressure transmitter has an engineering range from 0 to 1,000 psi, but the process normally operates between 495 and 505 psi. A change from 500 to 503 psi is small relative to the full transmitter span, but it may still be important to the process.

If compression settings are based only on the full 0 to 1,000 psi range, the configuration may not reflect the resolution the process team needs. When reviewing compression, consider both the instrument's range and the variation that matters during normal operation.

This becomes more important when PI history is used for analytics.

How compression affects PI trends

For normal trending, well-configured compression should be almost invisible. A PI Vision trend should still show the meaningful shape of the process.

Aggressive compression can make historical data appear smoother or less variable than the measurements originally collected. That difference may have little impact when an operator is checking a routine trend from yesterday, but it can matter much more when an engineer investigates an equipment event from six months ago.

Before using historical PI data for detailed analysis, check whether the available resolution fits the investigation.

How compression affects analytics

Analytics teams increasingly use years of PI history to find correlations, process transitions, rate-of-change behavior, equipment degradation, anomalies, failure signatures, and differences between operating modes.

Those applications can only analyze information that exists in the archive. If useful process variation was not archived, later analytics cannot recreate the original samples.

Compression does not need to be disabled for analytics. The settings need to preserve the process behavior required by the analysis.

How compression affects AI and machine learning

The same issue applies to AI and machine learning. A model does not observe the physical equipment directly. It learns from the historical data provided to it.

An equipment model may depend on subtle changes in pressure, temperature, flow, vibration, or power consumption. If the PI history does not preserve enough resolution to show those patterns, the model does not have access to them.

An important AI readiness question is whether the historical PI data contains enough process detail for the intended use case. Ten years of history can sound impressive, but the length of the history does not tell you whether it contains the information the model needs. That depends on how the data was collected, compressed, and stored.

Should compression be disabled for AI?

Usually, no. AI models do not automatically need every sensor sample, and raw data can contain redundant values and instrument noise.

A better approach is to work backward from the use case. Determine what behavior the model needs to detect, what time resolution it needs, what process variation matters, what measurement uncertainty exists, and whether the archived history preserves that information.

For some tags, the existing compression settings may already be suitable. Others may need closer review.

How can you identify questionable compression settings?

There is no single rule that identifies a bad compression setting, but several patterns are worth investigating:

  • Similar sensors have very different compression settings.

  • The Point span does not reflect normal process operation.

  • Important tags use unusually aggressive compression.

  • Tags store large amounts of measurement noise.

  • Compression settings changed without a documented reason.

  • Historical trends do not contain the detail users need.

  • Tags feed analytics or AI, but their historical resolution has never been reviewed.

  • One compression configuration was copied across unrelated measurement types.

Comparing similar assets is especially useful. Suppose twenty similar pumps use similar pressure transmitters. Nineteen have comparable compression settings and one is configured very differently. The different configuration may be correct, but someone should still be able to explain why it is different.

Compression is also a governance issue

Compression settings affect future users of the archived data. For important tags, teams should know what the current settings are, whether they are appropriate for the measurement, whether similar tags use similar configurations, when the settings changed, who changed them, and why.

They should also know which applications use the tag and what downstream systems could be affected by another change. A small PI Point configuration change can influence years of future historical data, which makes compression part of PI System governance for critical measurements.

PI Tag compression review checklist

You do not need to review every PI Tag at once. Start with data that supports important operational or analytical workflows.

For each tag, review the following areas.

Process

  • What does the measurement represent?

  • What variation matters to the process?

  • What is normal behavior?

  • What is the smallest meaningful change?

  • How quickly can an important event occur?

Instrumentation

  • What is the sensor accuracy?

  • What is the engineering range?

  • What is the normal operating range?

  • How frequently is the value collected?

PI configuration

  • Is compression enabled?

  • What is CompDev?

  • What is CompMax?

  • Does the Point span make sense?

  • Does exception reporting occur upstream?

  • How are similar tags configured?

Historical data

  • Does the archive preserve meaningful process changes?

  • Does it contain unnecessary noise?

  • Are short events visible?

  • Does the historical trend match the behavior engineers expect?

Usage

  • Is the tag used in PI Vision?

  • Does it feed PI Analyses?

  • Is it used for reliability or condition monitoring?

  • Does it feed analytics?

  • Is the history used for machine learning or AI?

Governance

  • Who owns the tag?

  • Who can change its settings?

  • Are configuration changes tracked?

  • Is the reason for the current configuration documented?

  • Can you identify downstream dependencies before changing it?

Where should you start?

A PI environment may contain hundreds of thousands of tags, so reviewing them one by one is rarely practical. Start with critical process measurements, tags used by important calculations, reliability and condition-monitoring data, tags used by critical PI Vision displays, data exported to analytics platforms, and data used for machine learning or AI.

Then compare similar measurements across similar equipment. Review pumps together, compressors together, pressure tags with other pressure tags, and temperature measurements with similar temperature measurements.

The configurations do not need to be identical. They should be intentional and explainable.

PI compression and data fidelity

Compression has traditionally been treated as a PI administration topic because it controls how historical time-series data is stored. The same data now feeds a much wider set of applications, which makes compression relevant to data quality and data fidelity.

Data fidelity describes how well digital data preserves and represents what happened in the physical process. Compression directly affects that historical representation.

A good configuration preserves the information users need without filling the archive with unnecessary events. A poor configuration can remove useful process behavior or retain large amounts of noise. The correct setting depends on the measurement, the process, and how the data will be used.

How Tycho Data helps

Tycho Data helps industrial teams understand the configuration, health, lineage, and use of operational data across the PI environment.

Osprey helps teams review PI configuration at scale, track changes, understand dependencies, and see which applications rely on important data. For compression, this gives teams more context than the configuration values alone.

Teams can see whether a setting changed, whether similar measurements are configured differently, and which calculations, displays, analytics, or AI applications depend on the affected data. This makes compression review part of a broader PI data quality and governance process.

Frequently asked questions

What is PI Tag compression?

PI Tag compression reduces the number of events stored in the PI Data Archive while preserving the useful shape of a process signal. It allows PI to maintain historical time-series data without storing every collected sample.

What is the difference between exception and compression in PI?

With traditional PI Interfaces, exception reporting can reduce the values sent toward the PI Server. Compression determines which eligible values are preserved in the PI Data Archive. They happen at different stages of the data path.

What is CompDev?

CompDev is the compression deviation used by the PI compression algorithm. It helps determine whether a signal has deviated enough from its expected trajectory to require another archived event. It is not simply a "store the value whenever it changes by X" setting.

What is CompMax?

CompMax is a time-based compression setting that helps ensure an event is periodically archived even when the signal stays within the compression criteria.

Can PI compression remove important data?

Yes. PI compression intentionally does not archive every collected value. When configured appropriately, the discarded values add little useful information. If compression is too aggressive for the process, the historical archive may not contain process behavior that users later need.

Does PI compression affect data quality?

It can. Compression becomes a data quality concern when the historical record no longer contains enough detail for its intended use.

Does PI compression affect AI?

It can. AI and machine learning models can only learn from information available in the historical data. If important process behavior was never archived, those original measurements cannot be recreated later from the PI Data Archive.

Should PI compression be disabled for AI?

Not automatically. Many AI use cases do not need every collected sensor value. Compression should be reviewed against the time resolution, process variation, and measurement accuracy required by the use case.

Should every PI Tag have the same compression settings?

No. Compression should reflect the measurement, instrument accuracy, process behavior, collection frequency, and downstream use. Standards can help manage large PI environments, but one setting should not be copied blindly across unrelated tags.

How often should compression settings be reviewed?

Review important compression settings when instrumentation, process behavior, or downstream requirements change. Teams should also track configuration changes and periodically compare settings across similar measurements.

Is PI compression only about storage?

No. Compression also affects the resolution of the historical process record, which makes it relevant to PI data quality, data fidelity, analytics, governance, and AI readiness.