How to Prepare Your PI System for AI: A Practical Readiness Checklist
Industrial companies are starting to connect operational data to AI, analytics, copilots, and other advanced applications. For many of them, the AVEVA PI System is one of the main sources of that operational data.
Having access to PI data does not automatically make it ready for AI. An AI application needs context to understand what a value represents, traceability to know where it came from, reliable access to the data, and enough confidence that the value still represents the physical process. Teams also need to understand where the data is used and how the PI environment has changed over time.
A practical PI System AI readiness assessment should review six areas:
Context
Traceability
Availability
Data fidelity
Usage and relevance
Auditability
Together, these areas help determine whether an AI application can understand and correctly use the PI data it depends on.
1. Start with context
Context is one of the most important requirements for industrial AI. A PI Tag can tell an application that a value is 162.4, but that number has limited meaning by itself.
An AI application may need to know what equipment the tag describes, what process that equipment supports, what the measurement represents, which engineering unit it uses, and whether the value is a measurement, calculation, setpoint, or another type of signal. Relationships to other measurements can matter as well.
The PI Asset Framework, or AF, can provide much of this context. AF organizes PI data around equipment, units, areas, sites, and other business structures. AF Attributes connect measurements to physical assets, while templates can provide a consistent structure across similar equipment.
For an AI use case, start with the PI data that matters to that specific application. Map important PI Tags to AF Attributes, review the AF Element hierarchy, verify engineering units, improve descriptions, and make sure equipment relationships are clear. Consistent AF Templates can also help when similar assets should have similar structures.
Trying to contextualize every historical tag before starting an AI project can create a lot of work without improving the use case. Focus first on the assets, measurements, and workflows that the AI application will actually use.
2. Understand tag traceability
AI applications should not have to treat every PI value as an isolated number. Teams need to understand where the value originated and what other systems or calculations depend on it.
A typical operational data path may look like this:
PLC or DCS → interface or connector → PI Data Archive → AF Attribute → PI Analysis → PI Vision → Power BI or AI application
A problem at one point in this chain can affect many downstream users. A PI Analysis, for example, may continue producing values after an upstream configuration changes. A PI Vision display may still look normal even though it references a different tag than an engineer expects.
Without traceability, it becomes difficult to determine whether an AI result is based on the correct source data. For important measurements, teams should be able to identify the source control system, source signal, PI Tag, AF Attribute, calculations, PI Vision displays, reports, analytics, and AI applications that use the data.
Indirect dependencies matter too. One calculated value may depend on several upstream measurements, some of which depend on other calculations or interfaces. Understanding the full dependency chain makes it much easier to troubleshoot unexpected results and assess the impact of a change.
3. Confirm data availability
The next question is whether the data is available when the AI application needs it. A tag can exist in PI and still be unsuitable for an AI workflow because it is stale, delayed, incomplete, or collected inconsistently.
Common availability problems include stale values, missing periods, interface failures, collection gaps, archive gaps, delayed data, incorrect scan rates, and intermittent communication failures.
The expected behavior depends on the measurement. A daily laboratory result should not be evaluated with the same availability rules as a pressure measurement that normally changes every few seconds.
For the data used by an AI application, define the expected update rate and acceptable latency. Check for stale tags, missing periods, unusual delays, and interface or connector problems. Compare actual collection behavior with the expected behavior of the measurement instead of applying one threshold across the entire PI environment.
This is especially important in large PI systems where tags represent very different types of process data. Availability should reflect how the measurement is supposed to behave.
4. Optimize tag compression
Availability tells you whether data is present. It does not tell you whether the value still represents what is happening in the physical process.
Data fidelity describes how well a digital value represents the real process or physical condition it is intended to measure. A PI Tag can update normally and still be wrong.
A sensor may be frozen at a plausible value. PLC scaling may be incorrect. A measurement can drift slowly away from the actual process condition. A tag can point to the wrong source, a calculation can use the wrong input, or a timestamp problem can shift data in time.
These problems are difficult because the data may still look valid. The value can remain inside a reasonable range and continue updating at the expected frequency while no longer representing the process correctly.
For important measurements, check for unexpected flatlines, abnormal ranges, sudden changes in normal behavior, scaling problems, timestamp issues, and sensor drift where it can be detected. Comparing related measurements and historical behavior can also help identify values that no longer make sense.
Context matters when evaluating compression. A stable tank level can legitimately remain flat for hours. A laboratory result may change only once per day. A pressure signal that normally varies every few seconds deserves more attention if it suddenly becomes constant.
Data quality rules need to reflect the process behind the signal.
5. Know what data matters and where it is used
Most PI environments contain a large number of tags, but they do not all have the same importance. Trying to clean, model, and govern every tag before an AI project usually creates more work than value.
Start with usage. Identify which PI data supports important operational decisions, calculations, displays, reports, analytics, or AI workflows.
A measurement used by several critical PI Vision displays and calculations deserves more attention than an old tag that no application uses. The same principle applies to AF Attributes, PI Analyses, and other PI resources.
Useful areas to review include frequently used PI Tags, critical AF Attributes, important PI Analyses, operational PI Vision displays, reliability workflows, Power BI reports, data exported to cloud platforms, and measurements used by machine learning or AI applications.
Classifying data by usage and business impact gives the AI readiness effort a practical scope. Teams can spend their time improving the data that would cause the most problems if it were wrong.
6. Maintain an audit trail
A PI environment changes continuously. Engineers modify PI Tag configuration, AF Elements and Attributes, templates, analyses, interfaces, connectors, data sources, displays, and calculation logic.
Those changes can alter the meaning or behavior of data that an AI application depends on. A system that is ready for an AI use case today may become less reliable after an upstream change tomorrow.
Teams should be able to determine what changed, when the change occurred, who made it, why it happened, and which downstream systems may be affected.
For important PI configuration, maintain an audit trail and connect changes to the data and applications that depend on them. If an engineer changes a PI Analysis, for example, the team should be able to identify the AF Attributes, PI Vision displays, reports, and AI workflows that may be affected.
This connection between change history and dependency information can reduce investigation time when data starts behaving differently.
PI System AI readiness checklist
Before connecting PI System data to an AI application, review the following areas.
Context
Are important PI Tags mapped to assets?
Are important measurements represented in AF?
Are equipment relationships clear?
Are engineering units correct?
Are descriptions useful?
Are AF Templates used consistently where appropriate?
Traceability
Can you identify the source of each important measurement?
Can you trace data through calculations?
Can you identify downstream PI Vision displays?
Can you identify reports and analytics that use the data?
Can you identify AI applications that consume it?
Availability
Are important tags updating as expected?
Are there stale tags?
Are there missing periods?
Are interfaces and connectors healthy?
Is data latency acceptable for the use case?
Data fidelity
Do measurements behave as expected?
Are there unexpected flatlines?
Are values inside reasonable ranges?
Is scaling correct?
Are timestamps correct?
Does the digital value still represent the physical process?
Usage and relevance
Do you know which tags are important?
Do you know which data supports critical workflows?
Can you identify unused or low-value data?
Can you prioritize AI readiness work based on business impact?
Auditability
Can you identify important PI configuration changes?
Do you know who made the change?
Do you know why it occurred?
Can you identify downstream systems that may be affected?
You do not need to clean your entire PI System before using AI
Treating AI readiness as a complete PI cleanup project can create a large amount of work before the organization gets any value from AI. For most companies, a use-case-driven approach is more practical.
Start with one AI use case and identify the assets, PI Tags, AF Attributes, calculations, displays, and applications that support it. Then evaluate those resources across context, traceability, availability, data fidelity, usage and relevance, and auditability.
This creates a defined scope while improving the operational data foundation around something the business is actually trying to accomplish. As new AI use cases are introduced, the same approach can expand to additional assets and datasets.
AI readiness needs continuous monitoring
Preparing PI data for an AI project is not something teams can complete once and ignore. Sensors fail, interfaces stop communicating, calculations change, tags are replaced, AF models evolve, and new applications begin consuming existing data.
Teams should continue checking whether critical data is available, whether signals behave as expected, whether dependencies have changed, and whether important PI configuration has been modified. They should also know when new consumers begin using data that was originally created for another purpose.
This is where operational data observability becomes useful. Continuous monitoring helps teams identify data problems and configuration changes before they spread into downstream displays, analytics, or AI applications.
How Tycho Data helps prepare PI System data for AI
Tycho Data helps industrial teams understand and monitor the operational data that supports analytics and AI. Osprey provides visibility into PI System data quality, dependencies, lineage, configuration changes, and downstream usage.
This gives teams a way to answer practical questions about their PI environment: Where did this value come from? Is the source data healthy? What changed? Which applications depend on it? What could be affected if the resource changes?
The objective is to build an operational data foundation that AI applications can understand, trace, and use with confidence, while focusing effort on the PI data that matters most.
Frequently asked questions
What does it mean for a PI System to be ready for AI?
A PI System is ready for an AI use case when the required operational data has enough context, traceability, availability, fidelity, relevance, and governance for the application to use it correctly. The requirements should match the specific AI use case, so readiness does not require every PI Tag to be cleaned or modeled.
Can AI use PI System data directly?
Yes. AI applications can use PI data directly or access it through other platforms and APIs. Direct access to values, however, may not provide enough information to interpret them correctly. Asset relationships, engineering units, source information, calculations, and data quality can all affect how a value should be understood.
Does PI data need to be in Asset Framework before it can be used by AI?
No. PI data can be used by AI without first being represented in Asset Framework. AF can still make the data easier to interpret because it connects measurements to assets, equipment structures, templates, and other context. Whether that context is required depends on the use case.
What PI System data quality problems can affect AI?
Common problems include stale data, missing periods, flatlined sensors, incorrect scaling, sensor drift, timestamp problems, failed calculations, and incorrect source mappings. Some problems are obvious, while others produce values that appear reasonable even though they no longer represent the process correctly.
What is the difference between data quality and data fidelity?
Data quality is a broad category that can include completeness, availability, consistency, validity, and other characteristics. Data fidelity focuses specifically on whether the digital value correctly represents the physical process or condition it is meant to represent.
A measurement can pass basic data quality checks and still have poor fidelity. For example, a drifting sensor can continue to update on time, remain within its expected range, and still report the wrong physical value.
Do I need to clean every PI Tag before I start an AI project?
No. Start with the data that supports the AI use case, then identify the related assets, tags, calculations, displays, and downstream applications. Assess those resources for context, traceability, availability, fidelity, usage, and auditability.
This keeps the work tied to business value and avoids spending months cleaning PI data that the AI application may never use.