INDUSTRIAL DATA
Data Engineering for Industrial AI: Getting Plant Data Ready for a Model
Before a model can learn anything, plant signals have to be collected, time-aligned, cleaned and labelled. The five jobs, the traps in each, and a free route to learn them.

Data engineering for industrial AI means getting plant signals out of PLCs, historians and maintenance systems into one clean, time-aligned, labelled store that a model can learn from. In practice it is five jobs: collect, contextualise, store, clean and serve. Most of the effort sits in cleaning and labelling, not in the model.
Where plant data actually lives
Before building anything, list the sources. On a typical site they are spread across several systems, often owned by different teams:
- PLC and DCS tags: live process values, states, counters and alarms, read over the control network.
- The historian: years of compressed time-series data, and usually the best source for training.
- SCADA alarms and events: when something happened, and what the operator did.
- MES or batch records: product, recipe, order and batch start and end times, following the ISA-88 batch model where one is used.
- The maintenance system (CMMS): work orders, failure codes and dates, which are where most labels come from.
- Quality and lab results: measurements taken hours after the process, matched back by batch or time.
- Spreadsheets: shift logs and manual readings that nobody else has digitised.
ISA-95 is useful here as a map, as our MES and ISA-95 guide explains. It separates the control levels (sensors, PLCs, SCADA) from operations management (MES) and business planning (ERP), which tells you roughly who owns each source and what timescale it runs on.

Job 1: collect
In this lesson by APMonitor.com, from our free Industrial Data with Python course, data transfer with Modbus, OPC and SQL is covered. The notes below add the details that decide whether that data is usable.
[OPC UA](https://opcfoundation.org/about/opc-technologies/opc-ua/), standardised as IEC 62541, is the common way to read structured data from controllers and servers. Prefer subscriptions, where the server sends a value when it changes, over polling every tag every second. Each value comes with a timestamp and a status code telling you whether it is good, uncertain or bad. OPC UA explained goes further into its address space and security.
[Modbus](/blog/modbus-communication-protocol-tutorial) is older and simpler: numbered registers with no names, units or timestamps. You need the device's register map, and you must get the details right: 16-bit registers, how two registers combine into a 32-bit value, the word order, and the scaling factor.
[MQTT](https://mqtt.org/) is a lightweight publish and subscribe protocol. A gateway reads OPC UA or Modbus and publishes values to topics on a broker, and any number of consumers subscribe. It is well suited to moving data from the plant to a database, a dashboard and a model at the same time. Sparkplug B adds a standard payload and device state on top of MQTT, and the idea of a unified namespace, one shared topic tree for the whole site, builds on the same pattern.
Two rules apply whatever the protocol. Read only: a data collector should never write to a controller. And do not load the controller: a PLC answering thousands of fast polls can have its communications, and sometimes its scan, affected.
Job 2: contextualise
A tag called TT_101_PV means little to a data scientist. Record, for every tag, the asset it belongs to, what it measures, its engineering units and range, and how it is scaled. A 4-20 mA transmitter spanning 0 to 10 bar gives 5 bar at 12 mA. Signals outside the normal band indicate a fault, not a process value: under the NAMUR NE 43 convention, currents at or below 3.6 mA or at or above 21 mA signal a failure.
Build this into an asset model, site, area, line, machine, signal, so data from two identical machines can be compared without guessing which tag is which.
Job 3: store
For learning and small projects, SQLite is enough. For continuous data, use a time-series database such as InfluxDB or TimescaleDB, or query the plant historian directly. Store data in long form, one row per timestamp, tag, value and quality, and pivot to wide tables only when you build features.
Timestamps deserve their own rule: store everything in UTC and record whether each time came from the device or the server. Local time with daylight saving creates a missing hour and a duplicated hour every year, and they will surface in your model.
Job 4: clean
This is where the time goes.
- Grid the data. Signals arrive at different rates. Resample to a fixed interval, and choose honestly how: last value for states, mean for fast analogue signals.
- Know your historian's compression. Many historians store a value only when it changes by more than a deadband and interpolate between stored points. A long flat line may be a steady process or a dead sensor. Check the quality flags.
- Mark gaps, do not paper over them. Forward-fill for a short, stated limit. Never interpolate across a shutdown.
- Tag the run state. Running, idle, starting, stopped, cleaning. A model trained on a mix of all of them learns the schedule, not the machine.
- Match the labels. Join work orders and batch records to the signals by asset and time, and check them by eye. A failure logged two days after it happened is common.
- Watch for leakage. Any column written after the event you are predicting, such as a downtime reason code, must not be a feature.

Job 5: serve
The same clean store feeds several consumers: features for a model, a dashboard for the production team, and reports. Grafana is a common choice for live time-series dashboards; Power BI is common for production and OEE reporting. When a model goes live, the pipeline that built its training data must also build its live inputs, the same way. The deployment side is covered in the MLOps guide, and a complete worked example of a model on sensor data is in predictive maintenance in Google Colab.
A free route to learn data engineering for industrial AI
- SQL for PLC and SCADA Engineers is a beginner course on tables, queries, joins and schema design, all on machine, alarm and downtime records.
- Node-RED for Industrial IoT is a beginner course on reading Modbus and PLCs, publishing over MQTT, logging to SQL and building dashboards.
- Industrial Data with Python is the core: Modbus TCP and RTU, OPC UA subscriptions, MQTT topic design with quality and timestamps, SQLite or InfluxDB with Grafana, honest cleaning and gridding, forecasting and anomaly detection, ending with a bench pipeline project.
- Power BI for Manufacturing turns the clean data into OEE, downtime and reject dashboards.
The Industrial AI engineer path and the IIoT and OT security path group related courses.
Start the courses
Each course is free in full, with video lessons from independent creators credited on every course page, EDWartens notes, a practice task per module and one final assessment of 15 questions (60% to pass, three attempts). Sign up for a free account so your progress is saved.
The optional EDWartens Certificate of Completion is a small one-off fee, US$8.99 for a beginner course such as SQL or Node-RED and a little more for an intermediate one such as Industrial Data with Python. It can be checked at edwartens.com/verification. It is not a vendor certification and not an accredited qualification.
Take the free course
Questions
What is data engineering in industrial AI?
It is the work of getting plant signals from PLCs, historians and maintenance systems into a clean, time-aligned, labelled store that a model or a dashboard can use. It usually takes more effort than training the model.
Should I use OPC UA or MQTT to get PLC data?
Often both. OPC UA is the common way to read structured data from controllers and servers, with subscriptions and quality codes. MQTT is a lightweight publish and subscribe protocol well suited to moving that data on to many consumers.
What database should I store sensor data in?
For learning, SQLite is enough. For live data, a time-series database such as InfluxDB or TimescaleDB, or the plant's existing historian, handles high-rate timestamped data better than a general table design.
Why is my historian data full of flat lines?
Many historians store a new value only when it changes by more than a deadband, and interpolate between stored points. A flat line may mean no change beyond the deadband, or a stale sensor. Check the tag's compression settings and quality flags before trusting it.
Where do labels for plant machine learning come from?
From dated records: maintenance work orders, downtime logs, batch and quality records, lab results. They are often incomplete, so reconciling them with the signals is part of the job.
Sources
- OPC Foundation: OPC Unified Architecture
- MQTT: the standard for IoT messaging
- Eclipse Foundation: Sparkplug working group
- NAMUR: current recommendations and worksheets, including NE 43
Written by the EDWartens engineering team for general education. Product names are trademarks of their owners; mentioning them does not imply endorsement. Prices and terms of other providers were checked on the date shown and can change.

Edge AI in Industrial Automation: When to Run Models on the Plant Floor

MQTT Sparkplug B Explained: Topics, Birth/Death and UNS (Video)

Node-RED PLC Dashboard: Read Live Data and Build an Operator Page (Video)

IIoT Training: How Plant Data Gets From a PLC to a Dashboard, and How to Learn It Free

Python for Automation Engineers: What to Learn and What to Build

Industry 4.0 and Smart Manufacturing: A Practical Guide for Engineers



