VIDEO TUTORIAL
Predict Machine Failure From Sensor Data in Google Colab (Video)
Watch the end-to-end workflow, then build a failure classifier on the public AI4I 2020 dataset step by step, without the data leakage that fools most tutorials.

This predictive maintenance Python tutorial trains a model that flags machine failure from sensor readings in a free Google Colab notebook, using the public AI4I 2020 dataset. Watch the end-to-end workflow lesson first. Then follow the written steps: load, engineer features, split, fit a baseline, evaluate on recall, and keep leakage out.
In this lesson by Victor Tan, a predictive maintenance project is worked through end to end in a Jupyter notebook. Colab is a hosted Jupyter service, so the same workflow runs in it with nothing to install. The lesson sits in the end-to-end workflow module of our free Predictive Maintenance with Machine Learning course, where the task is to frame a project before modelling: failure modes in scope, sensors, warning time, metric, and who acts. Follow along, then check your work against these steps.
The dataset: AI4I 2020
The AI4I 2020 Predictive Maintenance Dataset is on the UCI Machine Learning Repository under a CC BY 4.0 licence. It is synthetic, built to reflect real predictive maintenance data because real failure data is hard to publish. It has 10,000 rows and 14 columns:
- Identifiers: UDI and Product ID.
- Type: L, M or H product quality.
- Five readings: air temperature, process temperature, rotational speed, torque and tool wear.
- Machine failure: the target, 1 or 0.
- Five failure-mode flags: TWF (tool wear), HDF (heat dissipation), PWF (power), OSF (overstrain) and RNF (random).
339 rows are failures, about 3.4%. That imbalance shapes every step that follows.

Step 1: load it in Colab
Open a new notebook at colab.research.google.com. Download the zip from the UCI page, upload ai4i2020.csv to the session, and read it with Pandas. Print the shape, the column types and the count of the target. If you do not see 10,000 rows and 339 failures, stop and find out why.
Step 2: drop identifiers and the answer columns
Drop UDI and Product ID. They identify rows; they do not describe the machine. Then drop TWF, HDF, PWF, OSF and RNF. These say which failure mode occurred. Leave them in and your model scores almost perfectly, because it is reading the answer. It is an easy mistake to make, and it is the first thing to check in any notebook on this dataset.
One-hot encode Type into three columns.
Step 3: features from physics
The UCI page documents how each failure mode is triggered, and each rule is physics an engineer already knows. Build three features from it:
- Temperature gap: process temperature minus air temperature. Heat dissipation failures happen when the gap is small and the speed is low.
- Power: torque times angular speed, where angular speed in rad/s is rpm times 2π divided by 60. Power failures sit outside a band.
- Strain: tool wear times torque. Overstrain failures happen above a limit that depends on the product type.
Be honest about what this means. Because the data was generated from these rules, these features make the problem easier than a real plant would. On real equipment you have to discover the physics yourself.
Step 4: split before you scale
Use a stratified 80/20 train/test split so both sets keep the same share of failures. Fit any scaler on the training set only, then apply it to the test set. Scaling the full dataset first lets test-set statistics leak into training.
Step 5: score the baseline
Before you train anything, score a model that always predicts "no failure". It gets 96.61% accuracy and a recall of zero. Every model you build has to beat that on recall, not on accuracy.
Then fit a logistic regression with balanced class weights, and after that a random forest or gradient boosting model. Keep the logistic regression: if the forest barely beats it, the simpler model is easier to explain to a maintenance manager.
Step 6: evaluate like a maintenance engineer
Print the confusion matrix. Read recall (the share of real failures you caught) and precision (the share of alarms that were real). Plot precision against recall across thresholds. Then choose the threshold by cost: a missed failure usually costs far more than an unnecessary inspection, so a lower threshold is often right. Choose it on validation data, not on the test set.
Pitfalls: where leakage hides

AI4I has no time order and no machine IDs to group by, so a stratified random split is acceptable here. Real plant data is different. When you move to NASA's C-MAPSS turbofan data, each engine starts healthy and develops a fault, and in the training set it runs until failure. The FD001 subset has 100 training and 100 test engines. Split by engine, not by row, or the model will recognise the engine instead of the wear. The course's remaining-useful-life module does exactly this with grouped cross-validation.
What to try next on the same data
Once the binary model works, turn the five failure-mode flags into targets instead of features. Train a model that predicts which mode is coming, not only that something is. Heat dissipation, power and overstrain failures each follow their own rule, so a model that separates them tells the maintenance team what to inspect.
Then check what the model relies on. If the forest's most important features are the three you built from physics, it has found the documented rules. If it leans on something odd, look for a leak before you celebrate.
Take it further in the free course
Predictive Maintenance with Machine Learning goes from the P-F curve and vibration spectra to anomaly detection on healthy data, fault classification grouped by motor, a 1-D CNN, C-MAPSS remaining useful life, and alerts that people act on. It needs Machine Learning with Python first; start with Python for AI and Engineering Data if you are new to code. All three sit on the Industrial AI engineer path, and a free account saves your progress.
The courses are free in full. The optional EDWartens Certificate of Completion for the intermediate predictive maintenance course is a small one-off fee, with a code anyone can check at edwartens.com/verification. It is not a vendor certification or an accredited qualification.
Take the free course
Questions
What dataset is best for a first predictive maintenance project?
The AI4I 2020 Predictive Maintenance Dataset on the UCI repository: 10,000 rows, clearly documented failure modes and a CC BY 4.0 licence. Move on to NASA's C-MAPSS engine data for remaining-useful-life work.
Why is accuracy misleading here?
Only 339 of the 10,000 rows are failures, so a model that always predicts no failure scores 96.61% accuracy and catches nothing. Report recall and precision on the failure class instead.
What is data leakage in this dataset?
The TWF, HDF, PWF, OSF and RNF columns record which failure mode occurred, so using them as features hands the model the answer. Drop them before training.
Can I follow the video in Google Colab?
Yes. Colab is a hosted Jupyter notebook service, so a Jupyter workflow runs the same way in it, free of charge and with nothing to install.
Sources
- UCI: AI4I 2020 Predictive Maintenance Dataset
- NASA Open Data Portal: CMAPSS Jet Engine Simulated Data
- Google Colab: Frequently Asked Questions
Written by the EDWartens engineering team for general education. Product names are trademarks of their owners; mentioning them does not imply endorsement. Prices and terms of other providers were checked on the date shown and can change.





