Price Anomaly Detection in Nordic E-commerce Using Statistical and Unsupervised Machine Learning Methods
Iho, Albert (2026)
Iho, Albert
2026
Master's Programme in Computing Sciences and Electrical Engineering
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Hyväksymispäivämäärä
2026-05-29
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202605256319
https://urn.fi/URN:NBN:fi:tuni-202605256319
Tiivistelmä
E-commerce price monitoring systems depend on timely and reliable competitor-pricing data, but the collected price records may contain both statistical outliers and deterministic data-quality errors. This thesis investigates anomaly detection methods for competitor-pricing data under practical business and deployment constraints. The study compares statistical detectors, unsupervised forest-based detectors, and rule-based sanity checks, with the aim of identifying a detector configuration suitable for a layered detection system.
The evaluated statistical methods include the standard z-score, modified z-scores based on MAD and Sn, and hybrid robust z-score variants. The unsupervised methods include Isolation Forest, Extended Isolation Forest, and Robust Random Cut Forest. Because complete ground-truth anomaly labels are not available, the evaluation uses a reproducible synthetic anomaly injection protocol covering multiple anomaly cases, including price spikes, price drops, decimal shifts, currency swaps, zero prices, negative prices, and list-price violations. Detector performance is assessed across evaluation splits, aggregation granularities, and minimum-history settings using precision, recall, F1 score, specificity, and geometric mean.
The results show that the standard z-score and Isolation Forest provide the strongest candidates among the single-layer statistical and unsupervised methods, while more complex forest variants do not justify broader tuning under the examined constraints. The final layered evaluation indicates that combining sanity checks, a z-score layer, and Isolation Forest improves recall stability across anomaly cases and operating settings. The thesis concludes that a layered detector architecture provides a practical approach for anomaly detection in competitor-pricing pipelines.
The evaluated statistical methods include the standard z-score, modified z-scores based on MAD and Sn, and hybrid robust z-score variants. The unsupervised methods include Isolation Forest, Extended Isolation Forest, and Robust Random Cut Forest. Because complete ground-truth anomaly labels are not available, the evaluation uses a reproducible synthetic anomaly injection protocol covering multiple anomaly cases, including price spikes, price drops, decimal shifts, currency swaps, zero prices, negative prices, and list-price violations. Detector performance is assessed across evaluation splits, aggregation granularities, and minimum-history settings using precision, recall, F1 score, specificity, and geometric mean.
The results show that the standard z-score and Isolation Forest provide the strongest candidates among the single-layer statistical and unsupervised methods, while more complex forest variants do not justify broader tuning under the examined constraints. The final layered evaluation indicates that combining sanity checks, a z-score layer, and Isolation Forest improves recall stability across anomaly cases and operating settings. The thesis concludes that a layered detector architecture provides a practical approach for anomaly detection in competitor-pricing pipelines.
