Impact of data preprocessing strategies on the accuracy of linear regression algorithms
DOI:
https://doi.org/10.55204/trc.v6i2.e724Keywords:
data preprocessing, linear regression, normalization, imputation, machine learning, systematic reviewAbstract
Data preprocessing is a crucial stage in the performance of linear regression models, where data quality including the handling of missing values, outlier detection, normalization, and variable encoding directly influences the predictive capability of machine learning algorithms. This systematic review, conducted under the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) methodology, analyzes 40 scientific articles published between 2021 and 2026 in well-recognized databases such as Scopus, Web of Science, Springer Nature, MDPI, IEEE, and ScienceDirect. Within this context, linear regression models are highlighted as a fundamental tool in artificial intelligence and data analysis, where the quality of preprocessing directly affects the accuracy of the results obtained. During the development of the study, key techniques were examined, including data cleaning, missing value imputation, normalization or standardization, as well as the encoding of categorical variables. Regarding practical applications and future research directions, the findings support their implementation in predictive models related to sales, price estimation, and academic indicators. They also encourage extending the analysis to non-linear models and higher-dimensional datasets. Ultimately, proper data preprocessing significantly improves the predictive accuracy of linear regression models, helping to reduce issues such as overfitting and underfitting.Downloads
References
Ahsan, M. M., Mahmud, M. A. P., Saha, P. K., Gupta, K. D., & Siddique, Z. (2021). Effect of Data Scaling Methods on Machine Learning Algorithms and Model Performance. Technologies, 9(3), 52. https://doi.org/10.3390/technologies9030052
Alshdaifat, E., Alshdaifat, D., Alsarhan, A., Hussein, F., & El-Salhi, S. M. F. S. (2021). The Effect of Preprocessing Techniques, Applied to Numeric Features, on Classification Algorithms' Performance. Data, 6(2), 11. https://doi.org/10.3390/data6020011
Amini, M., Roozbeh, M., & Mohamed, N. A. (2024). Separation of the Linear and Nonlinear Covariates in the Sparse Semi-Parametric Regression Model in the Presence of Outliers. Mathematics, 12(2), 172. https://doi.org/10.3390/math12020172
Anelli, D., Morano, P., Tajani, F., & Guarini, M. R. (2025). The Interpretative Effects of Normalization Techniques on Complex Regression Modeling. Information, 16(6), 486. https://doi.org/10.3390/info16060486
Avelino, J. G., Cavalcanti, G. D. C., & Cruz, R. M. O. (2024). Resampling strategies for imbalanced regression: A survey and empirical analysis. Artificial Intelligence Review, 57(4), 82. https://doi.org/10.1007/s10462-024-10724-3
Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553
Buczak, P., Chen, J.-J., & Pauly, M. (2023). Analyzing the Effect of Imputation on Classification Performance under MCAR and MAR Missing Mechanisms. Entropy, 25(3), 521. https://doi.org/10.3390/e25030521
Cabello-Solorzano, K., Ortigosa de Araujo, I., Peña, M., Correia, L., & J. Tallón-Ballesteros, A. (2023). The Impact of Data Normalization on the Accuracy of Machine Learning Algorithms: A Comparative Analysis. En 18th International Conference on Soft Computing Models (SOCO 2023) (pp. 344-353). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-42536-3_33
Chen, J., & Yang, Z. (2024). Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms. Applied Sciences, 14(19), 8922. https://doi.org/10.3390/app14198922
Chowdhury, S., Lin, Y., Liaw, B., & Kerby, L. (2022). Evaluation of Tree Based Regression over Multiple Linear Regression for Non-normally Distributed Data in Battery Performance. 2022 International Conference on Intelligent Data Science Technologies and Applications (IDSTA), 17-25. https://doi.org/10.1109/IDSTA55301.2022.9923169
de Amorim, L. B. V., Cavalcanti, G. D. C., & Cruz, R. M. O. (2023). The choice of scaling technique matters for classification performance. Applied Soft Computing, 133, 109924. https://doi.org/10.1016/j.asoc.2022.109924
Demir, S., & Sahin, E. K. (2024). The effectiveness of data pre-processing methods on the performance of machine learning techniques using RF, SVR, Cubist and SGB. Stochastic Environmental Research and Risk Assessment, 38(8), 3273-3290. https://doi.org/10.1007/s00477-024-02745-9
Gan, Q., Gong, L., Hu, D., Jiang, Y., & Ding, X. (2023). A Hybrid Missing Data Imputation Method for Batch Process Monitoring Dataset. Sensors, 23(21), 8678. https://doi.org/10.3390/s23218678
Habib, M., & Okayli, M. (2024). Evaluating the Sensitivity of Machine Learning Models to Data Preprocessing Technique in Concrete Compressive Strength Estimation. Arabian Journal for Science and Engineering, 49(10), 13709-13727. https://doi.org/10.1007/s13369-024-08776-2
Hassanat, A. B., Alqaralleh, M. K., Tarawneh, A. S., Almohammadi, K., Alamri, M., Alzahrani, A., Altarawneh, G. A., & Alhalaseh, R. (2024). A Novel Outlier-Robust Accuracy Measure for Machine Learning Regression Using a Non-Convex Distance Metric. Mathematics, 12(22), 3623. https://doi.org/10.3390/math12223623
Hippel, P. T. von, & Bartlett, J. W. (2021). Maximum Likelihood Multiple Imputation: Faster Imputations and Consistent Standard Errors Without Posterior Draws. Statistical Science, 36(3), 400-420. https://doi.org/10.1214/20-STS793
Kappatou, C. D., Odgers, J., García-Muñoz, S., & Misener, R. (2023). An Optimization Approach Coupling Preprocessing with Model Regression for Enhanced Chemometrics. Industrial & Engineering Chemistry Research. https://doi.org/10.1021/acs.iecr.2c04583
Koukaras, P., & Tjortjis, C. (2025). Data Preprocessing and Feature Engineering for Data Mining: Techniques, Tools, and Best Practices. AI, 6(10), 257. https://doi.org/10.3390/ai6100257
Li, C., Ren, X., & Zhao, G. (2023). Machine-Learning-Based Imputation Method for Filling Missing Values in Ground Meteorological Observation Data. Algorithms, 16(9), 422. https://doi.org/10.3390/a16090422
Li, F., Sun, H., Gu, Y., & Yu, G. (2023). A Noise-Aware Multiple Imputation Algorithm for Missing Data. Mathematics, 11(1), 73. https://doi.org/10.3390/math11010073
Maharana, K., Mondal, S., & Nemade, B. (2022). A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 3(1), 91-99. https://doi.org/10.1016/j.gltp.2022.04.020
Mahmud Sujon, K., Binti Hassan, R., Tusnia Towshi, Z., Othman, M. A., Abdus Samad, M., & Choi, K. (2024). When to Use Standardization and Normalization: Empirical Evidence From Machine Learning Models and XAI. IEEE Access, 12, 135300-135314. https://doi.org/10.1109/ACCESS.2024.3462434
Mallikharjuna Rao, K., Saikrishna, G., & Supriya, K. (2023). Data preprocessing techniques: Emergence and selection towards machine learning models – a practical review using HPA dataset. Multimedia Tools and Applications, 82(24), 37177-37196. https://doi.org/10.1007/s11042-023-15087-5
Mu, W., Cardelli, R., & Ferrari, S. (2026). Data Preprocessing Techniques for Machine Learning Towards Improving Building Energy Performance: A Systematic Review. Energies, 19(6), 1561. https://doi.org/10.3390/en19061561
Ou, H., Yao, Y., & He, Y. (2024). Missing Data Imputation Method Combining Random Forest and Generative Adversarial Imputation Network. Sensors, 24(4), 1112. https://doi.org/10.3390/s24041112
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … Moher, D. (2021a). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
Page, M. J., Moher, D., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … McKenzie, J. E. (2021b). PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ, 372, n160. https://doi.org/10.1136/bmj.n160
Pandit, P., Dey, P., & Krishnamurthy, K. N. (2021). Comparative Assessment of Multiple Linear Regression and Fuzzy Linear Regression Models. SN Computer Science, 2(2), 76. https://doi.org/10.1007/s42979-021-00473-3
Pargent, F., Pfisterer, F., Thomas, J., & Bischl, B. (2022). Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features. Computational Statistics, 37(5), 2671-2692. https://doi.org/10.1007/s00180-022-01207-6
Park, H.-J., Koo, Y.-S., Yang, H.-Y., Han, Y.-S., & Nam, C.-S. (2024). Study on Data Preprocessing for Machine Learning Based on Semiconductor Manufacturing Processes. Sensors, 24(17), 5461. https://doi.org/10.3390/s24175461
Pinheiro, J. M. H., Oliveira, S. V. B. de, Silva, T. H. S., Saraiva, P. A. R., Souza, E. F. de, Godoy, R. V., Ambrosio, L. A., & Becker, M. (2025). The Impact of Feature Scaling in Machine Learning: Effects on Regression and Classification Tasks. IEEE Access, 13, 199903-199931. https://doi.org/10.1109/ACCESS.2025.3635541
Rasyidah, Efendi, R., Nawi, N. Mohd., Deris, M. M., & Burney, S. M. A. (2023). Cleansing of inconsistent sample in linear regression model based on rough sets theory. Systems and Soft Computing, 5, 200046. https://doi.org/10.1016/j.sasc.2022.200046
Rauf, R. I., Alrasheedi, M. A., Sadiq, R., & Aldawsari, A. M. A. (2024). Evaluating Predictive Accuracy of Regression Models with First-Order Autoregressive Disturbances: A Comparative Approach Using Artificial Neural Networks and Classical Estimators. Mathematics, 12(24), 3966. https://doi.org/10.3390/math12243966
Shin, W. G., Lee, J. S., Ju, Y. C., Hwang, H. M., & Ko, S. W. (2025). Data preprocessing and machine learning method based on ameliorated mathematical models for inferring the power generation of photovoltaic system. Energy Conversion and Management, 333, 119793. https://doi.org/10.1016/j.enconman.2025.119793
Tawakuli, A., Havers, B., Gulisano, V., Kaiser, D., & Engel, T. (2025). Survey: Time-series data preprocessing: A survey and an empirical analysis. Journal of Engineering Research, 13(2), 674-711. https://doi.org/10.1016/j.jer.2024.02.018
Veltri, G. A. (2025). The Effects of Data Preprocessing Choices on Behavioral RCT Outcomes: A Multiverse Analysis. Multivariate Behavioral Research, 0(0), 1-16. https://doi.org/10.1080/00273171.2025.2575399
Vescan, A., Găceanu, R., & Şerban, C. (2024). Exploring the impact of data preprocessing techniques on composite classifier algorithms in cross-project defect prediction. Automated Software Engineering, 31(2), 47. https://doi.org/10.1007/s10515-024-00454-9
Wanyonyi, E. N., & Masinde, N. W. (2025). The Impact of Data Preprocessing on Machine Learning Model Performance: A Comprehensive Examination. IJSRCSEIT, 11(2), 3814-3827. https://doi.org/10.32628/CSEIT25112854
Zeng, G., & Tao, S. (2023). A Generalized Linear Transformation and Its Effects on Logistic Regression. Mathematics, 11(2), 467. https://doi.org/10.3390/math11020467
Zheng, D., Hao, X., Khan, M., Wang, L., Li, F., Xiang, N., Kang, F., Hamalainen, T., Cong, F., Song, K., & Qiao, C. (2022). Comparison of machine learning and logistic regression as predictive models for adverse maternal and neonatal outcomes of preeclampsia: A retrospective study. Frontiers in Cardiovascular Medicine, 9. https://doi.org/10.3389/fcvm.2022.959649
Zhong, J., Ma, C., Hou, L., Yin, Y., Zhao, F., Hu, Y., Song, A., Wang, D., Li, L., Cheng, X., & Qiu, L. (2023). Utilization of five data mining algorithms combined with simplified preprocessing to establish reference intervals of thyroid-related hormones for non-elderly adults. BMC Medical Research Methodology, 23(1), 108. https://doi.org/10.1186/s12874-023-01898-5
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Jaime David Camacho Castillo, Mauricio Alexander Álvarez Ortega, Iván Alejandro Daqui Cabezas, Alex Nicolas Mazacon Guamingo, Juan Enrique Olivo Quintero, Jordy Joel Vele Carrera

This work is licensed under a Creative Commons Attribution 4.0 International License.
The authors retain the moral and patrimonial rights of their works. They only give to the magazine Tesla Revista Científica the right to the first publication of this. Since Tesla Revista Científica is an open access publication, readers can fully or partially reproduce its content as long as they properly credit the corresponding authors and the journal itself. Tesla Revista Científica undertakes not to make commercial use of the texts it receives and/or publishes.
Our journal is governed by the international policies SHERPA/RoMEO: Green journal: They allow the self-archiving of both the pre-print (draft of a paper) and the post-print (the version corrected and reviewed by peers) and even the final version ( layout as it will be published in the journal).
See also "Copyright and licences".



