Impact of data preprocessing strategies on the accuracy of linear regression algorithms

Authors

  • Jaime David Camacho Castillo
  • Mauricio Alexander Álvarez Ortega https://orcid.org/0009-0005-4850-9425
  • Iván Alejandro Daqui Cabezas
  • Alex Nicolas Mazacon Guamingo
  • Juan Enrique Olivo Quintero
  • Jordy Joel Vele Carrera

DOI:

https://doi.org/10.55204/trc.v6i2.e724

Keywords:

data preprocessing, linear regression, normalization, imputation, machine learning, systematic review

Abstract

Data preprocessing is a crucial stage in the performance of linear regression models, where data quality including the handling of missing values, outlier detection, normalization, and variable encoding directly influences the predictive capability of machine learning algorithms. This systematic review, conducted under the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) methodology, analyzes 40 scientific articles published between 2021 and 2026 in well-recognized databases such as Scopus, Web of Science, Springer Nature, MDPI, IEEE, and ScienceDirect. Within this context, linear regression models are highlighted as a fundamental tool in artificial intelligence and data analysis, where the quality of preprocessing directly affects the accuracy of the results obtained. During the development of the study, key techniques were examined, including data cleaning, missing value imputation, normalization or standardization, as well as the encoding of categorical variables. Regarding practical applications and future research directions, the findings support their implementation in predictive models related to sales, price estimation, and academic indicators. They also encourage extending the analysis to non-linear models and higher-dimensional datasets. Ultimately, proper data preprocessing significantly improves the predictive accuracy of linear regression models, helping to reduce issues such as overfitting and underfitting.

Downloads

Download data is not yet available.

References

Ahsan, M. M., Mahmud, M. A. P., Saha, P. K., Gupta, K. D., & Siddique, Z. (2021). Effect of Data Scaling Methods on Machine Learning Algorithms and Model Performance. Technologies, 9(3), 52. https://doi.org/10.3390/technologies9030052

Alshdaifat, E., Alshdaifat, D., Alsarhan, A., Hussein, F., & El-Salhi, S. M. F. S. (2021). The Effect of Preprocessing Techniques, Applied to Numeric Features, on Classification Algorithms' Performance. Data, 6(2), 11. https://doi.org/10.3390/data6020011

Amini, M., Roozbeh, M., & Mohamed, N. A. (2024). Separation of the Linear and Nonlinear Covariates in the Sparse Semi-Parametric Regression Model in the Presence of Outliers. Mathematics, 12(2), 172. https://doi.org/10.3390/math12020172

Anelli, D., Morano, P., Tajani, F., & Guarini, M. R. (2025). The Interpretative Effects of Normalization Techniques on Complex Regression Modeling. Information, 16(6), 486. https://doi.org/10.3390/info16060486

Avelino, J. G., Cavalcanti, G. D. C., & Cruz, R. M. O. (2024). Resampling strategies for imbalanced regression: A survey and empirical analysis. Artificial Intelligence Review, 57(4), 82. https://doi.org/10.1007/s10462-024-10724-3

Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553

Buczak, P., Chen, J.-J., & Pauly, M. (2023). Analyzing the Effect of Imputation on Classification Performance under MCAR and MAR Missing Mechanisms. Entropy, 25(3), 521. https://doi.org/10.3390/e25030521

Cabello-Solorzano, K., Ortigosa de Araujo, I., Peña, M., Correia, L., & J. Tallón-Ballesteros, A. (2023). The Impact of Data Normalization on the Accuracy of Machine Learning Algorithms: A Comparative Analysis. En 18th International Conference on Soft Computing Models (SOCO 2023) (pp. 344-353). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-42536-3_33

Chen, J., & Yang, Z. (2024). Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms. Applied Sciences, 14(19), 8922. https://doi.org/10.3390/app14198922

Chowdhury, S., Lin, Y., Liaw, B., & Kerby, L. (2022). Evaluation of Tree Based Regression over Multiple Linear Regression for Non-normally Distributed Data in Battery Performance. 2022 International Conference on Intelligent Data Science Technologies and Applications (IDSTA), 17-25. https://doi.org/10.1109/IDSTA55301.2022.9923169

de Amorim, L. B. V., Cavalcanti, G. D. C., & Cruz, R. M. O. (2023). The choice of scaling technique matters for classification performance. Applied Soft Computing, 133, 109924. https://doi.org/10.1016/j.asoc.2022.109924

Demir, S., & Sahin, E. K. (2024). The effectiveness of data pre-processing methods on the performance of machine learning techniques using RF, SVR, Cubist and SGB. Stochastic Environmental Research and Risk Assessment, 38(8), 3273-3290. https://doi.org/10.1007/s00477-024-02745-9

Gan, Q., Gong, L., Hu, D., Jiang, Y., & Ding, X. (2023). A Hybrid Missing Data Imputation Method for Batch Process Monitoring Dataset. Sensors, 23(21), 8678. https://doi.org/10.3390/s23218678

Habib, M., & Okayli, M. (2024). Evaluating the Sensitivity of Machine Learning Models to Data Preprocessing Technique in Concrete Compressive Strength Estimation. Arabian Journal for Science and Engineering, 49(10), 13709-13727. https://doi.org/10.1007/s13369-024-08776-2

Hassanat, A. B., Alqaralleh, M. K., Tarawneh, A. S., Almohammadi, K., Alamri, M., Alzahrani, A., Altarawneh, G. A., & Alhalaseh, R. (2024). A Novel Outlier-Robust Accuracy Measure for Machine Learning Regression Using a Non-Convex Distance Metric. Mathematics, 12(22), 3623. https://doi.org/10.3390/math12223623

Hippel, P. T. von, & Bartlett, J. W. (2021). Maximum Likelihood Multiple Imputation: Faster Imputations and Consistent Standard Errors Without Posterior Draws. Statistical Science, 36(3), 400-420. https://doi.org/10.1214/20-STS793

Kappatou, C. D., Odgers, J., García-Muñoz, S., & Misener, R. (2023). An Optimization Approach Coupling Preprocessing with Model Regression for Enhanced Chemometrics. Industrial & Engineering Chemistry Research. https://doi.org/10.1021/acs.iecr.2c04583

Koukaras, P., & Tjortjis, C. (2025). Data Preprocessing and Feature Engineering for Data Mining: Techniques, Tools, and Best Practices. AI, 6(10), 257. https://doi.org/10.3390/ai6100257

Li, C., Ren, X., & Zhao, G. (2023). Machine-Learning-Based Imputation Method for Filling Missing Values in Ground Meteorological Observation Data. Algorithms, 16(9), 422. https://doi.org/10.3390/a16090422

Li, F., Sun, H., Gu, Y., & Yu, G. (2023). A Noise-Aware Multiple Imputation Algorithm for Missing Data. Mathematics, 11(1), 73. https://doi.org/10.3390/math11010073

Maharana, K., Mondal, S., & Nemade, B. (2022). A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 3(1), 91-99. https://doi.org/10.1016/j.gltp.2022.04.020

Mahmud Sujon, K., Binti Hassan, R., Tusnia Towshi, Z., Othman, M. A., Abdus Samad, M., & Choi, K. (2024). When to Use Standardization and Normalization: Empirical Evidence From Machine Learning Models and XAI. IEEE Access, 12, 135300-135314. https://doi.org/10.1109/ACCESS.2024.3462434

Mallikharjuna Rao, K., Saikrishna, G., & Supriya, K. (2023). Data preprocessing techniques: Emergence and selection towards machine learning models – a practical review using HPA dataset. Multimedia Tools and Applications, 82(24), 37177-37196. https://doi.org/10.1007/s11042-023-15087-5

Mu, W., Cardelli, R., & Ferrari, S. (2026). Data Preprocessing Techniques for Machine Learning Towards Improving Building Energy Performance: A Systematic Review. Energies, 19(6), 1561. https://doi.org/10.3390/en19061561

Ou, H., Yao, Y., & He, Y. (2024). Missing Data Imputation Method Combining Random Forest and Generative Adversarial Imputation Network. Sensors, 24(4), 1112. https://doi.org/10.3390/s24041112

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … Moher, D. (2021a). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71

Page, M. J., Moher, D., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … McKenzie, J. E. (2021b). PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ, 372, n160. https://doi.org/10.1136/bmj.n160

Pandit, P., Dey, P., & Krishnamurthy, K. N. (2021). Comparative Assessment of Multiple Linear Regression and Fuzzy Linear Regression Models. SN Computer Science, 2(2), 76. https://doi.org/10.1007/s42979-021-00473-3

Pargent, F., Pfisterer, F., Thomas, J., & Bischl, B. (2022). Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features. Computational Statistics, 37(5), 2671-2692. https://doi.org/10.1007/s00180-022-01207-6

Park, H.-J., Koo, Y.-S., Yang, H.-Y., Han, Y.-S., & Nam, C.-S. (2024). Study on Data Preprocessing for Machine Learning Based on Semiconductor Manufacturing Processes. Sensors, 24(17), 5461. https://doi.org/10.3390/s24175461

Pinheiro, J. M. H., Oliveira, S. V. B. de, Silva, T. H. S., Saraiva, P. A. R., Souza, E. F. de, Godoy, R. V., Ambrosio, L. A., & Becker, M. (2025). The Impact of Feature Scaling in Machine Learning: Effects on Regression and Classification Tasks. IEEE Access, 13, 199903-199931. https://doi.org/10.1109/ACCESS.2025.3635541

Rasyidah, Efendi, R., Nawi, N. Mohd., Deris, M. M., & Burney, S. M. A. (2023). Cleansing of inconsistent sample in linear regression model based on rough sets theory. Systems and Soft Computing, 5, 200046. https://doi.org/10.1016/j.sasc.2022.200046

Rauf, R. I., Alrasheedi, M. A., Sadiq, R., & Aldawsari, A. M. A. (2024). Evaluating Predictive Accuracy of Regression Models with First-Order Autoregressive Disturbances: A Comparative Approach Using Artificial Neural Networks and Classical Estimators. Mathematics, 12(24), 3966. https://doi.org/10.3390/math12243966

Shin, W. G., Lee, J. S., Ju, Y. C., Hwang, H. M., & Ko, S. W. (2025). Data preprocessing and machine learning method based on ameliorated mathematical models for inferring the power generation of photovoltaic system. Energy Conversion and Management, 333, 119793. https://doi.org/10.1016/j.enconman.2025.119793

Tawakuli, A., Havers, B., Gulisano, V., Kaiser, D., & Engel, T. (2025). Survey: Time-series data preprocessing: A survey and an empirical analysis. Journal of Engineering Research, 13(2), 674-711. https://doi.org/10.1016/j.jer.2024.02.018

Veltri, G. A. (2025). The Effects of Data Preprocessing Choices on Behavioral RCT Outcomes: A Multiverse Analysis. Multivariate Behavioral Research, 0(0), 1-16. https://doi.org/10.1080/00273171.2025.2575399

Vescan, A., Găceanu, R., & Şerban, C. (2024). Exploring the impact of data preprocessing techniques on composite classifier algorithms in cross-project defect prediction. Automated Software Engineering, 31(2), 47. https://doi.org/10.1007/s10515-024-00454-9

Wanyonyi, E. N., & Masinde, N. W. (2025). The Impact of Data Preprocessing on Machine Learning Model Performance: A Comprehensive Examination. IJSRCSEIT, 11(2), 3814-3827. https://doi.org/10.32628/CSEIT25112854

Zeng, G., & Tao, S. (2023). A Generalized Linear Transformation and Its Effects on Logistic Regression. Mathematics, 11(2), 467. https://doi.org/10.3390/math11020467

Zheng, D., Hao, X., Khan, M., Wang, L., Li, F., Xiang, N., Kang, F., Hamalainen, T., Cong, F., Song, K., & Qiao, C. (2022). Comparison of machine learning and logistic regression as predictive models for adverse maternal and neonatal outcomes of preeclampsia: A retrospective study. Frontiers in Cardiovascular Medicine, 9. https://doi.org/10.3389/fcvm.2022.959649

Zhong, J., Ma, C., Hou, L., Yin, Y., Zhao, F., Hu, Y., Song, A., Wang, D., Li, L., Cheng, X., & Qiu, L. (2023). Utilization of five data mining algorithms combined with simplified preprocessing to establish reference intervals of thyroid-related hormones for non-elderly adults. BMC Medical Research Methodology, 23(1), 108. https://doi.org/10.1186/s12874-023-01898-5

Downloads

Published

2026-09-08

Issue

Section

Original Research Articles

How to Cite

Camacho Castillo, J. D., Álvarez Ortega, M. A., Daqui Cabezas, I. A., Mazacon Guamingo, A. N., Olivo Quintero, J. E., & Vele Carrera, J. J. (2026). Impact of data preprocessing strategies on the accuracy of linear regression algorithms. Tesla Revista Científica, 6(2), e724. https://doi.org/10.55204/trc.v6i2.e724

Most read articles by the same author(s)