Impacto de las estrategias de preprocesamiento de datos en la precisión de algoritmos de regresión lineal
DOI:
https://doi.org/10.55204/trc.v6i2.e724Palabras clave:
preprocesamiento de datos, regresión lineal, normalización, imputación, aprendizaje automático, revisión sistemáticaResumen
El preprocesamiento de datos es una etapa determinante en el rendimiento de los modelos de regresión lineal, donde la calidad de los datos, incluyendo el tratamiento de valores faltantes, la detección de outliers, la normalización y la codificación de variables, influyen directamente en la capacidad predictiva de los algoritmos de aprendizaje automático. Esta revisión sistemática, conducida bajo la metodología PRISMA 2020 (Preferred Reporting Items for Systematic reviews and Meta-Analyses), bajo esta misma óptica, se analiza 40 artículos científicos publicados entre 2021 y 2026 en bases de datos reconocidas como Scopus, Web of Science, Springer Nature, MDPI, IEEE y ScienceDirect. Dentro de este contexto, se destaca que los modelos de regresión lineal constituyen una herramienta fundamental en la inteligencia artificial y el análisis de datos, donde la calidad del preprocesamiento influye directamente en la precisión de los resultados obtenidos. En el desarrollo del estudio, se analizaron técnicas clave como la limpieza de datos, la imputación de valores faltantes, la normalización o estandarización y a esto se suma la codificación de variables categóricas. En cuanto a las aplicaciones prácticas y futuras líneas de investigación, los hallazgos permiten su implementación en modelos predictivos relacionados con ventas, estimación de precios e indicadores académicos, como también, plantear la extensión del análisis hacia modelos no lineales y conjuntos de datos de mayor dimensionalidad. No obstante, un adecuado preprocesamiento de datos mejora significativamente la precisión predictiva de los modelos de regresión lineal, contribuyendo a reducir problemas de overfitting y underfitting.Descargas
Referencias
Ahsan, M. M., Mahmud, M. A. P., Saha, P. K., Gupta, K. D., & Siddique, Z. (2021). Effect of Data Scaling Methods on Machine Learning Algorithms and Model Performance. Technologies, 9(3), 52. https://doi.org/10.3390/technologies9030052
Alshdaifat, E., Alshdaifat, D., Alsarhan, A., Hussein, F., & El-Salhi, S. M. F. S. (2021). The Effect of Preprocessing Techniques, Applied to Numeric Features, on Classification Algorithms' Performance. Data, 6(2), 11. https://doi.org/10.3390/data6020011
Amini, M., Roozbeh, M., & Mohamed, N. A. (2024). Separation of the Linear and Nonlinear Covariates in the Sparse Semi-Parametric Regression Model in the Presence of Outliers. Mathematics, 12(2), 172. https://doi.org/10.3390/math12020172
Anelli, D., Morano, P., Tajani, F., & Guarini, M. R. (2025). The Interpretative Effects of Normalization Techniques on Complex Regression Modeling. Information, 16(6), 486. https://doi.org/10.3390/info16060486
Avelino, J. G., Cavalcanti, G. D. C., & Cruz, R. M. O. (2024). Resampling strategies for imbalanced regression: A survey and empirical analysis. Artificial Intelligence Review, 57(4), 82. https://doi.org/10.1007/s10462-024-10724-3
Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553
Buczak, P., Chen, J.-J., & Pauly, M. (2023). Analyzing the Effect of Imputation on Classification Performance under MCAR and MAR Missing Mechanisms. Entropy, 25(3), 521. https://doi.org/10.3390/e25030521
Cabello-Solorzano, K., Ortigosa de Araujo, I., Peña, M., Correia, L., & J. Tallón-Ballesteros, A. (2023). The Impact of Data Normalization on the Accuracy of Machine Learning Algorithms: A Comparative Analysis. En 18th International Conference on Soft Computing Models (SOCO 2023) (pp. 344-353). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-42536-3_33
Chen, J., & Yang, Z. (2024). Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms. Applied Sciences, 14(19), 8922. https://doi.org/10.3390/app14198922
Chowdhury, S., Lin, Y., Liaw, B., & Kerby, L. (2022). Evaluation of Tree Based Regression over Multiple Linear Regression for Non-normally Distributed Data in Battery Performance. 2022 International Conference on Intelligent Data Science Technologies and Applications (IDSTA), 17-25. https://doi.org/10.1109/IDSTA55301.2022.9923169
de Amorim, L. B. V., Cavalcanti, G. D. C., & Cruz, R. M. O. (2023). The choice of scaling technique matters for classification performance. Applied Soft Computing, 133, 109924. https://doi.org/10.1016/j.asoc.2022.109924
Demir, S., & Sahin, E. K. (2024). The effectiveness of data pre-processing methods on the performance of machine learning techniques using RF, SVR, Cubist and SGB. Stochastic Environmental Research and Risk Assessment, 38(8), 3273-3290. https://doi.org/10.1007/s00477-024-02745-9
Gan, Q., Gong, L., Hu, D., Jiang, Y., & Ding, X. (2023). A Hybrid Missing Data Imputation Method for Batch Process Monitoring Dataset. Sensors, 23(21), 8678. https://doi.org/10.3390/s23218678
Habib, M., & Okayli, M. (2024). Evaluating the Sensitivity of Machine Learning Models to Data Preprocessing Technique in Concrete Compressive Strength Estimation. Arabian Journal for Science and Engineering, 49(10), 13709-13727. https://doi.org/10.1007/s13369-024-08776-2
Hassanat, A. B., Alqaralleh, M. K., Tarawneh, A. S., Almohammadi, K., Alamri, M., Alzahrani, A., Altarawneh, G. A., & Alhalaseh, R. (2024). A Novel Outlier-Robust Accuracy Measure for Machine Learning Regression Using a Non-Convex Distance Metric. Mathematics, 12(22), 3623. https://doi.org/10.3390/math12223623
Hippel, P. T. von, & Bartlett, J. W. (2021). Maximum Likelihood Multiple Imputation: Faster Imputations and Consistent Standard Errors Without Posterior Draws. Statistical Science, 36(3), 400-420. https://doi.org/10.1214/20-STS793
Kappatou, C. D., Odgers, J., García-Muñoz, S., & Misener, R. (2023). An Optimization Approach Coupling Preprocessing with Model Regression for Enhanced Chemometrics. Industrial & Engineering Chemistry Research. https://doi.org/10.1021/acs.iecr.2c04583
Koukaras, P., & Tjortjis, C. (2025). Data Preprocessing and Feature Engineering for Data Mining: Techniques, Tools, and Best Practices. AI, 6(10), 257. https://doi.org/10.3390/ai6100257
Li, C., Ren, X., & Zhao, G. (2023). Machine-Learning-Based Imputation Method for Filling Missing Values in Ground Meteorological Observation Data. Algorithms, 16(9), 422. https://doi.org/10.3390/a16090422
Li, F., Sun, H., Gu, Y., & Yu, G. (2023). A Noise-Aware Multiple Imputation Algorithm for Missing Data. Mathematics, 11(1), 73. https://doi.org/10.3390/math11010073
Maharana, K., Mondal, S., & Nemade, B. (2022). A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 3(1), 91-99. https://doi.org/10.1016/j.gltp.2022.04.020
Mahmud Sujon, K., Binti Hassan, R., Tusnia Towshi, Z., Othman, M. A., Abdus Samad, M., & Choi, K. (2024). When to Use Standardization and Normalization: Empirical Evidence From Machine Learning Models and XAI. IEEE Access, 12, 135300-135314. https://doi.org/10.1109/ACCESS.2024.3462434
Mallikharjuna Rao, K., Saikrishna, G., & Supriya, K. (2023). Data preprocessing techniques: Emergence and selection towards machine learning models – a practical review using HPA dataset. Multimedia Tools and Applications, 82(24), 37177-37196. https://doi.org/10.1007/s11042-023-15087-5
Mu, W., Cardelli, R., & Ferrari, S. (2026). Data Preprocessing Techniques for Machine Learning Towards Improving Building Energy Performance: A Systematic Review. Energies, 19(6), 1561. https://doi.org/10.3390/en19061561
Ou, H., Yao, Y., & He, Y. (2024). Missing Data Imputation Method Combining Random Forest and Generative Adversarial Imputation Network. Sensors, 24(4), 1112. https://doi.org/10.3390/s24041112
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … Moher, D. (2021a). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
Page, M. J., Moher, D., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … McKenzie, J. E. (2021b). PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ, 372, n160. https://doi.org/10.1136/bmj.n160
Pandit, P., Dey, P., & Krishnamurthy, K. N. (2021). Comparative Assessment of Multiple Linear Regression and Fuzzy Linear Regression Models. SN Computer Science, 2(2), 76. https://doi.org/10.1007/s42979-021-00473-3
Pargent, F., Pfisterer, F., Thomas, J., & Bischl, B. (2022). Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features. Computational Statistics, 37(5), 2671-2692. https://doi.org/10.1007/s00180-022-01207-6
Park, H.-J., Koo, Y.-S., Yang, H.-Y., Han, Y.-S., & Nam, C.-S. (2024). Study on Data Preprocessing for Machine Learning Based on Semiconductor Manufacturing Processes. Sensors, 24(17), 5461. https://doi.org/10.3390/s24175461
Pinheiro, J. M. H., Oliveira, S. V. B. de, Silva, T. H. S., Saraiva, P. A. R., Souza, E. F. de, Godoy, R. V., Ambrosio, L. A., & Becker, M. (2025). The Impact of Feature Scaling in Machine Learning: Effects on Regression and Classification Tasks. IEEE Access, 13, 199903-199931. https://doi.org/10.1109/ACCESS.2025.3635541
Rasyidah, Efendi, R., Nawi, N. Mohd., Deris, M. M., & Burney, S. M. A. (2023). Cleansing of inconsistent sample in linear regression model based on rough sets theory. Systems and Soft Computing, 5, 200046. https://doi.org/10.1016/j.sasc.2022.200046
Rauf, R. I., Alrasheedi, M. A., Sadiq, R., & Aldawsari, A. M. A. (2024). Evaluating Predictive Accuracy of Regression Models with First-Order Autoregressive Disturbances: A Comparative Approach Using Artificial Neural Networks and Classical Estimators. Mathematics, 12(24), 3966. https://doi.org/10.3390/math12243966
Shin, W. G., Lee, J. S., Ju, Y. C., Hwang, H. M., & Ko, S. W. (2025). Data preprocessing and machine learning method based on ameliorated mathematical models for inferring the power generation of photovoltaic system. Energy Conversion and Management, 333, 119793. https://doi.org/10.1016/j.enconman.2025.119793
Tawakuli, A., Havers, B., Gulisano, V., Kaiser, D., & Engel, T. (2025). Survey: Time-series data preprocessing: A survey and an empirical analysis. Journal of Engineering Research, 13(2), 674-711. https://doi.org/10.1016/j.jer.2024.02.018
Veltri, G. A. (2025). The Effects of Data Preprocessing Choices on Behavioral RCT Outcomes: A Multiverse Analysis. Multivariate Behavioral Research, 0(0), 1-16. https://doi.org/10.1080/00273171.2025.2575399
Vescan, A., Găceanu, R., & Şerban, C. (2024). Exploring the impact of data preprocessing techniques on composite classifier algorithms in cross-project defect prediction. Automated Software Engineering, 31(2), 47. https://doi.org/10.1007/s10515-024-00454-9
Wanyonyi, E. N., & Masinde, N. W. (2025). The Impact of Data Preprocessing on Machine Learning Model Performance: A Comprehensive Examination. IJSRCSEIT, 11(2), 3814-3827. https://doi.org/10.32628/CSEIT25112854
Zeng, G., & Tao, S. (2023). A Generalized Linear Transformation and Its Effects on Logistic Regression. Mathematics, 11(2), 467. https://doi.org/10.3390/math11020467
Zheng, D., Hao, X., Khan, M., Wang, L., Li, F., Xiang, N., Kang, F., Hamalainen, T., Cong, F., Song, K., & Qiao, C. (2022). Comparison of machine learning and logistic regression as predictive models for adverse maternal and neonatal outcomes of preeclampsia: A retrospective study. Frontiers in Cardiovascular Medicine, 9. https://doi.org/10.3389/fcvm.2022.959649
Zhong, J., Ma, C., Hou, L., Yin, Y., Zhao, F., Hu, Y., Song, A., Wang, D., Li, L., Cheng, X., & Qiu, L. (2023). Utilization of five data mining algorithms combined with simplified preprocessing to establish reference intervals of thyroid-related hormones for non-elderly adults. BMC Medical Research Methodology, 23(1), 108. https://doi.org/10.1186/s12874-023-01898-5
Descargas
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Jaime David Camacho Castillo, Mauricio Alexander Álvarez Ortega, Iván Alejandro Daqui Cabezas, Alex Nicolas Mazacon Guamingo, Juan Enrique Olivo Quintero, Jordy Joel Vele Carrera

Esta obra está bajo una licencia internacional Creative Commons Atribución 4.0.
Los autores conservan los derechos morales y patrimoniales de sus obras. Puesto que Tesla Revista Científica es una publicación de acceso abierto, los lectores pueden reproducir total o parcialmente su contenido siempre y cuando proporcionen adecuadamente el crédito a los autores correspondientes y a la revista misma. Tesla Revista Científica se compromete a no hacer uso comercial de los textos que recibe y/o publica.
Nuestra revista se rige por las politicas internacionales SHERPA/RoMEO: Revista verde: Permiten el autoarchivo tanto del pre-print (borrador de un trabajo) como del post-print (la versión corregida y revisada por pares) y hasta de la versión final (maquetada tal como saldrá publicada en la revista).
Véase también "Derechos de autor y licencias".



