Impacto de las estrategias de preprocesamiento de datos en la precisión de algoritmos de regresión lineal

Autores/as

  • Jaime David Camacho Castillo
  • Mauricio Alexander Álvarez Ortega https://orcid.org/0009-0005-4850-9425
  • Iván Alejandro Daqui Cabezas
  • Alex Nicolas Mazacon Guamingo
  • Juan Enrique Olivo Quintero
  • Jordy Joel Vele Carrera

DOI:

https://doi.org/10.55204/trc.v6i2.e724

Palabras clave:

preprocesamiento de datos, regresión lineal, normalización, imputación, aprendizaje automático, revisión sistemática

Resumen

El preprocesamiento de datos es una etapa determinante en el rendimiento de los modelos de regresión lineal, donde la calidad de los datos, incluyendo el tratamiento de valores faltantes, la detección de outliers, la normalización y la codificación de variables, influyen directamente en la capacidad predictiva de los algoritmos de aprendizaje automático. Esta revisión sistemática, conducida bajo la metodología PRISMA 2020 (Preferred Reporting Items for Systematic reviews and Meta-Analyses), bajo esta misma óptica, se analiza 40 artículos científicos publicados entre 2021 y 2026 en bases de datos reconocidas como Scopus, Web of Science, Springer Nature, MDPI, IEEE y ScienceDirect. Dentro de este contexto, se destaca que los modelos de regresión lineal constituyen una herramienta fundamental en la inteligencia artificial y el análisis de datos, donde la calidad del preprocesamiento influye directamente en la precisión de los resultados obtenidos. En el desarrollo del estudio, se analizaron técnicas clave como la limpieza de datos, la imputación de valores faltantes, la normalización o estandarización y a esto se suma la codificación de variables categóricas. En cuanto a las aplicaciones prácticas y futuras líneas de investigación, los hallazgos permiten su implementación en modelos predictivos relacionados con ventas, estimación de precios e indicadores académicos, como también, plantear la extensión del análisis hacia modelos no lineales y conjuntos de datos de mayor dimensionalidad. No obstante, un adecuado preprocesamiento de datos mejora significativamente la precisión predictiva de los modelos de regresión lineal, contribuyendo a reducir problemas de overfitting y underfitting.

Descargas

Los datos de descarga aún no están disponibles.

Referencias

Ahsan, M. M., Mahmud, M. A. P., Saha, P. K., Gupta, K. D., & Siddique, Z. (2021). Effect of Data Scaling Methods on Machine Learning Algorithms and Model Performance. Technologies, 9(3), 52. https://doi.org/10.3390/technologies9030052

Alshdaifat, E., Alshdaifat, D., Alsarhan, A., Hussein, F., & El-Salhi, S. M. F. S. (2021). The Effect of Preprocessing Techniques, Applied to Numeric Features, on Classification Algorithms' Performance. Data, 6(2), 11. https://doi.org/10.3390/data6020011

Amini, M., Roozbeh, M., & Mohamed, N. A. (2024). Separation of the Linear and Nonlinear Covariates in the Sparse Semi-Parametric Regression Model in the Presence of Outliers. Mathematics, 12(2), 172. https://doi.org/10.3390/math12020172

Anelli, D., Morano, P., Tajani, F., & Guarini, M. R. (2025). The Interpretative Effects of Normalization Techniques on Complex Regression Modeling. Information, 16(6), 486. https://doi.org/10.3390/info16060486

Avelino, J. G., Cavalcanti, G. D. C., & Cruz, R. M. O. (2024). Resampling strategies for imbalanced regression: A survey and empirical analysis. Artificial Intelligence Review, 57(4), 82. https://doi.org/10.1007/s10462-024-10724-3

Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553

Buczak, P., Chen, J.-J., & Pauly, M. (2023). Analyzing the Effect of Imputation on Classification Performance under MCAR and MAR Missing Mechanisms. Entropy, 25(3), 521. https://doi.org/10.3390/e25030521

Cabello-Solorzano, K., Ortigosa de Araujo, I., Peña, M., Correia, L., & J. Tallón-Ballesteros, A. (2023). The Impact of Data Normalization on the Accuracy of Machine Learning Algorithms: A Comparative Analysis. En 18th International Conference on Soft Computing Models (SOCO 2023) (pp. 344-353). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-42536-3_33

Chen, J., & Yang, Z. (2024). Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms. Applied Sciences, 14(19), 8922. https://doi.org/10.3390/app14198922

Chowdhury, S., Lin, Y., Liaw, B., & Kerby, L. (2022). Evaluation of Tree Based Regression over Multiple Linear Regression for Non-normally Distributed Data in Battery Performance. 2022 International Conference on Intelligent Data Science Technologies and Applications (IDSTA), 17-25. https://doi.org/10.1109/IDSTA55301.2022.9923169

de Amorim, L. B. V., Cavalcanti, G. D. C., & Cruz, R. M. O. (2023). The choice of scaling technique matters for classification performance. Applied Soft Computing, 133, 109924. https://doi.org/10.1016/j.asoc.2022.109924

Demir, S., & Sahin, E. K. (2024). The effectiveness of data pre-processing methods on the performance of machine learning techniques using RF, SVR, Cubist and SGB. Stochastic Environmental Research and Risk Assessment, 38(8), 3273-3290. https://doi.org/10.1007/s00477-024-02745-9

Gan, Q., Gong, L., Hu, D., Jiang, Y., & Ding, X. (2023). A Hybrid Missing Data Imputation Method for Batch Process Monitoring Dataset. Sensors, 23(21), 8678. https://doi.org/10.3390/s23218678

Habib, M., & Okayli, M. (2024). Evaluating the Sensitivity of Machine Learning Models to Data Preprocessing Technique in Concrete Compressive Strength Estimation. Arabian Journal for Science and Engineering, 49(10), 13709-13727. https://doi.org/10.1007/s13369-024-08776-2

Hassanat, A. B., Alqaralleh, M. K., Tarawneh, A. S., Almohammadi, K., Alamri, M., Alzahrani, A., Altarawneh, G. A., & Alhalaseh, R. (2024). A Novel Outlier-Robust Accuracy Measure for Machine Learning Regression Using a Non-Convex Distance Metric. Mathematics, 12(22), 3623. https://doi.org/10.3390/math12223623

Hippel, P. T. von, & Bartlett, J. W. (2021). Maximum Likelihood Multiple Imputation: Faster Imputations and Consistent Standard Errors Without Posterior Draws. Statistical Science, 36(3), 400-420. https://doi.org/10.1214/20-STS793

Kappatou, C. D., Odgers, J., García-Muñoz, S., & Misener, R. (2023). An Optimization Approach Coupling Preprocessing with Model Regression for Enhanced Chemometrics. Industrial & Engineering Chemistry Research. https://doi.org/10.1021/acs.iecr.2c04583

Koukaras, P., & Tjortjis, C. (2025). Data Preprocessing and Feature Engineering for Data Mining: Techniques, Tools, and Best Practices. AI, 6(10), 257. https://doi.org/10.3390/ai6100257

Li, C., Ren, X., & Zhao, G. (2023). Machine-Learning-Based Imputation Method for Filling Missing Values in Ground Meteorological Observation Data. Algorithms, 16(9), 422. https://doi.org/10.3390/a16090422

Li, F., Sun, H., Gu, Y., & Yu, G. (2023). A Noise-Aware Multiple Imputation Algorithm for Missing Data. Mathematics, 11(1), 73. https://doi.org/10.3390/math11010073

Maharana, K., Mondal, S., & Nemade, B. (2022). A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 3(1), 91-99. https://doi.org/10.1016/j.gltp.2022.04.020

Mahmud Sujon, K., Binti Hassan, R., Tusnia Towshi, Z., Othman, M. A., Abdus Samad, M., & Choi, K. (2024). When to Use Standardization and Normalization: Empirical Evidence From Machine Learning Models and XAI. IEEE Access, 12, 135300-135314. https://doi.org/10.1109/ACCESS.2024.3462434

Mallikharjuna Rao, K., Saikrishna, G., & Supriya, K. (2023). Data preprocessing techniques: Emergence and selection towards machine learning models – a practical review using HPA dataset. Multimedia Tools and Applications, 82(24), 37177-37196. https://doi.org/10.1007/s11042-023-15087-5

Mu, W., Cardelli, R., & Ferrari, S. (2026). Data Preprocessing Techniques for Machine Learning Towards Improving Building Energy Performance: A Systematic Review. Energies, 19(6), 1561. https://doi.org/10.3390/en19061561

Ou, H., Yao, Y., & He, Y. (2024). Missing Data Imputation Method Combining Random Forest and Generative Adversarial Imputation Network. Sensors, 24(4), 1112. https://doi.org/10.3390/s24041112

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … Moher, D. (2021a). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71

Page, M. J., Moher, D., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … McKenzie, J. E. (2021b). PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ, 372, n160. https://doi.org/10.1136/bmj.n160

Pandit, P., Dey, P., & Krishnamurthy, K. N. (2021). Comparative Assessment of Multiple Linear Regression and Fuzzy Linear Regression Models. SN Computer Science, 2(2), 76. https://doi.org/10.1007/s42979-021-00473-3

Pargent, F., Pfisterer, F., Thomas, J., & Bischl, B. (2022). Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features. Computational Statistics, 37(5), 2671-2692. https://doi.org/10.1007/s00180-022-01207-6

Park, H.-J., Koo, Y.-S., Yang, H.-Y., Han, Y.-S., & Nam, C.-S. (2024). Study on Data Preprocessing for Machine Learning Based on Semiconductor Manufacturing Processes. Sensors, 24(17), 5461. https://doi.org/10.3390/s24175461

Pinheiro, J. M. H., Oliveira, S. V. B. de, Silva, T. H. S., Saraiva, P. A. R., Souza, E. F. de, Godoy, R. V., Ambrosio, L. A., & Becker, M. (2025). The Impact of Feature Scaling in Machine Learning: Effects on Regression and Classification Tasks. IEEE Access, 13, 199903-199931. https://doi.org/10.1109/ACCESS.2025.3635541

Rasyidah, Efendi, R., Nawi, N. Mohd., Deris, M. M., & Burney, S. M. A. (2023). Cleansing of inconsistent sample in linear regression model based on rough sets theory. Systems and Soft Computing, 5, 200046. https://doi.org/10.1016/j.sasc.2022.200046

Rauf, R. I., Alrasheedi, M. A., Sadiq, R., & Aldawsari, A. M. A. (2024). Evaluating Predictive Accuracy of Regression Models with First-Order Autoregressive Disturbances: A Comparative Approach Using Artificial Neural Networks and Classical Estimators. Mathematics, 12(24), 3966. https://doi.org/10.3390/math12243966

Shin, W. G., Lee, J. S., Ju, Y. C., Hwang, H. M., & Ko, S. W. (2025). Data preprocessing and machine learning method based on ameliorated mathematical models for inferring the power generation of photovoltaic system. Energy Conversion and Management, 333, 119793. https://doi.org/10.1016/j.enconman.2025.119793

Tawakuli, A., Havers, B., Gulisano, V., Kaiser, D., & Engel, T. (2025). Survey: Time-series data preprocessing: A survey and an empirical analysis. Journal of Engineering Research, 13(2), 674-711. https://doi.org/10.1016/j.jer.2024.02.018

Veltri, G. A. (2025). The Effects of Data Preprocessing Choices on Behavioral RCT Outcomes: A Multiverse Analysis. Multivariate Behavioral Research, 0(0), 1-16. https://doi.org/10.1080/00273171.2025.2575399

Vescan, A., Găceanu, R., & Şerban, C. (2024). Exploring the impact of data preprocessing techniques on composite classifier algorithms in cross-project defect prediction. Automated Software Engineering, 31(2), 47. https://doi.org/10.1007/s10515-024-00454-9

Wanyonyi, E. N., & Masinde, N. W. (2025). The Impact of Data Preprocessing on Machine Learning Model Performance: A Comprehensive Examination. IJSRCSEIT, 11(2), 3814-3827. https://doi.org/10.32628/CSEIT25112854

Zeng, G., & Tao, S. (2023). A Generalized Linear Transformation and Its Effects on Logistic Regression. Mathematics, 11(2), 467. https://doi.org/10.3390/math11020467

Zheng, D., Hao, X., Khan, M., Wang, L., Li, F., Xiang, N., Kang, F., Hamalainen, T., Cong, F., Song, K., & Qiao, C. (2022). Comparison of machine learning and logistic regression as predictive models for adverse maternal and neonatal outcomes of preeclampsia: A retrospective study. Frontiers in Cardiovascular Medicine, 9. https://doi.org/10.3389/fcvm.2022.959649

Zhong, J., Ma, C., Hou, L., Yin, Y., Zhao, F., Hu, Y., Song, A., Wang, D., Li, L., Cheng, X., & Qiu, L. (2023). Utilization of five data mining algorithms combined with simplified preprocessing to establish reference intervals of thyroid-related hormones for non-elderly adults. BMC Medical Research Methodology, 23(1), 108. https://doi.org/10.1186/s12874-023-01898-5

Descargas

Publicado

2026-09-08

Número

Sección

Artículos de Investigación Original

Cómo citar

Camacho Castillo, J. D., Álvarez Ortega, M. A., Daqui Cabezas, I. A., Mazacon Guamingo, A. N., Olivo Quintero, J. E., & Vele Carrera, J. J. (2026). Impacto de las estrategias de preprocesamiento de datos en la precisión de algoritmos de regresión lineal. Tesla Revista Científica, 6(2), e724. https://doi.org/10.55204/trc.v6i2.e724

Artículos más leídos del mismo autor/a