Use of data mining for intelligent evaluation of imputation methods

Autores
La Red Martinez, David; Primorac, Carlos
Año de publicación
2023
Idioma
inglés
Tipo de recurso
artículo
Estado
versión publicada
Descripción
In real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori.
Fil: La Red Martinez, David. Universidad Tecnológica Nacional. Facultad Regional Resistencia; Argentina.
Fil: Primorac, Carlos. Universidad Nacional del Nordeste. Departamento de Informática; Argentina
Peer Reviewed
Materia
Computer Science
Data Imputation
Data Mining
Interdisciplinary Applications
Performance Evaluation of Imputation Methods
Nivel de accesibilidad
acceso abierto
Condiciones de uso
2023-11-21T20:26:38Z
Repositorio
Repositorio Institucional Abierto (UTN)
Institución
Universidad Tecnológica Nacional
OAI Identificador
oai:ria.utn.edu.ar:20.500.12272/8873

id RIAUTN_2c88dd56847bd346ddb99f174044e0c2
oai_identifier_str oai:ria.utn.edu.ar:20.500.12272/8873
network_acronym_str RIAUTN
repository_id_str a
network_name_str Repositorio Institucional Abierto (UTN)
spelling Use of data mining for intelligent evaluation of imputation methodsLa Red Martinez, DavidPrimorac, CarlosComputer ScienceData ImputationData MiningInterdisciplinary ApplicationsPerformance Evaluation of Imputation MethodsIn real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori.Fil: La Red Martinez, David. Universidad Tecnológica Nacional. Facultad Regional Resistencia; Argentina.Fil: Primorac, Carlos. Universidad Nacional del Nordeste. Departamento de Informática; ArgentinaPeer Reviewed2023-11-21T20:26:38Z2023-11-21T20:26:38Z2023-03-23info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttp://purl.org/coar/resource_type/c_6501info:ar-repo/semantics/articulopdfapplication/pdfhttp://hdl.handle.net/20.500.12272/887310.9781/ijimai.2023.03.002enginfo:eu-repo/semantics/openAccess2023-11-21T20:26:38ZAcceso Abiertoreponame:Repositorio Institucional Abierto (UTN)instname:Universidad Tecnológica Nacional2026-09-24T12:45:33Zoai:ria.utn.edu.ar:20.500.12272/8873instacron:UTNInstitucionalhttp://ria.utn.edu.ar/Universidad públicaNo correspondehttp://ria.utn.edu.ar/oaigestionria@rec.utn.edu.ar; fsuarez@rec.utn.edu.arArgentinaNo correspondeNo correspondeNo correspondeopendoar:a2026-09-24 12:45:34.151Repositorio Institucional Abierto (UTN) - Universidad Tecnológica Nacionalfalse
dc.title.none.fl_str_mv Use of data mining for intelligent evaluation of imputation methods
title Use of data mining for intelligent evaluation of imputation methods
spellingShingle Use of data mining for intelligent evaluation of imputation methods
La Red Martinez, David
Computer Science
Data Imputation
Data Mining
Interdisciplinary Applications
Performance Evaluation of Imputation Methods
title_short Use of data mining for intelligent evaluation of imputation methods
title_full Use of data mining for intelligent evaluation of imputation methods
title_fullStr Use of data mining for intelligent evaluation of imputation methods
title_full_unstemmed Use of data mining for intelligent evaluation of imputation methods
title_sort Use of data mining for intelligent evaluation of imputation methods
dc.creator.none.fl_str_mv La Red Martinez, David
Primorac, Carlos
author La Red Martinez, David
author_facet La Red Martinez, David
Primorac, Carlos
author_role author
author2 Primorac, Carlos
author2_role author
dc.subject.none.fl_str_mv Computer Science
Data Imputation
Data Mining
Interdisciplinary Applications
Performance Evaluation of Imputation Methods
topic Computer Science
Data Imputation
Data Mining
Interdisciplinary Applications
Performance Evaluation of Imputation Methods
dc.description.none.fl_txt_mv In real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori.
Fil: La Red Martinez, David. Universidad Tecnológica Nacional. Facultad Regional Resistencia; Argentina.
Fil: Primorac, Carlos. Universidad Nacional del Nordeste. Departamento de Informática; Argentina
Peer Reviewed
description In real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori.
publishDate 2023
dc.date.none.fl_str_mv 2023-11-21T20:26:38Z
2023-11-21T20:26:38Z
2023-03-23
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
http://purl.org/coar/resource_type/c_6501
info:ar-repo/semantics/articulo
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/20.500.12272/8873
10.9781/ijimai.2023.03.002
url http://hdl.handle.net/20.500.12272/8873
identifier_str_mv 10.9781/ijimai.2023.03.002
dc.language.none.fl_str_mv eng
language eng
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
2023-11-21T20:26:38Z
Acceso Abierto
eu_rights_str_mv openAccess
rights_invalid_str_mv 2023-11-21T20:26:38Z
Acceso Abierto
dc.format.none.fl_str_mv pdf
application/pdf
dc.source.none.fl_str_mv reponame:Repositorio Institucional Abierto (UTN)
instname:Universidad Tecnológica Nacional
reponame_str Repositorio Institucional Abierto (UTN)
collection Repositorio Institucional Abierto (UTN)
instname_str Universidad Tecnológica Nacional
repository.name.fl_str_mv Repositorio Institucional Abierto (UTN) - Universidad Tecnológica Nacional
repository.mail.fl_str_mv gestionria@rec.utn.edu.ar; fsuarez@rec.utn.edu.ar
_version_ 1877230900817690624
score 13.265058