Use of data mining for intelligent evaluation of imputation methods
- Autores
- La Red Martinez, David; Primorac, Carlos
- Año de publicación
- 2023
- Idioma
- inglés
- Tipo de recurso
- artículo
- Estado
- versión publicada
- Descripción
- In real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori.
Fil: La Red Martinez, David. Universidad Tecnológica Nacional. Facultad Regional Resistencia; Argentina.
Fil: Primorac, Carlos. Universidad Nacional del Nordeste. Departamento de Informática; Argentina
Peer Reviewed - Materia
-
Computer Science
Data Imputation
Data Mining
Interdisciplinary Applications
Performance Evaluation of Imputation Methods - Nivel de accesibilidad
- acceso abierto
- Condiciones de uso
- 2023-11-21T20:26:38Z
- Repositorio
.jpg)
- Institución
- Universidad Tecnológica Nacional
- OAI Identificador
- oai:ria.utn.edu.ar:20.500.12272/8873
Ver los metadatos del registro completo
| id |
RIAUTN_2c88dd56847bd346ddb99f174044e0c2 |
|---|---|
| oai_identifier_str |
oai:ria.utn.edu.ar:20.500.12272/8873 |
| network_acronym_str |
RIAUTN |
| repository_id_str |
a |
| network_name_str |
Repositorio Institucional Abierto (UTN) |
| spelling |
Use of data mining for intelligent evaluation of imputation methodsLa Red Martinez, DavidPrimorac, CarlosComputer ScienceData ImputationData MiningInterdisciplinary ApplicationsPerformance Evaluation of Imputation MethodsIn real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori.Fil: La Red Martinez, David. Universidad Tecnológica Nacional. Facultad Regional Resistencia; Argentina.Fil: Primorac, Carlos. Universidad Nacional del Nordeste. Departamento de Informática; ArgentinaPeer Reviewed2023-11-21T20:26:38Z2023-11-21T20:26:38Z2023-03-23info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttp://purl.org/coar/resource_type/c_6501info:ar-repo/semantics/articulopdfapplication/pdfhttp://hdl.handle.net/20.500.12272/887310.9781/ijimai.2023.03.002enginfo:eu-repo/semantics/openAccess2023-11-21T20:26:38ZAcceso Abiertoreponame:Repositorio Institucional Abierto (UTN)instname:Universidad Tecnológica Nacional2026-09-24T12:45:33Zoai:ria.utn.edu.ar:20.500.12272/8873instacron:UTNInstitucionalhttp://ria.utn.edu.ar/Universidad públicaNo correspondehttp://ria.utn.edu.ar/oaigestionria@rec.utn.edu.ar; fsuarez@rec.utn.edu.arArgentinaNo correspondeNo correspondeNo correspondeopendoar:a2026-09-24 12:45:34.151Repositorio Institucional Abierto (UTN) - Universidad Tecnológica Nacionalfalse |
| dc.title.none.fl_str_mv |
Use of data mining for intelligent evaluation of imputation methods |
| title |
Use of data mining for intelligent evaluation of imputation methods |
| spellingShingle |
Use of data mining for intelligent evaluation of imputation methods La Red Martinez, David Computer Science Data Imputation Data Mining Interdisciplinary Applications Performance Evaluation of Imputation Methods |
| title_short |
Use of data mining for intelligent evaluation of imputation methods |
| title_full |
Use of data mining for intelligent evaluation of imputation methods |
| title_fullStr |
Use of data mining for intelligent evaluation of imputation methods |
| title_full_unstemmed |
Use of data mining for intelligent evaluation of imputation methods |
| title_sort |
Use of data mining for intelligent evaluation of imputation methods |
| dc.creator.none.fl_str_mv |
La Red Martinez, David Primorac, Carlos |
| author |
La Red Martinez, David |
| author_facet |
La Red Martinez, David Primorac, Carlos |
| author_role |
author |
| author2 |
Primorac, Carlos |
| author2_role |
author |
| dc.subject.none.fl_str_mv |
Computer Science Data Imputation Data Mining Interdisciplinary Applications Performance Evaluation of Imputation Methods |
| topic |
Computer Science Data Imputation Data Mining Interdisciplinary Applications Performance Evaluation of Imputation Methods |
| dc.description.none.fl_txt_mv |
In real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori. Fil: La Red Martinez, David. Universidad Tecnológica Nacional. Facultad Regional Resistencia; Argentina. Fil: Primorac, Carlos. Universidad Nacional del Nordeste. Departamento de Informática; Argentina Peer Reviewed |
| description |
In real-world situations, researchers frequently face the difficulty of missing values (MV), i.e., values not observed in a data set. Data imputation techniques allow the estimation of MV using different algorithms, by means of which important data can be imputed for a particular instance. Most of the literature in this field deals with different imputation methods. However, few studies deal with a comparative evaluation of the different methods as to provide more appropriate guidelines for the selection of the method to be applied to impute data for specific situations. The objective of this work is to show a methodology for evaluating the performance of imputation methods by means of new metrics derived from data mining processes, using quality metrics of data mining models. We started from the complete dataset that was amputated with different amputation mechanisms to generate 63 datasets with MV; these were imputed using Median, k-NN, k-Means and Hot-Deck imputation methods. The performance of the imputation methods was evaluated using new metrics derived from quality metrics of the data mining processes, performed with the original full file and with the imputed files. This evaluation is not based on measuring the error when imputing (usual operation), but on considering the similarity of the values of the quality metrics of the data mining processes obtained with the original file and with the imputed files. The results show that –globally considered and according to the new proposed metric, the imputation methods that showed the best performance were k-NN and k-Means. An additional advantage of the proposed methodology is that it provides predictive data mining models that can be used a posteriori. |
| publishDate |
2023 |
| dc.date.none.fl_str_mv |
2023-11-21T20:26:38Z 2023-11-21T20:26:38Z 2023-03-23 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/publishedVersion http://purl.org/coar/resource_type/c_6501 info:ar-repo/semantics/articulo |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/20.500.12272/8873 10.9781/ijimai.2023.03.002 |
| url |
http://hdl.handle.net/20.500.12272/8873 |
| identifier_str_mv |
10.9781/ijimai.2023.03.002 |
| dc.language.none.fl_str_mv |
eng |
| language |
eng |
| dc.rights.none.fl_str_mv |
info:eu-repo/semantics/openAccess 2023-11-21T20:26:38Z Acceso Abierto |
| eu_rights_str_mv |
openAccess |
| rights_invalid_str_mv |
2023-11-21T20:26:38Z Acceso Abierto |
| dc.format.none.fl_str_mv |
pdf application/pdf |
| dc.source.none.fl_str_mv |
reponame:Repositorio Institucional Abierto (UTN) instname:Universidad Tecnológica Nacional |
| reponame_str |
Repositorio Institucional Abierto (UTN) |
| collection |
Repositorio Institucional Abierto (UTN) |
| instname_str |
Universidad Tecnológica Nacional |
| repository.name.fl_str_mv |
Repositorio Institucional Abierto (UTN) - Universidad Tecnológica Nacional |
| repository.mail.fl_str_mv |
gestionria@rec.utn.edu.ar; fsuarez@rec.utn.edu.ar |
| _version_ |
1877230900817690624 |
| score |
13.265058 |