Filtering useless data at the source

Autores
Pessolani, Pablo; Quaglia, Constanza; Nou, Ramón
Año de publicación
2019
Idioma
inglés
Tipo de recurso
documento de conferencia
Estado
versión publicada
Descripción
There are some processing environments where an application reads remote sequential files with a large number of records only to use some of them. Examples of those environments are servers, proxies, firewall and intrusion detection log analysis tools, sensor log analysis, large scientific datasets processing, etc. To be processed, all file records must be transferred through the network, and all of them must be processed by the application. Some of the transferred records would be discarded immediately by the application because it has no interest in them, but they just consumed network bandwidth and operating system’s cache buffers. This article proposes to filter records from the source of data but without changing the application. Those records of interest will be transferred without modifications but only references to the other records will be transferred from the source to the consuming application. At the application side, the sequence of records is rebuilt, keeping the content of records of interest and filling the others with dummy values which will be discarded by the application. As the number and length of records are preserved (and therefore the file size), it is not necessary to modify the application. Once a filtering rule is applied to a file, only the useful records and references to unuseful ones will be transferred to the application side reducing network usage, transfer time, and cache utilization. A modified (but compatible) version of NFS protocol was developed as a proof of concept.
Fil: Pessolani, Pablo Andrés. Universidad Tecnológica Nacional. Facultad Regional Santa Fe. Departamento de Ingeniería en Sistemas de Información; Argentina.
Fil: Quaglia, Constanza. Universidad Tecnológica Nacional. Facultad Regional Santa Fe. Departamento de Ingeniería en Sistemas de Información; Argentina.
Fil: Nou Castell, Ramón. Universitat Politécnica de Catalunya. Departamento de Arquitectura de Computadores. Barcelona Supercomputing Center; España.
Materia
Logging
Network File System
NFS
Nivel de accesibilidad
acceso abierto
Condiciones de uso
2021-08-03T20:07:56Z
Repositorio
Repositorio Institucional Abierto (UTN)
Institución
Universidad Tecnológica Nacional
OAI Identificador
oai:ria.utn.edu.ar:20.500.12272/5360

id RIAUTN_7f89eff3e03ffd8a76eb0500620c043a
oai_identifier_str oai:ria.utn.edu.ar:20.500.12272/5360
network_acronym_str RIAUTN
repository_id_str a
network_name_str Repositorio Institucional Abierto (UTN)
spelling Filtering useless data at the sourcePessolani, PabloQuaglia, ConstanzaNou, RamónLoggingNetwork File SystemNFSThere are some processing environments where an application reads remote sequential files with a large number of records only to use some of them. Examples of those environments are servers, proxies, firewall and intrusion detection log analysis tools, sensor log analysis, large scientific datasets processing, etc. To be processed, all file records must be transferred through the network, and all of them must be processed by the application. Some of the transferred records would be discarded immediately by the application because it has no interest in them, but they just consumed network bandwidth and operating system’s cache buffers. This article proposes to filter records from the source of data but without changing the application. Those records of interest will be transferred without modifications but only references to the other records will be transferred from the source to the consuming application. At the application side, the sequence of records is rebuilt, keeping the content of records of interest and filling the others with dummy values which will be discarded by the application. As the number and length of records are preserved (and therefore the file size), it is not necessary to modify the application. Once a filtering rule is applied to a file, only the useful records and references to unuseful ones will be transferred to the application side reducing network usage, transfer time, and cache utilization. A modified (but compatible) version of NFS protocol was developed as a proof of concept.Fil: Pessolani, Pablo Andrés. Universidad Tecnológica Nacional. Facultad Regional Santa Fe. Departamento de Ingeniería en Sistemas de Información; Argentina.Fil: Quaglia, Constanza. Universidad Tecnológica Nacional. Facultad Regional Santa Fe. Departamento de Ingeniería en Sistemas de Información; Argentina.Fil: Nou Castell, Ramón. Universitat Politécnica de Catalunya. Departamento de Arquitectura de Computadores. Barcelona Supercomputing Center; España.XXV Congreso Argentino de Ciencias de la Computación2021-08-03T20:07:56Z2021-08-03T20:07:56Z2019-10info:eu-repo/semantics/conferenceObjectinfo:eu-repo/semantics/publishedVersionhttp://purl.org/coar/resource_type/c_5794info:ar-repo/semantics/documentoDeConferenciaapplication/pdfapplication/pdfCongreso Argentino de Ciencias de la Computación, CACIC (25º : 2019 oct. 14-18 : Río Cuarto, Argentina) pp. 879-888http://hdl.handle.net/20.500.12272/5360engengSistema de Virtualización Distribuido. SIUTNFE0005162info:eu-repo/semantics/openAccess2021-08-03T20:07:56Zhttp://creativecommons.org/licenses/by-nc-sa/4.0/Atribución-NoComercial-CompartirIgual 4.0 InternacionalLos autoresCreativeCommonsreponame:Repositorio Institucional Abierto (UTN)instname:Universidad Tecnológica Nacional2026-10-01T11:59:45Zoai:ria.utn.edu.ar:20.500.12272/5360instacron:UTNInstitucionalhttp://ria.utn.edu.ar/Universidad públicaNo correspondehttp://ria.utn.edu.ar/oaigestionria@rec.utn.edu.ar; fsuarez@rec.utn.edu.arArgentinaNo correspondeNo correspondeNo correspondeopendoar:a2026-10-01 11:59:46.074Repositorio Institucional Abierto (UTN) - Universidad Tecnológica Nacionalfalse
dc.title.none.fl_str_mv Filtering useless data at the source
title Filtering useless data at the source
spellingShingle Filtering useless data at the source
Pessolani, Pablo
Logging
Network File System
NFS
title_short Filtering useless data at the source
title_full Filtering useless data at the source
title_fullStr Filtering useless data at the source
title_full_unstemmed Filtering useless data at the source
title_sort Filtering useless data at the source
dc.creator.none.fl_str_mv Pessolani, Pablo
Quaglia, Constanza
Nou, Ramón
author Pessolani, Pablo
author_facet Pessolani, Pablo
Quaglia, Constanza
Nou, Ramón
author_role author
author2 Quaglia, Constanza
Nou, Ramón
author2_role author
author
dc.subject.none.fl_str_mv Logging
Network File System
NFS
topic Logging
Network File System
NFS
dc.description.none.fl_txt_mv There are some processing environments where an application reads remote sequential files with a large number of records only to use some of them. Examples of those environments are servers, proxies, firewall and intrusion detection log analysis tools, sensor log analysis, large scientific datasets processing, etc. To be processed, all file records must be transferred through the network, and all of them must be processed by the application. Some of the transferred records would be discarded immediately by the application because it has no interest in them, but they just consumed network bandwidth and operating system’s cache buffers. This article proposes to filter records from the source of data but without changing the application. Those records of interest will be transferred without modifications but only references to the other records will be transferred from the source to the consuming application. At the application side, the sequence of records is rebuilt, keeping the content of records of interest and filling the others with dummy values which will be discarded by the application. As the number and length of records are preserved (and therefore the file size), it is not necessary to modify the application. Once a filtering rule is applied to a file, only the useful records and references to unuseful ones will be transferred to the application side reducing network usage, transfer time, and cache utilization. A modified (but compatible) version of NFS protocol was developed as a proof of concept.
Fil: Pessolani, Pablo Andrés. Universidad Tecnológica Nacional. Facultad Regional Santa Fe. Departamento de Ingeniería en Sistemas de Información; Argentina.
Fil: Quaglia, Constanza. Universidad Tecnológica Nacional. Facultad Regional Santa Fe. Departamento de Ingeniería en Sistemas de Información; Argentina.
Fil: Nou Castell, Ramón. Universitat Politécnica de Catalunya. Departamento de Arquitectura de Computadores. Barcelona Supercomputing Center; España.
description There are some processing environments where an application reads remote sequential files with a large number of records only to use some of them. Examples of those environments are servers, proxies, firewall and intrusion detection log analysis tools, sensor log analysis, large scientific datasets processing, etc. To be processed, all file records must be transferred through the network, and all of them must be processed by the application. Some of the transferred records would be discarded immediately by the application because it has no interest in them, but they just consumed network bandwidth and operating system’s cache buffers. This article proposes to filter records from the source of data but without changing the application. Those records of interest will be transferred without modifications but only references to the other records will be transferred from the source to the consuming application. At the application side, the sequence of records is rebuilt, keeping the content of records of interest and filling the others with dummy values which will be discarded by the application. As the number and length of records are preserved (and therefore the file size), it is not necessary to modify the application. Once a filtering rule is applied to a file, only the useful records and references to unuseful ones will be transferred to the application side reducing network usage, transfer time, and cache utilization. A modified (but compatible) version of NFS protocol was developed as a proof of concept.
publishDate 2019
dc.date.none.fl_str_mv 2019-10
2021-08-03T20:07:56Z
2021-08-03T20:07:56Z
dc.type.none.fl_str_mv info:eu-repo/semantics/conferenceObject
info:eu-repo/semantics/publishedVersion
http://purl.org/coar/resource_type/c_5794
info:ar-repo/semantics/documentoDeConferencia
format conferenceObject
status_str publishedVersion
dc.identifier.none.fl_str_mv Congreso Argentino de Ciencias de la Computación, CACIC (25º : 2019 oct. 14-18 : Río Cuarto, Argentina) pp. 879-888
http://hdl.handle.net/20.500.12272/5360
identifier_str_mv Congreso Argentino de Ciencias de la Computación, CACIC (25º : 2019 oct. 14-18 : Río Cuarto, Argentina) pp. 879-888
url http://hdl.handle.net/20.500.12272/5360
dc.language.none.fl_str_mv eng
eng
language eng
dc.relation.none.fl_str_mv Sistema de Virtualización Distribuido. SIUTNFE0005162
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
2021-08-03T20:07:56Z
http://creativecommons.org/licenses/by-nc-sa/4.0/
Atribución-NoComercial-CompartirIgual 4.0 Internacional
Los autores
CreativeCommons
eu_rights_str_mv openAccess
rights_invalid_str_mv 2021-08-03T20:07:56Z
http://creativecommons.org/licenses/by-nc-sa/4.0/
Atribución-NoComercial-CompartirIgual 4.0 Internacional
Los autores
CreativeCommons
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv XXV Congreso Argentino de Ciencias de la Computación
publisher.none.fl_str_mv XXV Congreso Argentino de Ciencias de la Computación
dc.source.none.fl_str_mv reponame:Repositorio Institucional Abierto (UTN)
instname:Universidad Tecnológica Nacional
reponame_str Repositorio Institucional Abierto (UTN)
collection Repositorio Institucional Abierto (UTN)
instname_str Universidad Tecnológica Nacional
repository.name.fl_str_mv Repositorio Institucional Abierto (UTN) - Universidad Tecnológica Nacional
repository.mail.fl_str_mv gestionria@rec.utn.edu.ar; fsuarez@rec.utn.edu.ar
_version_ 1877862418679332864
score 13.365483