Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues

Autores
Amalfitano, Agustin; Stocchi, Nicolas; Atencio, Hugo Marcelo; Villarreal, Fernando Daniel; Ten Have, Arjen
Año de publicación
2024
Idioma
inglés
Tipo de recurso
artículo
Estado
versión publicada
Descripción
Seqrutinator is an objective, flexible pipeline that removes sequences with sequencing and/or gene model errors and sequences from pseudogenes from complex, eukaryotic protein superfamilies. Testing Seqrutinator on major superfamilies BAHD, CYP, and UGT removes only 1.94% of SwissProt entries, 14% of entries from the model plant Arabidopsis thaliana, but 80% of entries from Pinus taeda’s recent complete proteome. Application of Seqrutinator on crude BAHDomes, CYPomes, and UGTomes obtained from 16 plant proteomes shows convergence of the numbers of paralogues. MSAs, phylogenies, and particularly functional clustering improve drastically upon Seqrutinator application, indicating good performance.
EEA Balcarce
Fil: Amalfitano, Agustín. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Científicas y Tecnológicas en Electrónica (ICYTE); Argentina
Fil: Amalfitano, Agustín. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Científicas y Tecnológicas en Electrónica (ICYTE); Argentina
Fil: Stocchi, Nicolas. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Stocchi, Nicolas. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Atencio, Hugo Marcelo. Instituto Nacional de Tecnología Agropecuaria (INTA). Estación Experimental Agropecuaria Balcarce; Argentina
Fil: Villarreal, Fernando Daniel. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Villarreal, Fernando Daniel. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Ten Have, Arjen. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Fil: Ten Have, Arjen. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fuente
Genome Biology 25 : 230 (agosto 2024)
Materia
Filogenética
Pseudogenes
Proteínas
Modelado de Homología
Phylogenetics
Proteins
Homology Modelling
Seqrutinator
Alineamiento Múltiple de Secuencias
Modelo Génico
Multiple Sequence Alignment
Gene Model
Nivel de accesibilidad
acceso abierto
Condiciones de uso
http://creativecommons.org/licenses/by-nc-sa/4.0/
Repositorio
INTA Digital (INTA)
Institución
Instituto Nacional de Tecnología Agropecuaria
OAI Identificador
oai:localhost:20.500.12123/27903

id INTADig_766df9584e699776f85e4de15227d584
oai_identifier_str oai:localhost:20.500.12123/27903
network_acronym_str INTADig
repository_id_str l
network_name_str INTA Digital (INTA)
spelling Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologuesAmalfitano, AgustinStocchi, NicolasAtencio, Hugo MarceloVillarreal, Fernando DanielTen Have, ArjenFilogenéticaPseudogenesProteínasModelado de HomologíaPhylogeneticsProteinsHomology ModellingSeqrutinatorAlineamiento Múltiple de SecuenciasModelo GénicoMultiple Sequence AlignmentGene ModelSeqrutinator is an objective, flexible pipeline that removes sequences with sequencing and/or gene model errors and sequences from pseudogenes from complex, eukaryotic protein superfamilies. Testing Seqrutinator on major superfamilies BAHD, CYP, and UGT removes only 1.94% of SwissProt entries, 14% of entries from the model plant Arabidopsis thaliana, but 80% of entries from Pinus taeda’s recent complete proteome. Application of Seqrutinator on crude BAHDomes, CYPomes, and UGTomes obtained from 16 plant proteomes shows convergence of the numbers of paralogues. MSAs, phylogenies, and particularly functional clustering improve drastically upon Seqrutinator application, indicating good performance.EEA BalcarceFil: Amalfitano, Agustín. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Científicas y Tecnológicas en Electrónica (ICYTE); ArgentinaFil: Amalfitano, Agustín. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Científicas y Tecnológicas en Electrónica (ICYTE); ArgentinaFil: Stocchi, Nicolas. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; ArgentinaFil: Stocchi, Nicolas. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; ArgentinaFil: Atencio, Hugo Marcelo. Instituto Nacional de Tecnología Agropecuaria (INTA). Estación Experimental Agropecuaria Balcarce; ArgentinaFil: Villarreal, Fernando Daniel. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; ArgentinaFil: Villarreal, Fernando Daniel. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; ArgentinaFil: Ten Have, Arjen. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; ArgentinaFil: Fil: Ten Have, Arjen. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; ArgentinaBioMed Central2026-09-23T12:36:22Z2026-09-23T12:36:22Z2024-08info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttp://purl.org/coar/resource_type/c_6501info:ar-repo/semantics/articuloapplication/pdfhttp://hdl.handle.net/20.500.12123/27903https://link.springer.com/article/10.1186/s13059-024-03371-y1474-760Xhttps://doi.org/10.1186/s13059-024-03371-yGenome Biology 25 : 230 (agosto 2024)reponame:INTA Digital (INTA)instname:Instituto Nacional de Tecnología Agropecuariaenginfo:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by-nc-sa/4.0/Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)2026-09-24T11:43:52Zoai:localhost:20.500.12123/27903instacron:INTAInstitucionalhttp://repositorio.inta.gob.ar/Organismo científico-tecnológicoNo correspondehttp://repositorio.inta.gob.ar/oai/requesttripaldi.nicolas@inta.gob.arArgentinaNo correspondeNo correspondeNo correspondeopendoar:l2026-09-24 11:43:52.941INTA Digital (INTA) - Instituto Nacional de Tecnología Agropecuariafalse
dc.title.none.fl_str_mv Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
title Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
spellingShingle Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
Amalfitano, Agustin
Filogenética
Pseudogenes
Proteínas
Modelado de Homología
Phylogenetics
Proteins
Homology Modelling
Seqrutinator
Alineamiento Múltiple de Secuencias
Modelo Génico
Multiple Sequence Alignment
Gene Model
title_short Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
title_full Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
title_fullStr Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
title_full_unstemmed Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
title_sort Seqrutinator: scrutiny of large protein superfamily sequence datasets for the identification and elimination of non‑functional homologues
dc.creator.none.fl_str_mv Amalfitano, Agustin
Stocchi, Nicolas
Atencio, Hugo Marcelo
Villarreal, Fernando Daniel
Ten Have, Arjen
author Amalfitano, Agustin
author_facet Amalfitano, Agustin
Stocchi, Nicolas
Atencio, Hugo Marcelo
Villarreal, Fernando Daniel
Ten Have, Arjen
author_role author
author2 Stocchi, Nicolas
Atencio, Hugo Marcelo
Villarreal, Fernando Daniel
Ten Have, Arjen
author2_role author
author
author
author
dc.subject.none.fl_str_mv Filogenética
Pseudogenes
Proteínas
Modelado de Homología
Phylogenetics
Proteins
Homology Modelling
Seqrutinator
Alineamiento Múltiple de Secuencias
Modelo Génico
Multiple Sequence Alignment
Gene Model
topic Filogenética
Pseudogenes
Proteínas
Modelado de Homología
Phylogenetics
Proteins
Homology Modelling
Seqrutinator
Alineamiento Múltiple de Secuencias
Modelo Génico
Multiple Sequence Alignment
Gene Model
dc.description.none.fl_txt_mv Seqrutinator is an objective, flexible pipeline that removes sequences with sequencing and/or gene model errors and sequences from pseudogenes from complex, eukaryotic protein superfamilies. Testing Seqrutinator on major superfamilies BAHD, CYP, and UGT removes only 1.94% of SwissProt entries, 14% of entries from the model plant Arabidopsis thaliana, but 80% of entries from Pinus taeda’s recent complete proteome. Application of Seqrutinator on crude BAHDomes, CYPomes, and UGTomes obtained from 16 plant proteomes shows convergence of the numbers of paralogues. MSAs, phylogenies, and particularly functional clustering improve drastically upon Seqrutinator application, indicating good performance.
EEA Balcarce
Fil: Amalfitano, Agustín. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Científicas y Tecnológicas en Electrónica (ICYTE); Argentina
Fil: Amalfitano, Agustín. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Científicas y Tecnológicas en Electrónica (ICYTE); Argentina
Fil: Stocchi, Nicolas. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Stocchi, Nicolas. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Atencio, Hugo Marcelo. Instituto Nacional de Tecnología Agropecuaria (INTA). Estación Experimental Agropecuaria Balcarce; Argentina
Fil: Villarreal, Fernando Daniel. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Villarreal, Fernando Daniel. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Ten Have, Arjen. Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
Fil: Fil: Ten Have, Arjen. Universidad Nacional de Mar del Plata (UNMdP). Instituto de Investigaciones Biológicas (IIB). Biología Computacional y Genómica Comparativa; Argentina
description Seqrutinator is an objective, flexible pipeline that removes sequences with sequencing and/or gene model errors and sequences from pseudogenes from complex, eukaryotic protein superfamilies. Testing Seqrutinator on major superfamilies BAHD, CYP, and UGT removes only 1.94% of SwissProt entries, 14% of entries from the model plant Arabidopsis thaliana, but 80% of entries from Pinus taeda’s recent complete proteome. Application of Seqrutinator on crude BAHDomes, CYPomes, and UGTomes obtained from 16 plant proteomes shows convergence of the numbers of paralogues. MSAs, phylogenies, and particularly functional clustering improve drastically upon Seqrutinator application, indicating good performance.
publishDate 2024
dc.date.none.fl_str_mv 2024-08
2026-09-23T12:36:22Z
2026-09-23T12:36:22Z
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
http://purl.org/coar/resource_type/c_6501
info:ar-repo/semantics/articulo
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/20.500.12123/27903
https://link.springer.com/article/10.1186/s13059-024-03371-y
1474-760X
https://doi.org/10.1186/s13059-024-03371-y
url http://hdl.handle.net/20.500.12123/27903
https://link.springer.com/article/10.1186/s13059-024-03371-y
https://doi.org/10.1186/s13059-024-03371-y
identifier_str_mv 1474-760X
dc.language.none.fl_str_mv eng
language eng
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
http://creativecommons.org/licenses/by-nc-sa/4.0/
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)
eu_rights_str_mv openAccess
rights_invalid_str_mv http://creativecommons.org/licenses/by-nc-sa/4.0/
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv BioMed Central
publisher.none.fl_str_mv BioMed Central
dc.source.none.fl_str_mv Genome Biology 25 : 230 (agosto 2024)
reponame:INTA Digital (INTA)
instname:Instituto Nacional de Tecnología Agropecuaria
reponame_str INTA Digital (INTA)
collection INTA Digital (INTA)
instname_str Instituto Nacional de Tecnología Agropecuaria
repository.name.fl_str_mv INTA Digital (INTA) - Instituto Nacional de Tecnología Agropecuaria
repository.mail.fl_str_mv tripaldi.nicolas@inta.gob.ar
_version_ 1877227359305728000
score 13.265058