AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data.

Gigascience

IEETA, Institute of Electronics and Informatics Engineering of Aveiro, and LASI, Intelligent Systems Associate Laboratory, University of Aveiro, Campus Universitário de Santiago, 3810-193 Aveiro, Portugal.

Published: December 2022


Category Ranking

98%

Total Visits

921

Avg Visit Duration

2 minutes

Citations

20

Article Abstract

Background: Low-complexity data analysis is the area that addresses the search and quantification of regions in sequences of elements that contain low-complexity or repetitive elements. For example, these can be tandem repeats, inverted repeats, homopolymer tails, GC-biased regions, similar genes, and hairpins, among many others. Identifying these regions is crucial because of their association with regulatory and structural characteristics. Moreover, their identification provides positional and quantity information where standard assembly methodologies face significant difficulties because of substantial higher depth coverage (mountains), ambiguous read mapping, or where sequencing or reconstruction defects may occur. However, the capability to distinguish low-complexity regions (LCRs) in genomic and proteomic sequences is a challenge that depends on the model's ability to find them automatically. Low-complexity patterns can be implicit through specific or combined sources, such as algorithmic or probabilistic, and recurring to different spatial distances-namely, local, medium, or distant associations.

Findings: This article addresses the challenge of automatically modeling and distinguishing LCRs, providing a new method and tool (AlcoR) for efficient and accurate segmentation and visualization of these regions in genomic and proteomic sequences. The method enables the use of models with different memories, providing the ability to distinguish local from distant low-complexity patterns. The method is reference and alignment free, providing additional methodologies for testing, including a highly flexible simulation method for generating biological sequences (DNA or protein) with different complexity levels, sequence masking, and a visualization tool for automatic computation of the LCR maps into an ideogram style. We provide illustrative demonstrations using synthetic, nearly synthetic, and natural sequences showing the high efficiency and accuracy of AlcoR. As large-scale results, we use AlcoR to unprecedentedly provide a whole-chromosome low-complexity map of a recent complete human genome and the haplotype-resolved chromosome pairs of a heterozygous diploid African cassava cultivar.

Conclusions: The AlcoR method provides the ability of fast sequence characterization through data complexity analysis, ideally for scenarios entangling the presence of new or unknown sequences. AlcoR is implemented in C language using multithreading to increase the computational speed, is flexible for multiple applications, and does not contain external dependencies. The tool accepts any sequence in FASTA format. The source code is freely provided at https://github.com/cobilab/alcor.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC10716826PMC
http://dx.doi.org/10.1093/gigascience/giad101DOI Listing

Publication Analysis

Top Keywords

low-complexity regions
8
genomic proteomic
8
proteomic sequences
8
low-complexity patterns
8
low-complexity
7
alcor
6
regions
6
sequences
6
method
5
alcor alignment-free
4

Similar Publications

Molecular characterization of the medaka (Oryzias latipes) Hsp90ab1.

Fish Shellfish Immunol

September 2025

State Key Laboratory of Mariculture Breeding, Engineering Research Center of the Modern Technology for Eel Industry, Ministry of Education, Key Laboratory of Healthy Mariculture for the East China Sea, Ministry of Agriculture and Rural Affairs, Fisheries College of Jimei University, Xiamen, 361021,

Nervous necrosis virus (NNV) can infect many species of fish and has caused severe economic losses to the aquaculture industry worldwide. Recently, there is a wealth of research indicating that heat shock protein 90ab1 (Hsp90ab1) is a receptor for red-spotted grouper nervous necrosis virus (RGNNV), however the potential function has not been examined so far. In this study, medaka (Oryzias latipes) hsp90ab1 gene (OlHsp90ab1) was identified, which contained 13 exons with a genome length of 5191 bp and coding sequences of 2175 bp.

View Article and Find Full Text PDF

Pathological aggregation of transactive response DNA binding protein of 43 kDa (TDP-43), primarily driven by its low-complexity domain, is closely associated with various neurodegenerative diseases, including amyotrophic lateral sclerosis (ALS) and frontotemporal lobar degeneration (FTLD). Despite the therapeutic potential of preventing TDP-43 aggregation, no effective small molecule or biomacromolecule therapeutics have been successfully developed so far. Here, we introduce a protein design strategy that yields de novo designed proteins capable of stabilizing the key amyloidogenic region of TDP-43 in its native helical conformation with nanomolar binding affinity.

View Article and Find Full Text PDF

Protein activities driven by amino acid composition.

J Biol Chem

August 2025

Department of Biochemistry and Molecular Biology, Colorado State University, Fort Collins, CO 80523, USA. Electronic address:

A foundational concept in protein biochemistry is that the primary sequence of a protein determines its structure and function. However, some protein regions defy this constraint, with protein activity dictated more by amino acid composition than primary sequence. In this review, we examine the concept of "composition-driven protein activities".

View Article and Find Full Text PDF

A Methionine-Rich Repeat Forms a Spiral Conformation That Guides Aragonite Nanofiber Organization in Molluscan Ligaments.

Biomacromolecules

August 2025

Department of Applied Biological Chemistry, Graduate School of Agricultural and Life Sciences, University of Tokyo, 1-1-1 Yayoi, Bunkyo-ku, Tokyo 113-8657, Japan.

The hinge ligament of contains aragonite nanofibers embedded in a dense organic matrix primarily composed of ligament methionine (Met)-rich protein (LMP). LMP features a low-complexity region with 30 repeats of the Met-Met-Met-lysine-proline-aspartic acid (MMMKPD) sequence; however, its structural and functional roles remain unclear. Using synthetic peptides and solution nuclear magnetic resonance with dispersive aragonite particles, we observed that the MMMKPD repeat formed a unique spiral conformation distinct from canonical secondary structures, which was supported by AlphaFold predictions.

View Article and Find Full Text PDF

The enjoyment of music involves a complex interplay between brain perceptual areas and the reward network. While previous studies have shown that musical liking is related to an enhancement of synchronization between the right temporal and frontal brain regions via theta frequency band oscillations, the underlying mechanisms of this interaction remain elusive. Specifically, a causal relationship between theta oscillations and musical pleasure has yet to be shown.

View Article and Find Full Text PDF