Probabilistic models for CRISPR spacer content evolutionReportar como inadecuado




Probabilistic models for CRISPR spacer content evolution - Descarga este documento en PDF. Documentación en PDF para descargar gratis. Disponible también para leer online.

BMC Evolutionary Biology

, 13:54

Theories and models

Abstract

BackgroundThe CRISPR-Cas system is known to act as an adaptive and heritable immune system in Eubacteria and Archaea. Immunity is encoded in an array of spacer sequences. Each spacer can provide specific immunity to invasive elements that carry the same or a similar sequence. Even in closely related strains, spacer content is very dynamic and evolves quickly. Standard models of nucleotide evolution cannot be applied to quantify its rate of change since processes other than single nucleotide changes determine its evolution.

MethodsWe present probabilistic models that are specific for spacer content evolution. They account for the different processes of insertion and deletion. Insertions can be constrained to occur on one end only or are allowed to occur throughout the array. One deletion event can affect one spacer or a whole fragment of adjacent spacers. Parameters of the underlying models are estimated for a pair of arrays by maximum likelihood using explicit ancestor enumeration.

ResultsSimulations show that parameters are well estimated on average under the models presented here. There is a bias in the rate estimation when including fragment deletions. The models also estimate times between pairs of strains. But with increasing time, spacer overlap goes to zero, and thus there is an upper bound on the distance that can be estimated. Spacer content similarities are displayed in a distance based phylogeny using the estimated times.

We use the presented models to analyze different Yersinia pestis data sets and find that the results among them are largely congruent. The models also capture the variation in diversity of spacers among the data sets. A comparison of spacer-based phylogenies and Cas gene phylogenies shows that they resolve very different time scales for this data set.

ConclusionsThe simulations and data analyses show that the presented models are useful for quantifying spacer content evolution and for displaying spacer content similarities of closely related strains in a phylogeny. This allows for comparisons of different CRISPR arrays or for comparisons between CRISPR arrays and nucleotide substitution rates.

KeywordsCRISPR-Cas Maximum Likelihood Microbial genome evolution Bacterial immunity Electronic supplementary materialThe online version of this article doi:10.1186-1471-2148-13-54 contains supplementary material, which is available to authorized users.

Download fulltext PDF



Autor: Anne Kupczok - Jonathan P Bollback

Fuente: https://link.springer.com/







Documentos relacionados