424 publications from this institution
Using direct access computer files of bibliographic information, an attempt is made to overcome one of the problems often associated with information retrieval, namely, the maintenance and use of large dictionaries, the greater part of which is used only infrequently. A novel method is presented, which maps the hyperbolic frequency distribution of text characteristics onto a rectangular distribution. This is more suited to implementation on storage devices. This method treats text as a string of characters rather than words bounded by spaces, and chooses subsets of strings such that their frequencies of occurrence are more even than those of word types. The members of this subset are then used as index keys for retrieval. The rectangular distribution of key frequencies results in a much simplified file organization and promises considerable cost advantages.
Variety Generation involves the selection of sets of character strings, or symbols, which are intended to occur with equal probabilities in bodies of text or sets of text units from a particular source. It is important that the sample used to generate the symbol set should be representative of the data with which the set will be used. An assessment is given here of the amount of variation in symbol sets generated from files of titles and author names from BNB MARC data over a five year period, and a comparison is made with LC MARC. Some of the BNB symbol sets are compared directly, and equifrequency statistics are obtained for the assignment of each symbol set to each file. The differences between the equifrequency statistics are examined by means of an analysis of variance technique.
The expected response in the mean of a trait subjected to a single generation of selection from an unselected base population is usually well-approximated by the breeder's equation, R = h2S. This chapter shows how this basic equation generalizes to an important number of settings, such as differential selection on the two sexes, overlapping generations, and selection based on measures other than just an individual's phenotype. It also briefly examines selection response on multiple traits.
Evolutionary processes have transformed simple cellular life into a great diversity of forms, ranging from the ubiquitous eukaryotic cell design to the more specific cellular forms of spirochetes, cyanobacteria, ciliates, heliozoans, amoeba, and many others. The cellular traits that constitute these forms require an evolutionary explanation. Ultimately, the persistence of a cellular trait depends on its net contribution to fitness, a quantitative measure. Independent of any positive effects, a cellular trait exhibits a baseline energetic cost that needs to be accounted for when quantitatively examining its net fitness effect. Here, we explore how the energetic burden introduced by a cellular trait quantitatively affects cellular fitness, describe methods for determining cell energy budgets, summarize the costs of cellular traits across the tree of life, and examine how the fitness impacts of these energetic costs compare to other evolutionary forces and trait benefits.
Abstract Catalytic subunits of DNA-dependent RNA polymerases of bacteria, archaea and eukaryotes share hundreds of ultra-conserved amino acids. Remarkably, the plant-specific RNA silencing enzymes, Pol IV and Pol V differ from Pols I, II and III at ~140 of these positions, yet remain capable of RNA synthesis. Whether these amino acid changes in Pols IV and V alter their catalytic properties in comparison to Pol II, from which they evolved, is unknown. Here, we show that Pols IV and V differ from one another, and Pol II, in nucleotide incorporation rate, transcriptional accuracy and the ability to discriminate between ribonucleotides and deoxyribonucleotides. Pol IV transcription is notably error-prone, which may be tolerable, or even beneficial, for biosynthesis of siRNAs targeting transposon families in trans. By contrast, Pol V exhibits high fidelity transcription, suggesting a need for Pol V transcripts to faithfully reflect the DNA sequence of target loci in order to recruit siRNA-Argonaute protein silencing complexes.
Abstract The ability to obtain genome-wide sequences of very large numbers of individuals from natural populations raises questions about optimal sampling designs and the limits to extracting information on key population-genetic parameters from temporal-survey data. Methods are introduced for evaluating whether observed temporal fluctuations in allele frequencies are consistent with the hypothesis of random genetic drift, and expressions for the expected sampling variances for the relevant statistics are given in terms of sample sizes and numbers. Estimation methods and aspects of statistical reliability are also presented for the mean and temporal variance of selection coefficients. For nucleotide sites that pass the test of neutrality, the current effective population size can be estimated by a method of moments, and expressions for its sampling variance provide insight into the degree to which such methodology can yield meaningful results under alternative sampling schemes. Finally, some caveats are raised regarding the use of the temporal covariance of allele-frequency change to infer selection. Taken together, these results provide a statistical view of the limits to population-genetic inference in even the simplest case of a closed population.
ABSTRACT The causes and consequences of spatiotemporal variation in mutation rates remain to be explored in nearly all organisms. Here we examine relationships between local mutation rates and replication timing in three bacterial species whose genomes have multiple chromosomes: Vibrio fischeri , Vibrio cholerae , and Burkholderia cenocepacia . Following five mutation accumulation experiments with these bacteria conducted in the near absence of natural selection, the genomes of clones from each lineage were sequenced and analyzed to identify variation in mutation rates and spectra. In lineages lacking mismatch repair, base substitution mutation rates vary in a mirrored wave-like pattern on opposing replichores of the large chromosomes of V. fischeri and V. cholerae , where concurrently replicated regions experience similar base substitution mutation rates. The base substitution mutation rates on the small chromosome are less variable in both species but occur at similar rates to those in the concurrently replicated regions of the large chromosome. Neither nucleotide composition nor frequency of nucleotide motifs differed among regions experiencing high and low base substitution rates, which along with the inferred ~800-kb wave period suggests that the source of the periodicity is not sequence specific but rather a systematic process related to the cell cycle. These results support the notion that base substitution mutation rates are likely to vary systematically across many bacterial genomes, which exposes certain genes to elevated deleterious mutational load. IMPORTANCE That mutation rates vary within bacterial genomes is well known, but the detailed study of these biases has been made possible only recently with contemporary sequencing methods. We applied these methods to understand how bacterial genomes with multiple chromosomes, like those of Vibrio and Burkholderia , might experience heterogeneous mutation rates because of their unusual replication and the greater genetic diversity found on smaller chromosomes. This study captured thousands of mutations and revealed wave-like rate variation that is synchronized with replication timing and not explained by sequence context. The scale of this rate variation over hundreds of kilobases of DNA strongly suggests that a temporally regulated cellular process may generate wave-like variation in mutation risk. These findings add to our understanding of how mutation risk is distributed across bacterial and likely also eukaryotic genomes, owing to their highly conserved replication and repair machinery.
Abstract Studies of closely related species with known ecological differences provide exceptional opportunities for understanding the genetic mechanisms of evolution. Here, we compared population-genomics data between D. pulex and D. pulicaria , two reproductively compatible sister species experiencing ecological speciation, the first largely confined to intermittent ponds and the second to permanent lakes in the same geographic region. D. pulicaria has lower genome-wide nucleotide diversity, a smaller effective population size, higher incidence of private alleles, and substantially more linkage-disequilibrium than D. pulex . Functional enrichment analysis revealed that positively selected genes in D. pulicaria are enriched in potentially aging-related categories such as cellular homeostasis, which may explain the extended lifespan in D. pulicaria . We also found that opsin-related genes, which may mediate photoperiodic responses, are under different selection pressures in these two species. Additionally, genes involved in mitochondrial functions, ribosomes, and responses to environmental stimuli are found to be under positive selection in both species. Our results provide insights into the physiological traits that differ within this regionally sympatric sister-species pair that occupies unique microhabitats.
The spontaneous deamination of cytosine produces uracil mispaired with guanine in DNA, which will produce a mutation, unless repaired. In all domains of life, uracil-DNA glycosylases (UDGs) are responsible for the elimination of uracil from DNA. Thus, UDGs contribute to the integrity of the genetic information and their loss results in mutator phenotypes. We are interested in understanding the role of UDG genes in the evolutionary variation of the rate and the spectrum of spontaneous mutations. To this end, we determined the presence or absence of the five main UDG families in more than 1,000 completely sequenced genomes and analyzed their patterns of gene loss and gain in eubacterial lineages. We observe nonindependent patterns of gene loss and gain between UDG families in Eubacteria, suggesting extensive functional overlap in an evolutionary timescale. Given that UDGs prevent transitions at G:C sites, we expected the loss of UDG genes to bias the mutational spectrum toward a lower equilibrium G + C content. To test this hypothesis, we used phylogenetically independent contrasts to compare the G + C content at intergenic and 4-fold redundant sites between lineages where UDG genes have been lost and their sister clades. None of the main UDG families present in Eubacteria was associated with a higher G + C content at intergenic or 4-fold redundant sites. We discuss the reasons of this negative result and report several features of the evolution of the UDG superfamily with implications for their functional study. uracil-DNA glycosylase, mutation rate evolution, mutational bias, GC content, DNA repair, mutator gene.
Much attention has been paid to translating isolated chemical names into forms such as connection tables, but less effort has been expended in identifying substance names in running text to make them available for processing. The requirement for automatic name identification becomes a more urgent priority today, not the least in light of the inherent importance of patents and the increasing complexity of newly synthesized substances and, with these, the need for error-free processing of information from patent and other documents. The elaboration of a methodology for isolating substance names in the text of English-language patents is described here, using, in part, the SGML (Standard Generalized Markup Language) of the patent text as an aid to this process. Evaluation of the procedures, which are still at an early stage of development, demonstrates that even simple methods can achieve very high degrees of success.