424 publications from this institution
Understanding the evolution of cellular features requires a catalog of costs of building, maintaining, and operating cell parts. Lane and Martin (1) define the cost of a gene as the ratio of a cell’s metabolic rate and total gene number. This is an ecologically and evolutionarily meaningless definition, revealing nothing about the incorporation of biomass into offspring and failing to account for differences in generation lengths among organisms. We directly derive the DNA-, mRNA-, and protein-level costs of a gene, dividing these by a cell’s lifetime energy expenditure and accounting for cell division time (2). Despite Lane and Martin’s (3) claim that their paper was not about the bioenergetic costs of a gene, they stated, “By ‘energy per gene’, we mean the cost of expressing the gene” (1).
Abstract Genotype calling plays important roles in population-genomic studies, which have been greatly accelerated by sequencing technologies. To take full advantage of the resultant information, we have developed maximum-likelihood (ML) methods for calling genotypes from high-throughput sequencing data. As the statistical uncertainties associated with sequencing data depend on depths of coverage, we have developed two types of genotype callers. One approach is appropriate for low-coverage sequencing data, and incorporates population-level information on genotype frequencies and error rates pre-estimated by an ML method. Performance evaluation using computer simulations and human data shows that the proposed framework yields less biased estimates of allele frequencies and more accurate genotype calls than current widely used methods. Another type of genotype caller applies to high-coverage sequencing data, requires no prior genotype-frequency estimates, and makes no assumption on the number of alleles at a polymorphic site. Using computer simulations, we determine the depth of coverage necessary to accurately characterize polymorphisms using this second method. We applied the proposed method to high-coverage (mean 18×) sequencing data of 83 clones from a population of Daphnia pulex. The results show that the proposed method enables conservative and reasonably powerful detection of polymorphisms with arbitrary numbers of alleles. We have extended the proposed method to the analysis of genomic data for polyploid organisms, showing that calling accurate polyploid genotypes requires much higher coverage than diploid genotypes.
The mutation process ultimately defines the genetic features of all populations and, hence, has a bearing on a wide range of issues involving evolutionary genetics, inheritance, and genetic disorders, including the predisposition to cancer. Nevertheless, formidable technical barriers have constrained our understanding of the rate at which mutations arise and the molecular spectrum of their effects. Here, we report on the use of complete-genome sequencing in the characterization of spontaneously arising mutations in the yeast Saccharomyces cerevisiae . Our results confirm some findings previously obtained by indirect methods but also yield numerous unexpected findings, in particular a very high rate of point mutation and skewed distribution of base-substitution types in the mitochondrion, a very high rate of segmental duplication and deletion in the nuclear genome, and substantial deviations in the mutational profile among various model organisms.
Abstract Rapidly improving high-throughput sequencing technologies provide unprecedented opportunities for carrying out population-genomic studies with various organisms. To take full advantage of these methods, it is essential to correctly estimate allele and genotype frequencies, and here we present a maximum-likelihood method that accomplishes these tasks. The proposed method fully accounts for uncertainties resulting from sequencing errors and biparental chromosome sampling and yields essentially unbiased estimates with minimal sampling variances with moderately high depths of coverage regardless of a mating system and structure of the population. Moreover, we have developed statistical tests for examining the significance of polymorphisms and their genotypic deviations from Hardy–Weinberg equilibrium. We examine the performance of the proposed method by computer simulations and apply it to low-coverage human data generated by high-throughput sequencing. The results show that the proposed method improves our ability to carry out population-genomic analyses in important ways. The software package of the proposed method is freely available from https://github.com/Takahiro-Maruki/Package-GFE.
Mildly deleterious mutation has been invoked as a leading explanation for a diverse array of observations in evolutionary genetics and molecular evolution and is thought to be a significant risk of extinction for small populations.However, much of the empirical evidence for the deleterious-mutation process derives from studies of Drosophila melanogaster, some of which have been called into question.We review a broad array of data that collectively support the hypothesis that deleterious mutations arise in flies at rate of about one per individual per generation, with the average mutation decreasing fitness by about only 2% in the heterozygous state.Empirical evidence from microbes, plants, and several other animal species provide further support for the idea that most mutations have only mildly deleterious effects on fitness, and several other species appear to have genomic mutation rates that are of the order of magnitude observed in Drosophila.However, there is mounting evidence that some organisms have genomic deleterious mutation rates that are substantially lower than one per individual per generation.These lower rates may be at least partially reconciled with the Drosophila data by taking into consideration the number of germline cell divisions per generation.To fully resolve the existing controversy over the properties of spontaneous mutations, a number of issues need to be clarified.These include the form of the distribution of mutational effects and the extent to which this is modified by the environmental and genetic background and the contribution of basic biological features such as generation length and genome size to interspecific differences in the genomic mutation rate.Once such information is available, it should be possible to make a refined statement about the long-term impact of mutation on the genetic integrity of human populations subject to relaxed selection resulting from modern medical procedures.
1 Abstract Although various empirical studies have reported a positive correlation between the specific growth rate and cell size across bacteria, it is currently unclear what causes this relationship. We conjecture that such scaling occurs because smaller cells have a larger surface-to-volume ratio and thus have to allocate a greater fraction of the total resources to the production of the cell envelope, leaving fewer resources for other biosynthetic processes. To test this theory, we developed a coarse-grained model of bacterial physiology composed of the proteome that converts nutrients into biomass, with the cell envelope acting as a resource sink. Assuming resources are partitioned to maximize the growth rate, the model yields expected scalings. Namely, the growth rate and ribosomal mass fraction scale negatively, while the mass fraction of envelope-producing enzymes scales positively with surface-to-volume. These relationships are compatible with growth measurements and quantitative proteomics data reported in the literature.
Abstract Histone proteins and the nucleosomal organization of chromatin are near-universal eukaroytic features, with the exception of dinoflagellates. Previous studies have suggested that histones do not play a major role in the packaging of dinoflagellate genomes, although several genomic and transcriptomic surveys have detected a full set of core histone genes. Here, transcriptomic and genomic sequence data from multiple dinoflagellate lineages are analyzed, and the diversity of histone proteins and their variants characterized, with particular focus on their potential post-translational modifications and the conservation of the histone code. In addition, the set of putative epigenetic mark readers and writers, chromatin remodelers and histone chaperones are examined. Dinoflagellates clearly express the most derived set of histones among all autonomous eukaryote nuclei, consistent with a combination of relaxation of sequence constraints imposed by the histone code and the presence of numerous specialized histone variants. The histone code itself appears to have diverged significantly in some of its components, yet others are conserved, implying conservation of the associated biochemical processes. Specifically, and with major implications for the function of histones in dinoflagellates, the results presented here strongly suggest that transcription through nucleosomal arrays happens in dinoflagellates. Finally, the plausible roles of histones in dinoflagellate nuclei are discussed.
Abstract Newly emerging data from genome sequencing projects suggest that gene duplication, often accompanied by genetic map changes, is a common and ongoing feature of all genomes. This raises the possibility that differential expansion/contraction of various genomic sequences may be just as important a mechanism of phenotypic evolution as changes at the nucleotide level. However, the population-genetic mechanisms responsible for the success vs. failure of newly arisen gene duplicates are poorly understood. We examine the influence of various aspects of gene structure, mutation rates, degree of linkage, and population size (N) on the joint fate of a newly arisen duplicate gene and its ancestral locus. Unless there is active selection against duplicate genes, the probability of permanent establishment of such genes is usually no less than 1/(4N) (half of the neutral expectation), and it can be orders of magnitude greater if neofunctionalizing mutations are common. The probability of a map change (reassignment of a key function of an ancestral locus to a new chromosomal location) induced by a newly arisen duplicate is also generally >1/(4N) for unlinked duplicates, suggesting that recurrent gene duplication and alternative silencing may be a common mechanism for generating microchromosomal rearrangements responsible for postreproductive isolating barriers among species. Relative to subfunctionalization, neofunctionalization is expected to become a progressively more important mechanism of duplicate-gene preservation in populations with increasing size. However, even in large populations, the probability of neofunctionalization scales only with the square of the selective advantage. Tight linkage also influences the probability of duplicate-gene preservation, increasing the probability of subfunctionalization but decreasing the probability of neofunctionalization.