Grant Information
| Knowledge Area | Subject of Investigation | Field of Science | Percent |
|---|---|---|---|
| 201 - Plant Genome, Genetics, and Genetic Mechanisms | 999 - Citrus, general/other | 1080 - Genetics (excludes breeding) | 50% |
| 202 - Plant Genetic Resources | 999 - Citrus, general/other | 1081 - Breeding | 50% |
Citrus fruits are a major crop in several states of the USA, but the recent introduction of Huanglongbing, a bacterial disease, and the Asian Citrus Psyllid insect which transmits it, pose the most severe threat that these industries have ever faced. To address HLB and other exotic diseases, breeders need tools to rapidly characterize citrus varieties and hybrids and to locate genes for disease resistance, fruit quality, and other essential traits. To address this problem, this project will develop a high-density SNP genotyping array for citrus. This tool allows investigators to rapidly determine the particular genetic variants present in a variety or hybrid. The 20,000 genetic variants detected with this SNP array can then be related to the traits displayed by the various individuals studied. We will use this tool to study essentially all trees in the largest and most diverse citrus variety collection in the US and several large families in which individuals vary for traits of economic importance. We will then correlate the particular genetic variants carried by each individual to measured traits such as disease resistance and fruit quality. The most valuable outcomes of this project will be a tool that citrus breeders can use to improve the efficiency of breeding, and a comprehensive understanding of relationships among citrus varieties and how these relate to economically valuable characters.
Objective 1. Develop sequences for a large and diverse set of citrus varieties. Collect all available citrus sequences from online databases and by contacting scientists that have reported sequencing but not yet released sequence. For species and other groups not adequately represented by available sequences, we will isolate DNA and sequence these to at least 12X depth using Illumina HiSeq. A total of about 50 sequences is expected. Objective 2. Design a 20,000 SNP Illumina Infinium SNP assay. Analyze all available sequence data to identify SNPs and design a robust 20,000-SNP assay, principally targeting genes while also maintaining reasonable coverage in gene-poor regions. New sequences will be assembled using SOAPdenovo and aligned to the reference sweet orange and Clementine sequences. Within and between individual SNPs will be called with several algorithms and a consensus used to select SNPs for further analysis. Initially we will select SNPs within genes and then add non-genic SNPs as needed to give good genome coverage. The genotyping array will then be designed by Illumina or another genotyping array provider. Objective 3. Apply the SNP assay to Citrus and closely related genera in the citrus germplasm collection to support phylogenetic clarifications and association mapping. We will isolate DNA from about 1000 accessions from the citrus germplasm collection and breeding parents and analyze these with the SNP array. Prior to association mapping, we will determine population structure by analysis with Structure and similar programs. Genome wide association mapping will be performed with TASSEL or similar programs with adjustments for population structure. Objective 4. Apply the SNP assay to breeding populations used for mapping disease resistance and tolerance, and populations already phenotyped for fruit quality traits. Several populations with about 500 total individuals that have already been phenotyped for various traits will be mapped using the SNP array. Additional array mapping will be made available to collaborators in Florida where plants are being phenotyped for resistance or tolerance to HLB.
Target Audience
The target audience for this final phase of the project is the community of citrus geneticists and breeders who will benefit from the resources generated. A secondary target is the broader community of plant and animal geneticists who may benefit from methods we have developed for genotyping single pollen grains and a novel method for inferring chromosome level haplotypes from analysis of a few haploid individuals.
Changes / Problems
This two-year project was extended to a third year because it took more time than expected to negotiate access to DNA sequences developed in other laboratories. We did not want to sequence the same genotypes as others. This was eventually resolved. Completion of sequencing also took more time than expected due to queues at the sequencing center and low coverage from some initial sequencing runs. Overall, we note four major improvements to the originally planned project: greater sequencing depth, development of SNP arrays with about 2.5 times the originally proposed SNP density, development of a high density (1.4M) SNP array (although the number of valid SNPs on this array will be less than 1.4M), and development of methods to infer chromosome level haplotypes for citrus.
Training & Professional Development
One Ph.D. student worked on the project. She developed improved DNA isolation methods for citrus and used these to collect most of the germplasm collection samples that we analyzed. She also learned bioinformatics methods used in selecting sequences for inclusion on the array. Training in plant genetics and biotechnology was provided to this student through personal meetings and presentations to the lab group. She expects to present a talk on the citrus array project at the PAG meeting in January 2017. A visiting scientist from Italy developed the pollen grain isolation and whole genome amplification methods used. He received training in bioinformatics use in SNP analysis, presented a poster at PAG in January 2016 and will do so again in 2017. A visiting scientist from India is analyzing array data for loss-of-heterozygosity and copy number variation and will present a poster at PAG in 2017.
Dissemination Streams
Results have been disseminated through talks and posters at scientific meetings. Manuscripts are being prepared for submission to journals. Posters were presented at the PAG meeting in San Diego in 2016 and a talk and two posters will be presented in January 2017. One talk and one poster developed from the project were presented at the International Citrus Congress in Brazil in September 2016. Sequence data will be released through the NCBI Sequence Read Archive and through the Citrus Genome Database (https://www.citrusgenomedb.org/). We expect to sign an agreement with Affymetrix to allow commercialization of the two Citrus arrays. Other citrus researchers will be able to submit larger numbers of samples directly to Affymetrix. The Roose lab will also collect and consolidate smaller numbers of samples from other labs and periodically submit these to Affymetrix for analysis. This will allow the tool to be used by those without sufficient funding to process large numbers of samples.
Next Reporting Steps
Nothing Reported
Target Audience
The target audience for this initial phase of the project is the community of citrus geneticists and breeders who will benefit from the resources generated.
Changes / Problems
This two-year project was extended to a third year because it took more time than expected to negotiate access to DNA sequences developed in other laboratories. We did not want to sequence the same genotypes as others. This was eventually resolved. Completion of sequencing also took more time than expected due to queues at the sequencing center and low coverage from some initial sequencing runs. Overall, we note four major improvements to the originally planned project: greater sequencing depth, development of SNP arrays with about 2.5 times the original SNP density, development of a high density (1.3M) SNP array (although the number of valid SNPs on this array will be less than 1.3M), and development of methods to infer chromosome level haplotypes for citrus. We expect to collect all planned marker data by the revised end date, but complete analysis of all of the data will likely require additional time.
Training & Professional Development
One Ph.D. student is working on the project. She is developing improved DNA isolation methods for citrus and using these to collect the large number of samples that we plan to analyze. She has also learned bioinformatics methods using in selecting sequences for inclusion on the array. Training in plant genetics and biotechnology is being provided to this student through personal meetings and presentations to the lab group. Two undergraduate students provided assistance in DNA isolation and received training in lab procedures.
Dissemination Streams
Nothing Reported
Next Reporting Steps
Objective 1) During the next reporting period we will conduct a phylogenetic analysis of citrus based on the citrus sequence data, and further analyze evidence for recent and ancient hybridization and introgression in the evolution of citrus. Sequences will be released in December 2015 or January 2016. Objective 2) We have designed a 1.3M SNP array for citrus and will use information from this to validate SNPs and design a robust, 50K SNP array. Objective 3) Genetic diversity analysis. We will analyze about 200 diverse citrus accessions using the 1.3M SNP array, and then analyze an additional 800 accessions using the 50K array. Association mapping for various traits will be explored using this database. We have developed methods for whole genome amplification of DNA from single pollen grains that are very effective for amplification for SSR markers and we expect this WGA DNA to also be suitable for SNP array analysis. Analysis of a few pollen grains from each genotype will allow assignment of haplotypes to selected accessions, increase our understanding of citrus evolution and improve accuracy and resolution of association mapping. Objective 4) The 50K array will be used to generate dense (at least 1000 marker) maps for several mapping populations. The maps will be combined with phenotypic data and QTL analysis performed. <br><br>
<br>What was accomplished under these goals? 1) Develop sequences for a large and diverse set of citrus varieties. Before deciding which samples we would sequence, we first determined what sequence data we could obtain from public repositories and, by agreement, from other researchers. We then identified 30 accessions that, together with the 12 sequenced accessions available from others, would represent diversity in citrus. These include the major commercial cultivar types as well as ancestral species and closely related and interfertile genera. Several accessions with apparent HLB tolerance were included. DNA was isolated and sequenced on the Illumina Hi-Seq 2500. The citrus genome is about 380 Mb but essentially all individuals are fairly heterozygous. To capture this heterozygosity we chose to sequence 2 x 100 bp reads and 30X nominal depth. Nominal sequence depth for the 30 accessions ranged from 5.7X to 53X, with a mean of 29.0. Three samples had less than 20X depth and 60% had greater than 25X depth. Across all samples an average of 7.8% of reads did not align with the 301 Mb reference genome. Variation in the percentage of reads that did not align was not related to taxonomic distance from the reference. About 19.9M variants were annotated: 17.7M SNPs, 0.9M insertions, and 1.2M deletions, for a total variant rate of about 1 per 14 bases. Among the variant positions, percentage of heterozygous variants was 2-3% in citrons, 4-5% in pummelos, trifoliate oranges, and kumquat, 3-8% in mandarins (some introgressed), and 12-19% in known interspecific hybrids. We also have chloroplast genome sequences for all accessions and are using these to develop a definitive chloroplast phylogeny of these accessions. 2) Design a 20,000 SNP Illumina Infinium SNP assay. We investigated SNP platforms from Illumina and Affymetrix and eventually chose Affymetrix as offering greater capacity for the funds available. We are developing two Affymetrix Axiom arrays. An initial 1.3M SNP array will be used to genotype 288 samples including about 200 accessions selected to represent diversity, parent-offspring trios, varieties that have diverged by mutation, and whole-genome amplified samples from single pollen grains. Result from this array will be used to validate SNPs selected from the sequence database and also provide high density coverage of selected accessions. We will then select about 50,000 SNPs to be tiled on a lower cost array that will be used for analysis of germplasm and mapping populations as outlined in 3) and 4) below. The 1.3M SNP arrays is now being manufactured, and the DNA samples for this analysis have been prepared. 3) Apply the SNP assay to Citrus and closely related genera in the citrus germplasm collection to support phylogenetic clarifications and association mapping. We have isolate DNA from nearly all germplasm accessions to be studied using a new protocol to first reduce the waxy coating on citrus leaves, dry the leaves with silica gel, and then isolate DNA with a commercial kit. 4) Apply the SNP assay to breeding populations used for mapping disease resistance and tolerance, and populations already phenotyped for fruit quality traits. Most leaf and/or DNA samples for this objective have been obtained. <br><br><b>Publications</b><br>
Target Audience
The target audience for this initial phase of the project is the community of citrus geneticists and breeders who will benefit from the resources generated.
Changes / Problems
This two-year project is somewhat behind schedule because it took more time than expected to negotiate access to DNA sequences developed in other laboratories. We did not want to sequence the same genotypes as others. This was eventually resolved. Completion of sequencing is also taking more time than expected due to queues at the sequencing center. A concern is that the approved budget does not include funding for personnel to analyze the SNP data to be generated by this project. Therefore we have explored less expensive alternatives to Illumina SNP arrays, specifically a variant of genotyping-by-sequencing that uses sequence capture methods to produce a targeted, representative set of sequences from each sample. However the bioinformatics workload for analyzing this type of data in heterozygous genotypes is likely to be higher than with Illumina SNP data. A final decision on the platform has not been made.
Training & Professional Development
One Ph.D. student is working on the project. She is developing improved DNA isolation methods for citrus and using these to collect the large number of samples that we plan to analyze. Training in plant genetics and biotechnology is being provided to this student through personal meetings and presentations to the lab group.
Dissemination Streams
No results suitable for dissemination at this point.
Next Reporting Steps
Objective 1) During the next reporting period we will complete DNA sequencing and analyze the DNA sequence database to identify SNPs. Objective 2) Design a 20,000 SNP assay. We will analyze the SNP database to identify a set of SNPs expected to maximize information about citrus germplasm. We will like run a set of already sequenced samples through a sequence-capture/HT Sequencing process to evaluate this system. Objective 3) Genetic diversity analysis. We will complete preparation of DNA from at least 600 of the 1000 germplasm accessions and at least 300 individuals from mapping populations. The SNP assay will be applied to these samples. A preliminary analysis of genetic diversity will be completed. Objective 4) The mapping population data will be analyzed to generate linkage maps and QTL analysis initiated. It is likely that a no-cost extension to the project will be required to complete the remaining samples. <br><br>
<br>What was accomplished under these goals? 1) Develop sequences for a large and diverse set of citrus varieties. Before deciding which samples we would sequence, we first determined what sequence data we could obtain from public repositories and, by agreement, from other researchers. We then identified 30 accessions that, together with the 13 sequenced accessions available from others, would represent diversity in citrus. These include the major commercial cultivar types as well as ancestral species and closely related and interfertile genera. Several accessions with apparent HLB tolerance were included. DNA was isolated and submitted for sequencing on the Illumina Hi-Seq 2500. The citrus genome is about 380 Mb but essentially all individuals are fairly heterozygous. To capture this heterozygosity we chose to sequence 2 x 100 bp reads and 30X nominal depth. We are currently waiting for the sequencing to be completed. 2) Design a 20,000 SNP Illumina Infinium SNP assay. Mainly because of cost issues, we are investigating alternatives to the Infinium assay, primarily a sequence capture method followed by sequencing. We developed a strategy to allocate the 20,000 SNPs within and between major species groups. Groups, such as mandarins, that are of greater commercial importance and have more diversity have deeper coverage. SNPs will be selected primarily from genes, but we will include at least one SNP per Mb for the major groups. 3) Apply the SNP assay to Citrus and closely related genera in the citrus germplasm collection to support phylogenetic clarifications and association mapping. We have begun to evaluate alternative DNA isolation methods than we have used previously for citrus. For the large number of samples to be studied with the array, a more rapid and reliable method would be quite valuable. 4) Apply the SNP assay to breeding populations used for mapping disease resistance and tolerance, and populations already phenotyped for fruit quality traits. No accomplishments yet in this area. <br><br><b>Publications</b><br>