Grant Information

ENABLING GENOMICS-ASSISTED SPECIALTY CROP BREEDING AND RESEARCH THROUGH ADVANCED DATABASE RESOURCES

Sponsoring Institution National Institute of Food and Agriculture
Program SCRI - Specialty Crop Research Initiative
Status ACTIVE
Funding Source OTHER GRANTS
Division NIFA Non Formula
Reporting Frequency Annual
Project Director Main, Doreen
Accession Number 1029319
Grant Number 2022-51181-38449
Project Number WNP00902
Proposal Number 2022-05305
Dates 2022-09-15 - 2026-09-14
Grant Year 2022
Cumulative Award Amount $5,176,559.00
Animal Health Component 70%
Recipient Organization WASHINGTON STATE UNIVERSITY
240 FRENCH ADMINISTRATION BLDG
PULLMAN,WA 99164-0001
Keywords big-data
breeding
citrus
databases
fruit
genetics
genomics
nut
pulse
rosaceae
vaccinium
Research Effort Applied (70%)
Basic (10%)
Developmental (20%)
Classification Parameters
Knowledge AreaSubject of InvestigationField of SciencePercent
201 - Plant Genome, Genetics, and Genetic Mechanisms 1119 - Deciduous tree fruits, general/other 1081 - Breeding 20%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1129 - Berries and cane fruits, general/other 1081 - Breeding 20%
201 - Plant Genome, Genetics, and Genetic Mechanisms 999 - Citrus, general/other 1081 - Breeding 20%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1212 - Almond 1081 - Breeding 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1419 - Leguminous vegetables, general/other 1081 - Breeding 10%
202 - Plant Genetic Resources 999 - Citrus, general/other 1080 - Genetics (excludes breeding) 10%
206 - Basic Plant Biology 1119 - Deciduous tree fruits, general/other 1040 - Molecular biology 10%
Non-technical Summary

This project will significantly expand existing online database resources for tree fruit, berry, nut, and pulse crops by providing access to usable big data aggregated, analyzed, integrated, and visualized, as well as data management and analysis tools. Together with personalized training, this will enable breeders and scientists to optimize the use of cutting-edge big data, tools, and predictive capability to accelerate the pace of research discovery and application in crop improvement. New and improved cultivars and management practices are essential for the specialty crop industry to mitigate a changing environment and provide consumers with access to highly nutritious specialty crop foods. ?

Goals / Objectives

Addressing the economic, environmental, and sustainability challenges facing the U.S. specialty crop industry requires us to accelerate the pace of meaningful discoveries and translate results into efficient and effective crop improvement solutions. Critical to this is turning information into knowledge using insights now afforded by the generation, analysis, and re-use of large and detailed sets of relevant "big data". Scientists in disciplines related to crop improvement in the 25 crops of this project (tree fruit, tree nut, berries, pulses) are now generating these kinds of complex data sets at an exponential rate. We propose expanding the rosaceae, citrus, vaccinium, and pulse crop databases to meet community need for resources with usable big-data aggregated, analyzed, integrated, and visualized as well as data management and analysis tools. Training and outreach in the efficient collection and use of big data, breeding data management systems and databases are integral to these efforts. We will meet these goals through the following objectives:

  1. DATA - Collect, curate, and integrate all types of genomics, genetics, and breeding big data in easy-to-use and robust crop-specific databases.
    1. Add new content and data types, such as Pan-Genome, Epigenome, and Other Large-Scale Genomic Data; Gene Transcript and Expression Data; Gene/Mutant/Disease Data; GWAS, Trait Locus, and Genetic Map Data; and Large-Scale Phenotypic and Genotypic Data and Germplasm
    2. Analyze big data, including pan-genome, functional annotation of whole genome sequences, biochemical pathways, and synteny among related crops
    3. Improve integration of diverse datasets through curation, analysis, and interface design
    4. Investigate the potential of AI to improve translation of global predictions for optimizing complex utilization in U.S. breeding programs.
  2. TOOLS - Develop and integrate new or improved tools to promote the collection, integration, and utilization of by scientists and breeders.
    1. Expand the Field Book App and BIMS to meet the specific need of specialty crop breeders for improved field data collection
    2. Modify and implement the GOAT community gene curation tool
    3. Develop a graphical and search viewer for exploring integrated data of genome to phenome across related crops
  3. OUTREACH - Provide scientists and breeders with personalized training in the use of big data, tools, and other resources of the Rosaceae, Citrus, Vaccinium, and Pulse crop databases, and develop and disseminate re-usable, extendable materials for training and broader community outreach.
    1. Expert trainers (crop postdocs) deliver on-site and online training on effective and efficient use of the Breeding Information Management System (BIMS), the Field Book App, and the databases
    2. Use communication platforms (social media, newsletters, tutorials, growers meetings, conference presentations, and publications) to disseminate information to all stakeholders
    3. Evaluate the impact of the databases to specialty crop breeding and research outcomes
    4. Provide support for other specialty crops in using the resource-efficient Tripal database platform
Methods (unparsed)

OBJECTIVE 1: DATA - This will involve (1a) Addition of Pan-genome, epigenome, other large-scale genomic data, gene transcript, expression data, gene, mutant, and disease data, GWAS, Trait Locus, and Genetic Map Data, large-scale phenotypic and genotypic data and germplasm data to the databases through collection from NCBI, publications and direct submission from researchers. Various ontologies such as Plant Ontology, Trait Ontology, and Environmental Ontology will be used to annotate plant tissue, traits, and conditions used in the experiments and efficiently integrated with all the other relevant genomic, genetic, and breeding data; (1b) Analysis of big data, including pan-genome, functional annotation of whole genome sequences, biochemical pathways, and synteny among related crops. In addition to hosting pan-genome data generated by researchers, we will perform pan-genome analysis for each species in our database using Pandagma, modified as needed for our analysis/system. We will perform additional analysis of new genome assemblies through computational annotation of the predicted genes with homology to genes of closely related or plant model species and assignment of InterPro protein domains, Gene Ontology terms, and ortholog terms through synteny analysis and KEGG categories. We will align any other genomic features, such as transcripts and genetic markers, to new or updated whole genome sequences when data become available. We will position QTLs from the literature within the current whole genome sequences of each crop. Genomic positions of QTLs will be identified using co-localizing or neighboring markers. We will perform synteny analysis to find conserved syntenic regions among the newest versions of all publicly available high-quality genomes in our databases using MCScanX and the data will be displayed using our Synteny Viewer Whenever a new whole genome data becomes available, we will also construct PlantCyc (metabolic pathway) databases using PathwayTools . The predicted genes from the whole genome sequences will be functionally annotated through a sequence homology search using BLASTX with the UniProtKB/Swiss-Prot and TAIR. BLASTX matches will be parsed, and EC codes transferred to our crop gene models. We will update the Pathway databases as enough new data become available both from the whole genome annotations of our crops and/or the model databases, UniProtKB/Swiss-Prot and TAIR; (1c) Improving the integration of diverse datasets through curation, analysis, and interface design;and (1d) Investigating the potential of AI to improve translation of global predictions for optimizing complex utilization in U.S. breeding programs. Global datasets for almond, blueberry, citrus, and lentil will be accumulated via outreach to U.S. and global communities. Preliminary apple, cherry, peach, and strawberry datasets will be used to develop imputation methods and pipelines, and to identify and resolve issues in global genomic prediction models. This approach will be applied to other project crops as datasets become available. For citrus, we will collaborate with an ongoing U.S. citrus GWAS and genomic selection project to pilot new phenotyping technology, the incorporation of existing data and new data types into the citrus database system, and the use of our technology for genomic prediction of performance for HLB tolerance or resistance, overall plant health, and fruit quality attributes.OBJECTIVE 2: TOOLS - This will involve (2a) Expansion of the Field Book App and BIMS to meet the specific need of specialty crop breeders for improved field data collection - Field Book App will add support for repeated measures/subsampling and create a framework that allows simplified addition of novel trait types such as integrated algorithms, customized crop-specific layouts, and external sensor hardware that are often used as a selection tool in specialty crops (e.g., NIRS, colorimeters). Based on direct feedback, additional apps or features that fill specific breeding informatics gaps will be developed or improved to ensure that breeders can manage their program and data entirely within the BIMS digital breeding ecosystem; (2b) Modifying and implementing the GOAT community gene curation tool - The GOAT software will be modified to make it a Tripal module and to integrate the gene data with other data types and ontologies; and (2c) Developing a graphical and search viewer for exploring integrated data of genome to phenome across related crops by connecting existing tools - Tools to integrate genetic and genomic data will include the extension of the MapViewer tool to display trait-associated markers from GWAS studies on the chromosome next to genetic maps. A GWAS Catalog will be implemented to provide browsing, searching, and graphical display of GWAS data, integrated with marker/gene pages. MapViewer will also be expanded to add an option to display syntenic regions next to the genome viewer, a functionality allowing users to view QTLs and GWAS markers in well-characterized genomes along with the syntenic regions in the crop of interest without the specific QTLs or GWAS markers.OBJECTIVE 3: OUTREACH - This will involve: (3a) Expert trainers (crop postdocs) delivering on-site and online training on effective and efficient use of the Breeding Information Management System (BIMS), the Field Book App, and the databases. This activity will include facilitating the conversion of legacy breeding data into the BIMS; (3b) Use of communication platforms (social media, newsletters, tutorials, growers meetings, conference presentations, and publications) to disseminate information to all stakeholders; (3c) Evaluation of the impact of the databases to specialty crop breeding and research outcomes; (3d) Support for other specialty crops to adopt the resource-efficient Tripal database platform, and; (3e) Implement sustainability plans for the databases (developed through sister project NRSP10).

Project Timeline Tracking

Outputs

Target Audience
The target audience is predominantly rseearch scientists, breeders, bioinformaticists, database developers, database curators, both national and international, as well as, U.S. industry stakeholders from the 25 crops covered in this project. Scientists have been engaged through training workshops, presentations at scientific conferences and meetings, monthly mailing list posts, quarterly newesletters, webinars, emails about the website, and meetings with advisory groups. Bioinformaticists and database developers have been engaged through participation in monthly meetings on Tripal software, the release of new versions of Tripal and extension modules, the tripal.info website, a two day Tripal codefest, a two day BrAPIhackathon, monthly Agricultural Biological Database (AgBioData) group meetings, peer-reviewed publications, and presentations at conferences and meetings. Industry stakeholders are being engaged through presentations and interactions at coferencesand advisory committee meetings.

Changes / Problems
Due to delays in getting funding to sub contract instituitions we were not able to hire the crop postdocs until later in year 1 and are still trying to recruit one for a joint pulse/citrus curation position at WSU. Hurricanes severly impacted the population of citrus being used for GWAS studies and we are discussing alternative populations as a substitute.

Training & Professional Development
Training opportunities for researchers (postdocs, early career scientists)anddevelopersincluded participation and presentation at workshops, conferences and meetings, as well as manuscript writing and participation in short courses. A BraPI hackathon and Tripal codefest enabled developers to contribute in person or remotely to work with other developers from across the world on shared projects.

Dissemination Streams
Through 4 training workshops at the major conferences for the community (PAG, ASHS, NAPB, Rosaceae biennial International Conference), 11 oral and 9 poster presentations at conferences,monthly community emails for each database (quarterly newsletters for each database, 7 "how to" videos, and a "year in review"), database and BIMS brochures at every conference or meeting, training webinars, 4 advisoryboard meetings, and twitter.

Next Reporting Steps
DATA: We will ramp up data curation and ingest of genomics, genetics, and breeding data to the databases now that we have most of the crop postdocs hired and trained. We will also collect and curate data for performance prediction modeling efforts and decide on a final plan to deal with issues of data collection for citrus phenotyping which was impacted by hurricanes in the research trial populations in Florida. TOOLS: We will start work on modifying and implementing the GOAT community gene curation tool, add repeated measures functionality to BIMS, add geo-navigation plot/tree identification to Field Book, and work on global performance prediction modeling. OUTREACH: We will conduct in-person program outreach training on using BIMS, FieldBook and the databases (made possible now that we have most of the crop postdocs hired and trained) as well as hold training workshops at conferences, presentations and posters at conferences and meetings, do "How to" videos, newsletter, meet with the project advisory committee, as well as database advisory committees, and submit manuscripts for CGD, GDV, PCD to peer-reviewed journals.

Outputs

Target Audience
The target audience is predominantly rseearch scientists, breeders, bioinformaticists, database developers, database curators, both national and international, as well as, U.S. industry stakeholders from the 25 crops covered in this project. Scientists have been engaged through training workshops, presentations at scientific conferences and meetings, monthly mailing list posts, quarterly newesletters, webinars, emails about the website, and meetings with advisory groups. Bioinformaticists and database developers have been engaged through participation in monthly meetings on Tripal software, the release of new versions of Tripal and extension modules, the tripal.info website, a two day Tripal codefest, a two day BrAPIhackathon, monthly Agricultural Biological Database (AgBioData) group meetings, peer-reviewed publications, and presentations at conferences and meetings. Industry stakeholders are being engaged through presentations and interactions at coferencesand advisory committee meetings.

Changes / Problems
Due to delays in getting funding to sub contract instituitions we were not able to hire the crop postdocs until later in year 1 and are still trying to recruit one for a joint pulse/citrus curation position at WSU. Hurricanes severly impacted the population of citrus being used for GWAS studies and we are discussing alternative populations as a substitute.

Training & Professional Development
Training opportunities for researchers (postdocs, early career scientists)anddevelopersincluded participation and presentation at workshops, conferences and meetings, as well as manuscript writing and participation in short courses. A BraPI hackathon and Tripal codefest enabled developers to contribute in person or remotely to work with other developers from across the world on shared projects.

Dissemination Streams
Through 4 training workshops at the major conferences for the community (PAG, ASHS, NAPB, Rosaceae biennial International Conference), 11 oral and 9 poster presentations at conferences,monthly community emails for each database (quarterly newsletters for each database, 7 "how to" videos, and a "year in review"), database and BIMS brochures at every conference or meeting, training webinars, 4 advisoryboard meetings, and twitter.

Next Reporting Steps
DATA: We will ramp up data curation and ingest of genomics, genetics, and breeding data to the databases now that we have most of the crop postdocs hired and trained. We will also collect and curate data for performance prediction modeling efforts and decide on a final plan to deal with issues of data collection for citrus phenotyping which was impacted by hurricanes in the research trial populations in Florida. TOOLS: We will start work on modifying and implementing the GOAT community gene curation tool, add repeated measures functionality to BIMS, add geo-navigation plot/tree identification to Field Book, and work on global performance prediction modeling. OUTREACH: We will conduct in-person program outreach training on using BIMS, FieldBook and the databases (made possible now that we have most of the crop postdocs hired and trained) as well as hold training workshops at conferences, presentations and posters at conferences and meetings, do "How to" videos, newsletter, meet with the project advisory committee, as well as database advisory committees, and submit manuscripts for CGD, GDV, PCD to peer-reviewed journals. <br><br>

Impacts (unparsed)

<br>What was accomplished under these goals? Objective 1:DATA -Collect, curate, and integrate all types of genomics, genetics, and breeding big data in easy-to-use and robust crop-specific databases. In year 1 of this award we added genomic, genetic and breeding data to our rosaceae (GDR), citrus (CGD), vaccinium (GDV) and pulse crop (PCD) databases. In total we have added 54 genome sequences, 3,773,154 genes, 4,090,430 mRNAs, 52,971 germplasm, 6,194 genotypes, 256,193 markers, 23 genetic maps, 1028 QTL/MTL, and 158 species. As per our stated objective 1A, we also added new data types for many of our databases. These included GWAS, methylation, expression, haplotype, and pan genome data. As part of objective 1B, we performed computational analysis on genes and synteny analysis for all new genomes which are then made searchable via our updated MegaSearch tool and synteny/genome viewers. Breeders using our Breeding Information Management System (BIMS) continued to add phenotypic data to their private breeding programs, with most using the Field Book App to collect and transfer these data to their BIMS accounts. Co-PI's Gmitter (Florida) and Ru (Auburn) started using the BIMS for their program with data upload and assessment ongoing, while Co-PI Gasic (Clemson), a long-term user of BIMS, continued to upload phenotypic and genotypic data to her program. High-throughput phenotyping for the citrus GWAS project was started with data collected by drone for canopy width and NDVI. For Objective 1D, work was started by Co-PI Hardner (Queensland) on working with collaborators to obtain phenotype and genotype datasets to assess inclusion in global performance prediction for various fruit crops. Work on crop ontology development for strawberry was initiated by Co-PI's Bassil and Ru (Auburn). Objective 2: TOOLS -Develop and integrate new or improved tools to promote the collection,integration, and utilization of by scientists and breeders. In year 1 of this award, Co-PI Rife was able to recruit a developer to start working on improvements to the Field Book App for specialty crop breeders. Field Book has been further developed with new features. Repeated measures have been added which allows plant breeders to easily collect multiple phenotypes per trait-entry. Themes have been added to adjust the size of fonts and modify the app color scheme which will allow better support on eInk devices that have better outdoor visibility and increased battery life. The delivery of new versions of Field Book have been automated using GitHub actions allowing faster turnaround to fix issues and integrate new features. A number of improvements were made for the Breeding Information Manage System (BIMS). These included, enabling Trait descriptor export (BIMS-Field Book) through BrAPI, implicit grant authorization added from BIMS to FieldBook, BIMS/MCL made compatible with PHP8, and a standalone BreedwithBIMS program (www.breedwithbims.org) was created to allow any plant breeder to use BIMS (so independent of our integrated database BIMS). Work was started to enable image upload in BIMS. In other work, the gene expression module was modified for biomaterial page inclusion, New data templates were created for GWAS data submission and GWAS data search was added to MegaSearch and GWAS viewing functionality in MapViewer. Objective 3. OUTREACH -Provide scientists and breeders with personalized training in the use of big data, tools, and other resources of the Rosaceae, Citrus, Vaccinium, and Pulse crop databases, and develop and disseminate re-usable, extendable materials for training and broader community outreach. Despite delays in hiring crop postdocs for data curation and outreach In year 1 of this award, considerable effort was made by the project team on outreach and training to showcase the databases and tools of this project. This included workshops at four conferences:" Database Resources for Crop Genomics, Genetics and Breeding Workshop": at the 2023 International Plant and Animal Genome Conference (January, 2023); "GDR Training Workshop" at the International Rosaceae Genomics Conference (March, 2023); "Breeding Tools Workshop" at the National Association of Plant Breeding Annual Meeting (July, 2023); and "Hands-on Training for Effective Use, Data Contribution, and Options for Long Term Sustainability of Specialty Crop Community Databases Workshop" at the American Society for Horticultural Science Annual Conference (August, 2023). These events were well attended, with excellent feedback provided by participants. To increase the connectivity between Field Book and external database systems like BIMS, Clemson hosted the Breeding API hackathon in March 2023. Field Book developers worked with WSU to ensure BrAPI compatibility and connectivity would remain robust after recent version updates. In addition to our training workshops, we also gave a total of 11 oral and 9 poster presentations at three of these conferences in relevant sessions. For all these conferences we brought and disseminated brochures for the databases. We have monthly emails to each of our database communities. These included 7 short "How to" videos regarding using the resources in the databases, 4 quarterly newsletters for each database highlighting new data, features, and community news, etc, and a happy holidays, end of calendar year flyer summarizing the year in review for each database. We also held online training for several groups who wanted dedicated training. Training materials were updated for all our databases, BIMS and FieldBook. The PhenoApps project website (phenoapps.org) was updated with hardware suggestions and a contact form to enable requests for new/improved features and get technical help. Field Book v5.4 documentation was moved online to https://docs.fieldbook.phenoapps.org. The manual for BIMS was updated and made available on the homepage of www.breedwithbims.org. We were happy to have 2 groups successfully create BIMS programs just by using the manual. BIMS is to our knowledge the only free, online breeding data management program that enables breeders to fully manage their own breeding programs, including all data upload. <br><br><b>Publications</b><br>


Publications Inventory

Conference Papers and Presentations

Other

Websites