Grant Information

DATABASE RESOURCES FOR CROP GENOMICS, GENETICS AND BREEDING RESEARCH

Sponsoring Institution National Institute of Food and Agriculture
Status COMPLETE
Funding Source HATCH
Division NIFA Formula
Reporting Frequency Annual
Project Director Main, Doreen
Accession Number 1005849
Project Number WNP00822
Multistate Number NRSP-_OLD_10
Dates 2015-01-20 - 2019-09-30
Animal Health Component 50%
Performing Department Horticulture & Landscape Architecture
Recipient Organization WASHINGTON STATE UNIVERSITY
240 FRENCH ADMINISTRATION BLDG
PULLMAN,WA 99164-0001
Keywords citrus
cotton
database
genetics
genomics and breeding
legumes
rosaceae
tripal
vaccinium
Research Effort Applied (50%)
Basic (50%)
Developmental (0%)
Classification Parameters
Knowledge AreaSubject of InvestigationField of SciencePercent
201 - Plant Genome, Genetics, and Genetic Mechanisms 1119 - Deciduous tree fruits, general/other 1080 - Genetics (excludes breeding) 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1119 - Deciduous tree fruits, general/other 1081 - Breeding 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1129 - Berries and cane fruits, general/other 1080 - Genetics (excludes breeding) 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1129 - Berries and cane fruits, general/other 1081 - Breeding 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1419 - Leguminous vegetables, general/other 1080 - Genetics (excludes breeding) 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1419 - Leguminous vegetables, general/other 1081 - Breeding 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1719 - Cotton, other 1080 - Genetics (excludes breeding) 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 1719 - Cotton, other 1081 - Breeding 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 999 - Citrus, general/other 1080 - Genetics (excludes breeding) 10%
201 - Plant Genome, Genetics, and Genetic Mechanisms 999 - Citrus, general/other 1081 - Breeding 10%
Non-technical Summary

This projects will establish a robust, dynamic, and widely available genome database platform immediately useful for crops of national significance currently underserved (Citrus, Cool Season Food Legumes, Cotton, Rosaceae, Vaccinium), and flexible enough to implement for other crops and organisms important to U.S. agriculture. To do this we will: Expand technical resources available to existing online databases currently housing genomics, genetics and breeding data and bioinformatics tools and establish consensual standards, protocols and applications for data collection, organization, storage, analysis, and curation so other crops/organisms can efficiently leverage genome database platforms and analytical tools.

Goals / Objectives
  1. Expand online community databases for Rosaceae, citrus, cotton, cool season food legumes and Vaccinium crops: We will continue ongoing curation to integrate new genomic, genetic, genotypic, phenotypic and germplasm data in all community databases. For example, in the GDR, we will add all the data from the large European FruitBreedOmics project (www.fruitbreedomics.com). We will also introduce software innovations first developed in the GDR to other community databases. New or updated genome sequence and annotation data will be added to the databases with additional computational analysis performed to identify predicted genes with homology to known genes in other databases (our standard analysis pipeline). Whole genome sequences will also be used to construct orthologous regions among closely related genomes. PlantCyc metabolic pathway databases will be displayed using GBrowse_Syn (McKay et al. 2010) and PathwayTools (Paley et al. 2012). Genetic data will continue to be integrated and the trait loci will be associated with the Trait Ontology (TO) terms (Jaiswal et al. 2002), with new terms added as necessary. Large scale phenotypic and genotypic data will continue to be integrated. Within the framework of the NRSP, these types of targets may be nominated by a specific crop but will be chosen for development mostly on the perceived need for comparable resources by the NRSP crop communities
  2. Develop a tablet application to collect phenotypic data from field and laboratory studies: A tablet application will be developed for field and laboratory studies to record phenotypes, take pictures and submit to the appropriate database. Users will be able to upload site information, trait descriptors, dataset names, germplasm, new observation values and comments. The app will have functionality to associate pictures with each observation of germplasm and will allow users to temporarily store data in the tablet for subsequent upload to the appropriate database. We will also add an option to store data from the tablet in the cloud so that users can keep the data until they are ready to upload to the database. We will review open-source tablet software applications, such as the one developed for maize by the Dr. Ed Buckler Lab at USDA (http://www.maizegenetics.net/field-informatics), to see if we can modify to work with Chado
  3. Develop a Tripal Application Programming Interface for building breeding databases: GDR has interfaces to search breeding data and decision tools such as Marker Converter and Cross Assist to help breeders refine markers and support crossing decisions, respectively. We will convert and expand functionalities of these interfaces into Tripal modules for transfer to our other databases or other Tripal-based databases
  4. Convert GenSAS, the community genome annotation tool, to Tripal: GenSAS (Main et al. 2013c) is a web-based Genome Sequence Annotation Server that provides a one-stop website with a single graphical interface for running multiple structural and functional annotation tools, enabling visualization and manual curation of genome sequences. The availability of a single web application where users can combine analyses and curation for their locus of interest will accelerate the refinement of whole genome sequence data by community experts. In this NRSP, we will develop a Chado exporter for GenSAS to allow the genome curator to export completed structural and functional annotations into Chado. GenSAS in Tripal provides a complete whole genome annotation and visualization platform for any research community Develop Web Services to promote database interoperability: Web services will be enabled for retrieval of sequence and annotation data by applying the most commonly used technologies for Web Services such as Representational State Transfer (REST) and the Simple Object Access Protocol (SOAP). We will follow developing standards and recommendations from BioHackathon, an organization who represent the major biological databases and cyber-infrastructure projects worldwide (Katayama et al. 2013)
Methods (unparsed)

Objective 1: Expand online community databases for Rosaceae, citrus, cotton, cool season food legumes and Vaccinium crops.We will continue ongoing curation to integrate new genomic, genetic, genotypic, phenotypic and germplasm data in all community databases. For example, in the GDR, we will add all the data from the large European FruitBreedOmics project (www.fruitbreedomics.com). We will also introduce software innovations first developed in the GDR to other community databases. New or updated genome sequence and annotation data will be added to the databases with additional computational analysis performed to identify predicted genes with homology to known genes in other databases (our standard analysis pipeline). Whole genome sequences will also be used to construct orthologous regions among closely related genomes. PlantCyc metabolic pathway databases will be displayed using GBrowse_Syn (McKay et al. 2010) and PathwayTools (Paley et al. 2012). Genetic data will continue to be integrated and the trait loci will be associated with the Trait Ontology (TO) terms (Jaiswal et al. 2002), with new terms added as necessary. Large scale phenotypic and genotypic data will continue to be integrated. Within the framework of the NRSP, these types of targets may be nominated by a specific crop but will be chosen for development mostly on the perceived need for comparable resources by the NRSP crop communities.Objective 2: Partner with Kansas State University (Dr. Jesse Poland) to further develop the mobile App Field Book, originally developed for wheat phenotype data collection, to be fully functional for Specialty Crop Breeders field and laboratory phenotype data collection.The coding modifications, which will include ability to synchronize between devices in a program will be carried out as a subcontract to Dr. Jesse Poland of KSU, the program who developed the FieldBook App (http://www.wheatgenetics.org/bioinformatics/22- android-field-book), with guidance on functionality provided by NRSP10 breeders. Working with Jesse Poland we already have over 20 breeders evaluating FieldBook and it has received high approval rating for its routine adoption in specialty crop phenotyping provided some new features of particular relevance to horticultural crops such as perennial nature and use of clones are developed, which Jesse's group will do with this subcontract. An API will be developed to upload data directly to Tripal databases and connect with the Breeding Information Management System that is currently under development. This subcontract will leverage resources spent on developing the existing FieldBook App functionality. Current features include site information, trait descriptors, dataset names, germplasm, new observation values, trial and comments. The app will have functionality to associate pictures with each observation of germplasm and will allow users to temporarily store data in the tablet for subsequent upload to the appropriate database. We will also add an option to store data from the tablet in the cloud so that users can keep the data until they are ready to upload to the database.Objective 3: Develop Tripal API for building breeding databases.GDR has interfaces to search breeding data and decision tools such as Marker Converter and Cross Assist to help breeders refine markers and support crossing decisions, respectively. We will convert and expand functionalities of these interfaces into Tripal modules for transfer to our other databases or other Tripal-based databases.Objective 4: Convert GenSAS, the community genome annotation tool, to Tripal.GenSAS (Main et al. 2013c) is a web-based Genome Sequence Annotation Server that provides a one-stop website with a single graphical interface for running multiple structural and functional annotation tools, enabling visualization and manual curation of genome sequences. The availability of a single web application where users can combine analyses and curation for their locus of interest will accelerate the refinement of whole genome sequence data by community experts. In this NRSP, we will develop a Chado exporter for GenSAS to allow the genome curator to export completed structural and functional annotations into Chado. GenSAS in Tripal provides a complete whole genome annotation and visualization platform for any research community.Objective 5: Develop Web Services to promote database interoperability.Web services will be enabled for retrieval of sequence and annotation data by applying the most commonly used technologies for Web Services such as Representational State Transfer (REST) and the Simple Object Access Protocol (SOAP). We will follow developing standards and recommendations from BioHackathon, an organization who represent the major biological databases and cyber-infrastructure projects worldwide (Katayama et al. 2013).Detailed deliverables and milestones are covered in the 'Projected Outcomes' and 'Timeline' under 'Management, Budget and Business Plan', respectively. The assessment methods of each objective are described in 'Outreach, Communications and Assessment' section below.

Project Timeline Tracking

Outputs

Target Audience
The target audience is predominantly scientists, breeders, bioinformaticists, database developers, database curators, both national and international, as well as, U.S. industry stakeholders from the 26 crops covered in this project. Scientists have been engaged through peer-reviewed publications, workshops, presentations at scientific conferences and meetings, newsletters to the mailing list, webinars, emails to the community, and meetings with advisory groups. Breeders have been engaged through Breeding Information Management System Training and a US breeding capacity survey with breeding program data available to view, search and visualizefrom the NRSP10 site. Bioinformaticists and database developers have been engaged through participation in monthly meetings on Tripal software, release of new versions of Tripal, the tripal.info website, a two day in-person hackathon, monthly Agricultural Biological Database (AgBioData) group meetings, peer-reviewed publications, and presentations at conferences and meetings. Industry stakeholders are being engaged through presentations and interactions at commodity group meetings and advisory committee meetings.

Changes / Problems
Nothing Reported

Training & Professional Development
Training opportunities for undergraduate, graduate and postdoctoral researchers included participation in hackathons, presentations and participation at workshops, conferences, and meetings, and authorship on peer-reviewed publications. In addition three developers attended the DrupalCon in Year 5 of this award.

Dissemination Streams
Disseminating the results of this project has been accomplished by the following activities in year 5 of NRSP10: 10 peer reviewed publications with one awaiting publication and the other in review, 1 book chapter, 25 presentations at 8 conferences/meetings (Crop Science Society of America Annual Conference, Washington State Tree Fruit Association Annual Meeting, Cotton Beltwide Conference, International Plant and Animal Genome Conference, International Research Conference on Huanglongbing, Annual Plant Biology Meeting, American Society for Horticultural Science Annual Meeting, and National Association of Plant Breeders Annual Meeting), as well as webinars, brochures, database booths, posters and a NRSP10 workshop at the 2019 PAG conference and two BIMS workshops at the Annual RosBREED participants meeting and the Cotton Beltwide Conference. All presentations/webinars are available from the database websites and the NRSP10 project website.

Next Reporting Steps
As the project has been recommended for renewal, we will be continuing to develop the NRSP10 tools and resources. This will include finishing the one objective we did not complete, integration of a light version of GenSAS within each of the databases.

Outputs

Target Audience
The target audience is predominantly scientists, bioinformaticists, database developers, database curators, both national and international, as well as, U.S. industry stakeholders from the 25 crops covered in this project. Scientists have been engaged through peer-reviewed publications, workshops, presentations at scientific conferences and meetings, newsletters to the mailing list, webinars, online surveys of users, emails about the website, and meetings with advisory groups. Bioinformaticists and database developers have been engaged through participation in monthly meetings on Tripal software, release of new versions of Tripal, the tripal.info website, a twoday in-person hackathon, monthly Agricultural Biological Database (AgBioData) group meetings, peer-reviewed publications, and presentations at conferences and meetings. Industry stakeholders are being engaged through presentations and interactions at commodity group meetings and advisory committee meetings.

Changes / Problems
Nothing Reported

Training & Professional Development
Training opportunities for undergraduate, graduate and postdoctoral researchers included participation and presentation at workshops, conferences and meetings, as well as manuscript writing and participation in short courses.

Dissemination Streams
Disseminating the results of this project has been accomplished by the following activities in year 4 of NRSP10: 3 peer-reviewed publications, 1 non-peer reviewed publications (Acta Hort), 2 book chapters and several papers under review, 27 presentations at 8 conferences (International Plant and Animal Genome Conference, Cotton Beltwide Conference, Bioinformatics Open Source Conference, International Cotton Genomics Initiative Conference, International Rosaceae Genomics Conference, Plant Biology Meeting, American Society for Horticultural Science Annual Meeting, National Association of Plant Breeders Annual Meeting), as well as webinars, brochures, database booths, posters and a NRSP10 workshop and annual meeting at the ASHS conference. All presentations/webinars are available from the database websites and the NRSP10 project website.

Next Reporting Steps
Our plans for Year 5 of NSRP10 are as follows: (1) Finish creating and implementing a GenSASLite version that directly plugs into Tripal for better community curation within Tripal databases (2) BIMS: Further develop BIMS by adding ability to upload genotype data and add new searching capability, as well as work more on providing dedicated training to breeding program participants. (3) Keep the NRSP10 databases current with tools and data while serving as community communication portals for their respective communities (4) Continue contributing to development of Tripal Core and Tripal Extension Module Development while providing advocacy and support for Tripal database adoption and module development by the developer community (help desk, web site, hackathon, monthly meetings) (5) Continue database advocacy by participating in database outreach activities of the AgBioData Consortium (6) present the outcomes of NRSP10 through peer-reviewed publications, conferences, meetings and training activities and specifically develop outreach materials to inform industry of NRSP10 activities and their contribution to basic, translational and applied research. <br><br>

Impacts (unparsed)

<br>What was accomplished under these goals? Tripal Progress: Major progress for Tripal in Yr 4 includes release of Tripal v3.0.rc3 and the first stable release of Tripal v3 as well as release of the Tripal MapViewer extension module (for Tripal v2.1 and v3.0) and several other extension modules released. Of the available modules 30 are already Tripal v3.0 compatible with the remainder being converted. In the Tripal website, the extensions modules are now classified by category into Administration, Analysis/Annotation, Data Loading/Collection, Developer Tools, Third Party Integration (BrAPI, JBrowse, Galaxy, BLAST, VCF Filter), Searching, Visualization/Display and Development. There are 12 groups active in Tripal development in three countries, over 150 downloads of the Tripal platform and over 500 help desk support questions submitted and answered over the last 4 years. We estimate that over 4000 crop and wild relative species are now served through a Tripal database. Monthly Tripal meetings are regularly attended by 20 to 30 developers and the yearly codefest (formerly hackathon) attracts a similar number of developers. A Tripal workshop is held yearly at PAG with participant numbers exceeding 60 in the last 2 years following the inaugural workshop in 2015. The Tripal website (https://tripal.info, Figure 2) is kept current with tutorials and documentation, and regular webinars are provided for the community. In year 4 of NRSP, the Tripal site has been accessed by 6,760 users from 163 countries (1,332 U.S.) with 10,085 visits and 23,574 page views. All code is checked and approved by the Project Management Committee (PMC) before it is released to ensure standards are maintained. For newer projects, when required, we provide images of established databases, such as CottonGen to accelerate the transition to production, or provide programmatic support to update to newer Tripal versions, e.g. PeanutBase. While all new versions of Tripal core are backward compatible, a lack of Drupal experience can make it a challenging for developers at times. Having access to an experienced support system led by initiators of Tripal, Drs. Stephen Ficklin (WSU) and Meg Staton (UTenn) has been highly effective when accompanied by training for new developers. NRSP10 Databases: In year 4 we upgraded of the Citrus Genome Database to Tripal v3.0. and begun converting the other 4 to Tripal 3.0. This involved converting our Mainlab search and display modules to this new version and releasing these to the Tripal community. Conversion to Tripal 3.0 means users are now able to cross query data from the TreeGenes and Hardwood Genomics databases from the Citrus database. Significant new data and functionality has been added to all five databases (documented in detail in the work completed section of each database). Highlights include (1) synteny analysis performed for all sequenced genomes in each database, with results visualized in the Tripal Synteny Viewer (Fei Lab, BTI) and linked to all ortholog and functional annotation, (2) major expansion of the map, marker, QTL and genome sequence datasets (curated) in GDR and CottonGen as well as keeping the Citrus, Vaccinium and Legume databases up to date with published data, (3) GRIN data added for all the databases, (4) Several new/modified search interfaces added such as genotype and marker searching and (5) Updated RefTrans expression data analyses and Pathway analyses. Usage of the databases continues to grow and in Year 4 was as follows: GDR - 24,331 users from 156 countries, 649,026 pages viewed; CottonGen - 12,098 users from 143 countries, 278,001 pages viewed; CSFL - 3,685 users from 108 countries, 47,203 pages viewed; CGD - 5,410 users from 127 countries, 76,653 pages viewed; and GDV ­­­- 2,664 users from 92 countries, 30,997 pages viewed. For the first four years of the NRSP10 project, over 2.5 million pages have been accessed. GenSAS Progress: GenSAS Progress: GenSAS v5.1 was released in January 2018. Major improvements include (1) Upgraded to Tomcat v8 and Apollo v2 and enabled GenSAS to submit jobs to computational cluster which increased performance (2) Implemented restrictions to user accounts to limit the number of jobs that can be run concurrently and also added PRINSEQ to check assembly quality (3) Gene models manually edited in Apollo have functional annotation jobs run on then during publish step and integration into final annotation. GenSAS v6.0 was released September 2018. Major improvements include () Addition of the tools BUSCO, HISAT2, DIAMOND, pBLAT (version of BLAT that runs multi-threaded), and BRAKER2 (2) GenSAS now allows tools that can run multi-threaded to run as such on cluster, this has decreased the time it takes for jobs to complete (3) At the final publish step, summary reports about the types of repeats and annotation metrics are produced along with the final annotation files. A file with a summary of which tools were used to generate the final annotation, and the tool settings, is also produced. Users also have the option of creating a merged GFF3 file of the final annotation that is suitable for submission to GenBank. GenSAS was presented at the International Plant & Animal Genome Conference in January 2018 and as a component in general NRSP10 presentations. A book chapter called "Structural and functional annotation of eukaryotic genomes with GenSAS" was accepted for publication in a "Eurkaryotic Gene Prediction" volume of the Methods in Molecular Biology series. A manuscript called "GenSAS v6.0: a web-based platform for computational and manual annotation of model and non-model organism genome sequences" was submitted in September 2018 for publication in Genome Biology. In Year 4, GenSAS was accessed by 1,424 visitors from 78 countries, with 2,252 sessions and 7,079 pages viewed. It was used to annotate bacteria, viruses, fungi, plants and animal species. Breeding Tools Progress: In year 4 of the NRSP10 we continued development of the Tripal Breeding Information Management System (BIMS) and the phenotype data collection tool FieldBook App (Poland Program). New functionality includes addition of: (1) ability to archive data , (2) graphical and tabular view of phenotype statistics, (3) configuration page to set column names for phenotype files, (4) new search/download, (5) ability to generate Field Book input file for progeny from a new cross, (6) ability to add more columns to the downloaded file, (6) generate list of accessions for searches, (7) Mean/max/min/std and frequency were added in the downloaded file from search, (8) Google Map embedded in BIMS to show the locations of the sites when breeding program coordinates are provided, (9) ability to change order of traits in ''Manage Breeding' section, (10) ability to compare trait statistics from different categories (year, cross, etc), (11) Frequency of each categories are now displayed for categorical traits for a given dataset, and (12) Updated User Manual and FAQ and added video tutorials. Two full day BIMS training workshops held, several webinars and BIMS presented at several conferences in year 4. <br><br><b>Publications</b><br>

Outputs

Target Audience
The target audience is predominantly scientists, bioinformaticists, database developers/curators, both national and international, as well as, U.S. industry stakeholders from the 25 crops covered in this project. Scientists have been engaged through peer-reviewed publications, workshops, presentations at scientific conferences and meetings, newsletters to the mailing list, online surveys of users, emails about the website, and meetings with advisory groups. Bioinformaticists and database developers have been engaged through participation in monthly meetings on Tripal software, release of new versions of Tripal, the tripal.info website, a one day in-person hackathon, monthly Agricultural Biological Database (AgBioData) group meetings and a 2 day AgBioData workshop, peer-reviewed publications, and presentations at conferences and meetings. Industry stakeholders are being engaged through presentations and interactions at commodity group meetings and advisory committee meetings.

Changes / Problems
Nothing Reported

Training & Professional Development
Training opportunities for undergraduate, graduate and postdoctoral researchers included participation and presentation at workshops, conferences and meetings, as well as manuscript writing and participation in short courses.

Dissemination Streams
Disseminating the results of this project has been accomplished by the following activities in year 3 of NRSP10: 4peer-reviewed publications, 2 non-peer reviewed publications (Acta Hort), 31 presentations at 6 international conferences (XXV International Plant and Animal Genome Conference, 5th International Research Conference on Huanglongbing, 2017 North American Pulse Improvement Association Conference, 2017 PAG Asia Conference, 2017 Plant and Breeding Symposium, XI International Peach Symposium), 4 national conferences (Cotton Beltwide, Cotton Breeders Tour, 2017 American Society for Horticultural Science conference, 2017 American Society Plant Biology conference, and several local meetings. These include holding user-taught dedicated training workshops at the 8th International Rosaceae Genomics Conference for GDR, the American Society of Horticultural Science conference NRSP10 Workshop, the Cotton Beltwide Conference CottonGen Workshop, and presentations at several annual meetings: the Plant and Animal Genome Conference, the 5th International Research Conference on Huanglongbing, American Association of Plant Biology, North American Pulse Improvement Conference, as well as webinars, brochures, database booths, and posters. All presentations/webinars are available from the database websites and the NRSP10 project website.

Next Reporting Steps
Our plans for Year 4 of NSRP10 are as follows: (1) GenSAS: Further develop GenSAS to add more tools and error checking of assemblies, and create a GenSASLite version that directly plugs into Tripal for better community curation within Tripal databases (2) BIMS: Further develop BIMS by adding more searching and analysis capability, with more testing by additional breeding programs, implement/further develop as needed a Tripal API version of BrAPI to allow breeders to use the various Breeding Apps being developed with BIMS, (3) Keep the NRSP10 databases current with tools and data while serving as community communication portals for their respective communities (4) Continue contributing to development of Tripal Core and Tripal Extension Module Development while providing advocacy and support for Tripal database adoption and module development by the developer community (help desk, web site, hackathon, monthly meetings) (5) Further participate in database advocacy by contributing to the publication of the Agricultural Biological Database (AgBioData) whitepaper, contributing to the NSF Data Repository Sustainability Process Guide and participating in other database outreach activities of the AgBioData Consortium (6) present the outcomes of NRSP10 through peer-reviewed publications, conferences, meetings and training activities and specifically develop outreach materials to inform industry of NRSP10 activities and their contribution to basic, translational and applied research. <br><br>

Impacts (unparsed)

<br>What was accomplished under these goals? Tripal Progress (Main, Ficklin, Staton and Wegrzyn Programs): Major progress for Tripal in Yr3 includes releases of Core Tripal (v2.1, v3.0.rc1, v3.0.rc2); and release of several Tripal Extension modules. These included modules to load, search and display sequence, map, marker, QTL, genotype, phenotype and germplasm data and implementation of the popular search engine Elasticsearch for the chado database tables that has now enabled cross-database site querying capability; Tripal Analysis Expression modules and Tripal Galaxy and associated workflows. There are now over 31 Tripal extension modules available to use or test from the Tripal.info site, with many more under development including the Tripal MapViewer, Tripal Synteny, Tripal Galaxy etc. Outreach efforts provided for Tripal include monthly conference calls, 2 day hackathon, Tripal workshop held at Plant and Animal Genome Conference, 258 correspondences through the Tripal support mailing list, ), keeping the Tripal website (https://tripal.info) current and providing programming support and code to other Tripal groups. Tripal is now being used for more than 100 species/clade/project databases. GenSAS Progress: GenSAS v5.0 released in January 2017. Major improvements completed include(1) ability to upload sequences before project creation so sequence subsets can be created from multiple-sequence fasta files. Sequence subsets can also be filtered by sequence name or minimum size (2) ability to upload RNA-seq reads use them to train gene model predictors (Augustus, during structural step) (3) addition of the tool Tophat to enable alignment specifically for the RNA-Seq reads. Alignment tools (blast, blat, PASA, tophat) are now included as a step before structural annotation (4)structural annotation (previously labelled as the "genes" step) now also has GeneMark for prokaryotic and eukaryotic gene prediction and (5)an official genes step (OGS) was added where users can either use EvidenceModeler to create a genes consensus and use it as the OGS, or can select another data track (from gene prediction program or alignment) as the OGS. Manual curation is merged with the OGS at the end of the GenSAS protocol. GenSAS was demonstrated/presented at several conferences. In Year 3 GenSAS was accessed by 1,779 visitors from 76 countries, with 6,434 sessions and 27,566 pages viewed. Breeding Tools Progress: In year 3 of the NRSP10 we continued development of the Tripal Breeding Information Management System (BIMS) and the phenotype data collection tool FieldBook App (Poland Program). New functionality includes the ability to view and download breeding data by cross population; generate trait and field files as input files for the Field Book App; ability for users to upload exported trait and field files to the database from excel templates or the FieldBook App; improved error checking for trait evaluation data to prevent outliers and flag invalid data prior to uploading; ability to create lists and generate statistics on them; and enhanced search stock functionality. We currently have peach and cotton breeding program data loaded in BIMS with functionality under testing by the respective breeding programs. Development on Field Book has primarily focused on adding user-requested features and patching user-reported bugs. A new trait format, 'Location', was added to facilitate collection of location point data and ability to add traits as a multi-trait category. A button was added to the main screen for missing values to help breeders distinguish between missing data and missing entries. Users can now load files directly from Dropbox, eliminating a file transfer step and streamlining the data collection process. Photos now also include the name of the trait to help researchers know better what they're looking at. A dedicated Android programmer was hired in the Poland lab in January 2017 to work on rewriting parts of the apps to fix bugs, increase efficiency, and better-adhere to best programming practices. Handheld Samsung tablets with Field Book have been provided to more than 50 NRSP10 associated breeders and allied researchers to test and use. A Field Book App and BIMS webinar was held in November, 2016 and FieldBook and BIMS were presented at several conferences. NRSP10 Databases: In year 3 we completed the upgrade of the Genome Database for Vaccinium, all five databases are now current in Tripal 2.1 and Drupal 7. CSFL and GDV have been converted to the next version, Tripal 3, on the development server and are being tested. Significant new data and functionality has been added to all five databases (documented in the work completed section of each database) and CGD has been expanded to include information specifically relevant to HLB research. Usage of the databases continues to grow and in 2017 was as follows: GDR - 21,845 users from 157 countries, 392,245 pages viewed; CottonGen - 11,328 users from 138 countries, 212,204 pages viewed; CSFL - 3,306 users from 110 countries, 34,475 pages viewed; CGD - 5,410 users from 127 countries, 76,653 pages viewed; and GDV ­­­- 2,265 users from 91 countries, 21,781 pages viewed. <br><br><b>Publications</b><br>

Outputs

Target Audience
The target audience is predominantly scientists, bioinformaticists, database developers/curators, both national and international, as well as, U.S. industry stakeholders from the 25 crops covered in this project. Scientists have been engaged through peer-reviewed publications, workshops, presentations at scientific conferences and meetings, newsletters to the mailing list, an online survey of users, emails about the website, and meetings with advisory groups. Bioinformaticists and database developers have been engaged through participation in monthly meetings on Tripal software, release of new versions of Tripal, the tripal.info website, a one day in-person hackathon, monthly Agricultural Biological Database (AgBioData) group meetings, peer-reviewed publications, and presentations at conferences and meetings. Industry stakeholders are being engaged through presentations and interactions at commodity group meetings.

Changes / Problems
Nothing Reported

Training & Professional Development
Training opportunities for undergraduate, graduate and postdoctoral researchers included participation and presentation at workshops, conferences and meetings as well as manuscript writing and participation in courses.

Dissemination Streams
Various components of NRSP10 project were presented at more than 25 conference/workshops/meetings/groups in Year 2. Major ones include the (1) International NAPIA biannual conference in November 2015; (2) Cotton Breeders Workshop at the Cotton Beltwide Conference in January 2016; (3) International Plant and Animal Genome Conference in January 2016 - presentations and demonstrations of individual databases, GenSAS and Tripal; 2nd Tripal workshop, database and GenSAS brochures and posters, plant genome database booth; (4) RosBREED participants meeting in March 2016; (5) World Cotton Conference and International Cotton Genomics Initiative Conference in May 2016; (6) 5thInternational Conference on Quantitative Genetics Conference in June 2016; (7) 8th International Conference on Rosaceae Genomics in June 2016; (8) American Society for Horticultural Science Annual Conference in August 2016; (9) 8th International Strawberry Symposium in August 2017; (10) AgBioData meetings; (11) Meta data ontology working group meetings; (12) Bioinformatics for Research graduate class at WSU (13) Peer-reviewed publications; and (14) Steering committee meetings. During the reporting period the databases were accessed as follows (as recorded by Google Analytics): GDR - accessed by 19,272 users from 152 countries, who viewed 232,162 pages; CottonGen - accessed by 9,422 users from 139 countries, who viewed 134,123 pages; CSFL - accessed by 2,982 users from 114 countries, who viewed 22,162 pages; CGD - accessed by 4,516 users from 123 countries, who viewed 24,064 pages; GDV- accessed by 1,349 users from 77 countries, who viewed 6,116 pages.

Next Reporting Steps
Continue to add new genomics, genetics and breeding data as they become available for each of the databases; go live with CGD and GDV update and redesign; release GenSAS v5.0; release Tripal v3.0, release new Mainlab Tripal extension modules; complete implementation of new Tripal RefTrans analysis for all databases; complete development of new Tripal Map extension module; continue developing Tripal BIMS; hold an NRSP10 workshop at ASHS in Sept 2017; continue providing Tripal support; continue actively participating in GGB database efforts at the national and international level including organizing the two day AgBioData conference for the database community to develop a white paper on Agricultural Biological Database data, code and communication sharing opportunities; hold database and fieldbook training sessions and online webinars; and continue to be very proactive in presenting NRSP10 activities/products at conferences, workshops and meetings. <br><br>

Impacts (unparsed)

<br>What was accomplished under these goals? Tripal Progress: In year 2 the major achievement in core Tripal development was the release of an alpha version of Tripal v3 in Jan 2016, set up of a demo site at http//demo.tripal.info/3.X for user exploration, and provision of a demo for stakeholders at the July Tripal users's conference call. Web service implementation is underway in Tripal 3 which makes Tripal back end database schema independent and helps overcome some of the limitations of having to use a chado database schema. This provides more flexibility for Tripal use and extends its utility to other genomic, genetic and breeding databases that do not want to be reliant on using Chado. In addition to work on version 3, we worked on finalizing extension modules for Tripal v2 which were developed to be easily upgraded to v3. These extension modules which were released in November 2016 included a Chado Loader, Chado Data Display and Chado Search. These modules provide site developers with tools 1) to collect and upload data 2) to organize and display data and 3) to enable advanced search functions. Supported data types include organism, marker, QTL, Mendelian Trait Loci, germplasm, map, project, phenotype, genotype and their associated metadata. The Chado Loader module provides data collection templates and PHP loaders to enable site developers to collect and store various types of data that are required to build comprehensive genomic and genetic databases. The Chado Data Display module contains a set of Drupal/PHP templates which can be used as is or customized as desired. The Chado Search module provides comprehensive search and download functionality for sequence, gene, marker, map, QTL, phenotype, genotype and germplasm. Also included are the tools to build map data and species summary pages. The use of materialized views in the Chado Search module enables better performance as well as flexibility of data modeling in Chado, allowing existing Tripal databases with their data stored in different ways to utilize the module. These Tripal Extension modules are implemented in the Genome Database for Rosaceae (rosaceae.org), CottonGEN (cottongen.org) and Cool Season Food Legume Database (coolseasonfoodlegume.org). Other ongoing Tripal extension work includes development of a MapViewer called TripalMap, designed to replace CMap. Outreach efforts to aid community development and adoption of Tripal included: Monthly Tripal developer meetings (includes international attendees), organizing the first Tripal hackathon and the 3rd Tripal workshop at PAG 2016 , keeping the Tripal website current and providing programming support and code to other Tripal groups. At the end of year 2, Tripal was being used by 90 databases. GenSAS Progress: GenSAS v4.0 was released in January 2016. Major improvements completed in year 2 include; the ability to use high performance computing where jobs can be submitted to a cluster rather than just a single server (enhanced speed); functional annotation capability through the addition of the BLAST,InterProScan,Pfam,SignalP, andTargetP tools to help assign a function to the proteins encoded by the predicted gene models; as well as additional tools added for structural annotation, and integration with JBrowse and WebApollo. Major work started in year 2 for a release of GenSAS v5.0 in January 2016. It includes: (1) ability to upload sequences before project creation so sequence subsets can be created from multiple-sequence fasta files. Sequence subsets can also be filtered by sequence name or minimum size (2) ability to upload RNA-seq reads use them to train gene model predictors (Augustus, during structural step) (3) addition of the tool Tophat to enable alignment specifically for the RNA-Seq reads. Alignment tools (blast, blat, PASA, tophat) are now included as a step before structural annotation (4)structural annotation (previously labelled as the "genes" step) now also has GeneMark for prokaryotic and eukaryotic gene prediction and (5)an official genes step (OGS) was added where users can either use EvidenceModeler to create a genes consensus and use it as the OGS, or can select another data track (from gene prediction program or alignment) as the OGS. Manual curation is merged with the OGS at the end of the GenSAS protocol. It was demonstrated at several conferences (see results disseminated sectionIn Year 2 GenSAS was been accessed by 1,135 visitors from 73 countries, with 2,333 sessions and 7,850 pages viewed. Breeding Tools Progress: We received encouraging feedback from our 20 stakeholder breeders who had been testing or deploying the FieldBook App for collection of phenotypic data in their breeding programs. As a result, we decided to more formally collaborate with the FieldBook PI, Jesse Poland at Kansas State University, and divert funds (0.5 FTE of a developer's salary) in years 3-5 of NRSP10 to help expedite further development relative to the expressed needs of our breeders. Following an initial needs assessment for breeding information data management tool functionality in year 1, we began BIMS design and implementation in year 2. BIMS will allow breeders to store, manage and analyze all their breeding data within a secure portal that connects their private data with all the data and tools available in their community genomics, genetics and breeding databases. We focused on working with two representative breeding programs from the 25 crops covered in the NRSP10; the peach breeding program of Dr. Kensija Gasic at Clemson University and the USDA ARS cotton breeding program of Dr. Todd Campbell in South Carolina. To gain further breeder input, we demonstrated goals and design of BIMS and FieldBook at several venues in year 2 (see the results disseminated section). These included a dedicated cotton breeder's workshop we organized at the Cotton Beltwide conference in January 2016, the International Plant and Animal Genome Conference in January 2016, the RosBREED project participants meeting in March 2016, the International Rosaceae Genomics Conference in June 2016 and the American Society for Horticultural Science annual meeting in August 2016. By the end of year 2 the design of BIMS was decided and programming initiated. NRSP10 Databases: In year 2 we completed upgrade of the CottonGen and GDR databases to Tripal 2 and Durpal 7 with a more user-friendly design on the live site and have done the same for the Citrus (CGD) and Vaccinium (GDV) databases on the development site. We are in the process of updating the data in the latter two and will go live with the new design in year 3. Significant data and new functionality has been added to all five databases. These are documented in the work completed page of the databases. <br><br><b>Publications</b><br>

Outputs

Target Audience
The target audiences are genomics, genetics and breeding scientists, bioinformaticists, and database curators and developers.

Changes / Problems
Nothing Reported

Training & Professional Development
Participating students, staff and postdoctoral researchers were all given opporunities to attend and present their work at conferences,workshops and meetings

Dissemination Streams
Various components of NRSP10 project were presented at 15 meetings/groups in 2015 (1) International Animal Plant and Genome Conference - presentations and demonstrations of individual databases, GenSAS and Tripal; inaugural Tripal workshop, database and GenSAS brochures and posters, plant genome database booth; (2) American Institute of Biological Sciences Complex Data Integration Workshop; (3) International Rubus and Ribes Symposium; (4) American Society for Horticultural Science Meeting; (5) Galaxy Community Conference (6) National Association of Plant Breeders Meeting; (6) NRSP10 Breeders Workshop; (7) WA Apple Genomics, Genetics and Genomics; (8) WA Apple Review; (9) US Legislative and Congressional Staff Meeting at WSU (10) United Kingdom Consulate Meeting at WSU; (10) Tripal database developer meetings; (11) AgBioData meetings; (12) Meta data ontology working group meetings; (13) Bioinformatics for Research graduate class at WSU (14) Peer-reviewed publications; and (15) Steering committee meetings. During the reporting period the databases were accessed as follows (as recorded by Google Analytics): GDR - accessed by 18,089 users from 145 countries, who viewed 196,562 pages; CottonGen - accessed by 19,760 users from 142 countries, who viewed 102,549 pages; CSFL - accessed by 2,977 users from 126 countries, who viewed 17,317 pages; CGD - accessed by 4,262 users from 117 countries, who viewed 21,961 pages; GDV- accessed by 1,324 users from 73 countries, who viewed 6,444 pages.

Next Reporting Steps
Continue to add new genomics, genetics and breeding data as they become available for each of the databases; migrate remaining databases to Tripal2; release GenSAS v4.0; release Tripal v3.0, release Mainlab Tripal extension modules; finish converting CottonGen, CSFL, GDR CGD, GDV to Tripal 2 with new theme implementation on live site; complete development of new Tripal RefTrans analysis workflow and extension module; complete development of new Tripal Map extension module; continue to be very proactive in presenting NRSP10 activities/products at conferences, workshops and meetings. <br><br>

Impacts (unparsed)

<br>What was accomplished under these goals? ore Tripal Progress: In year 1 the major achievement in core Tripal development was the release of Tripal v2.0-rc1 in Sept 2014 followed by v2.0 on June 1, 2015. Improvements in v2.0 included: (1) Improved ease of finding extension modules by addition of an interface in Tripal to directly query the tripal.info site and find available modules; (2) improved performance by fixing the memory leaks in the loaders; (3) updated the GFF loader to support automatic creation of protein sequences for genes (4) created an initial implementation of web services based on Chado structure then decided to change direction to make Tripal database backend independent (released in year 2 of the project) and therefore widen access to use Tripal for databases who use a different database schema and help resolve possible performance issues using Chado with very large datasets; (5) fixed bugs as they were found; (6) Improved administrative pages within Tripal; (7) regularly updated the Tripal website (http://tripal.info); (8) Held the first Tripal Workshop at the 2015 International Plant and Animal Genome Conference; (9) Held monthly Tripal developer meetings; (10) provided programming support and code to other Tripal groups. Breeding Tools Progress: In year 1 the major achievement in breeding tools proress was indidntifying NRSP10 breeders database needs. A NRSP Breeders Database Needs Assessment Workshop was held in July where representative breeders (invitation only) of the major NRSP crops and industry representatives were presented with ideas for development of a more comprehensive breeding information management system called TripalBIMS. Building on the GDR and RosBREED project breeding tools TripalBIMS will integrate all aspects of breeding program management, public and private data integration, with analysis capabilities from a single portal, to aid in breeding program decision-making. For more efficient collection of phenotype data the FieldBook App is being evaluated by several NRSP10 crop breeders. The feedback so far indicates FieldBook is an excellent resource but will need some modifications (working with the developer from Jess Polands lab on these). Information gathered from this workshop, plus discussion held at the annual RosBREED participants meeting held in March 2015 are being used to develop TripalBIMS. GenSAS Progress: In Year 1 our major achievement regarding GenSAS: the Genome Sequence Annotation Server, was the release of GenSAS v3.0 in Decemebr 2014. GenSAS v3.0 has several major improvements over the previous v2.0 release; (1) interface redesign combining the project creation options and results browsing from v2.0 into one webpage. The new interface has a flow chart of the annotation process in the webpage header as well as a right-side menu for quick access to the job queue, JBrowse and project sharing. Users interact with GenSAS through a series of tabs that open in the center part of the screen. The redesign has improved the user experience by not only making it more intuitive, but also includes integrated instructions on how to use GenSAS; (2) Integration of JBrowse and Apollo for viewing results and manual curation of the annotation. By using an already developed genome browser and annotation editor, users can easily view and curate the data. Apollo also is configured for community annotation and works well with the collaborative annotation tools that were already a part of GenSAS; (3) Addition of Evidence Modeler as a tool in GenSAS v3.0 to allow for generation of a gene model consensus from multiple gene predictor tool outputs; (4) Addition of a GenSAS Administration page under the Drupal Administration Menu so administrators on the GenSAS website can monitor job activity, configure GenSAS settings, and add/remove/configure tools that are available for use in GenSAS; (5) Presentation of GenSAS at the 2015 Plant and Animal Genome Conference, 2015 American Society of Horticultural Science Annual Meeting, and the 2015 WSU Academic Showcase and created a 15 minute video tutorial for GenSAS. We also answered questions and fixed program issues for about 15 users through the year; (6) By August 31, 2015 we had 110 GenSAS v3.0 user accounts created. In March 2015 we also started measuring GenSAS site visits using Google Analytics. In the 6th month period to August 31, 2015, GenSAS was accessed by 595 visitors from 61 countries with 1,132 sessions and 3,969 pages viewed. Almost half of the site visitors (47.7%) were returning users (47.7%). NRSP10 Databases Progress: In year 1 our major achievements regarding the 5 NRSP10 databases was improving the data and functionality of the databases on the live sites with further improvements being tested on the development server. (1) Cool Season Food Legume Database (CSFL, https://www.coolseasonfoodlegume.org) In Year 1 the CSFL was fully converted to Tripal v2.0 on the development server and made current with genome, gene, trait, map, marker, germplasm, phenotype, genotype and publications data for pea, lentil, chickpea and fava bean. A new analysis pipeline was developed to combine all publicly available published RNASeq datasets (different platforms) with EST data into reference transcriptomes which were then functionally characterized and mapped. As part of the conversion to Tripal v2.0 a cleaner, more user-friendly interface has been developed which makes it easier to get to the data and tools. CSFL currently contains: 1 genome; 233,191 genes; 3,777 markers; 38 genetic maps; 215,713 transcripts; 2,255 Trait Loci; and 4,225 publications. (2) Genome Database for Rosaceae (GDR, https://www.rosaceae.org): In year 1 new functionality and dataadded tothe GDR included: 5 annotated genome sequences; SNP array data for apple, strawberry and rose and a new interface developed to access the SNP data; establishment of a gene database and a Tripal extension module developed to enable the genes to be searched by gene symbol, function, species, and genome location; establishment of a community agreed standard nomenclatures for gene designation in the Rosaceae and a template and a gene submission page; and map, marker and QTL data added. Several new interfaces were added to improve querying capability. GDR currently contains: 10 genomes; 233,191 genes; 2,193,827 markers; 123 genetic maps; 518,586 transcripts; 912 Trait Loci; 13,500 germplasm; and 6,319 publications. (3) CottonGen: In Year 1, the conversion of CottonGen (https://www.cottongen.org) to Tripal v2.0 was initiated on the development server, with design and implementation of a new theme providing a cleaner, more user-friendly interface which makes it easier to get to the data and tools. On the live site major new data and functionality in CottonGen included: two cotton genome sequences and gene models with additional annotation performed though our annotation workflow. The annotated sequences are searchable by species, name, function, location, mapped marker, viewable through GBrowse and JBrowse and made available to BLAST; 6000 SNP marker primers; and hosted ICGI elections. CottonGen currently contains : 5 genomes; 323,180 genes; 276,221 markers; 50 genetic maps; 442,954 transcripts; 988 trait loci; 106,664 genotypes; 130,756 phenotypes; 14,077 germplasm; and 15,494 publications. (4) Citrus Genome Database (CGD, https://www.citrusgenomedb.org): In Year 1 of this project map, marker and QTL data have been collected. GCD currently contains: 3 genomes; 2,192 genes; 3,427 markers; 51 genetic maps; 386,445 transcripts; 86 trait loci and 2,955 publications. (5) Genome Database for Vaccinium (https://www.vaccinium.org, GDV): In Year 1 of this project sequence, map, marker and QTL data have been collected. GDV currently contains: 261,775 transcripts and 402 publications. <br><br><b>Publications</b><br>


Publications Inventory

Conference Papers and Presentations

Websites

Journal Articles

Book Chapters

Other