Grant Information
| Knowledge Area | Subject of Investigation | Field of Science | Percent |
|---|---|---|---|
| 212 - Pathogens and Nematodes Affecting Plants | 920 - Orange | 1090 - Immunology | 100% |
There is a high priority, yet unmet need for improved microbe resistance in crops to ensure food security. Plant immunity against pathogenic microbes is strongly impacted by the specificity of pattern recognition receptors such as FLS2. Our lab is particularly well suited to address this challenge. Our extensive protein engineering expertise in high-throughput cell surface display, library design, directed evolution, and machine learning enables us to capture plant-pest interactions with unprecedented detail. With the advantage of generating mutagenic libraries with up to 109 unique variants, our approach to narrow down and design improved immune receptors will ease the burden of characterization, leading to a wealth of new knowledge of plant-pathogen interactions that will shape how we engineer sustainable crops in the future.
Efforts and Evaluation:We willcharacterize FLS2 function in yeast. We recently validated an FLS2 expression platform in yeastthat enables the characterization of large libraries with >109 unique FLS2s. This will consist of cloning libraries of FLS2s into yeast, then characterized and sorted via flow cytometry using a panel of fluorescently labelled pathogenic peptides (flg-X). Numerous synthetic peptides with diverse amino acid length and species-specific detection have been identified as PAMPs for a variety of crops. Each flg-X will be individually screened against the entire population of FLS2 variants. Binding kinetics of the ligand to FLS2 variants will be measured via affinity titration curve. The rank order binding affinity will be efficiently quantified for all FLS2 variants simultaneously using the MassTitr approach by binning FLS2 variants according to binding intensity at multiple flg-X peptide concentrations. Each bin is separately deep sequenced for downstream analysis and training machine learning models. Our lab has extensive experience conducting high-throughput screens of yeast displayed libraries, analyzing deep sequencing data, and developing machine learning models.Compile 'labeled' FLS2 sequence data by isolating FLS2 variants functional against a panel of pathogens. The NLP and deep learning models will be trained for the purposes of classification (e.g. Does the FLS2 sequence detect flg-X? yes/no) and regression (e.g. How well does the FLS2 sequence detect flg-X, on a continuous scale from 0-100?). To accommodate for classification, a single gate can be drawn during cytometry sorting so that all events with positive fluorescence are isolated and labeled as 'good' sequences. Alternatively, kinetic parameters for binding selectivity can be estimated for each individual FLS2 variant using a MassTitr approachwhereby the population of FLS2 expressing cells is incubated at a range of peptide concentrations, then sorted into multiple gates followed by deep sequencing with barcoding for each cytometry gate and peptide concentration. Train machine learning models on the labeled receptor sequences. Yeast surface display experimentsprovidelabeled data for training machine learning algorithms. This facilitates going beyond the obtained data and subsequently understanding the design principles for functional FLS2. By applying generative models, we identify features within the amino acid sequences that are critical for FLS2 function. This enables us to generate novel and improved sequences by leveraging the lessons stored in our optimized models. Importantly, we will implement two distinct and complementary approaches: variational auto-encoder (VAE) and a de-noising transformer (BART). The two approaches taken for generation tasks offer unique advantages as VAE is probabilistic[61] and BART is a de-noising model for improving 'bad' sequences. VAE consists of the encoder (E) and the decoder (D) to generate new datasets. It minimizes the error between initial data, X, and encoded-decoded data, D(E(X)). In BART, we corrupt the labeled data and minimize the loss in reconstructing the correct data. The minimized loss has learned the signal and will be able to generate de-noised sequences.Validation of newly discovered FLS2 variants in plant hosts. The ability of FLS2 to prompt an immune response upon binding flg22-X peptides will be evaluated in plant hosts. We will measure metabolic changes associated with defense response following agrobacterium-mediated Arabidopsis cloning. We will assay candidate-transformed plant material for seedling growth inhibition, ethylene production, and callose deposition. These responses occur upon formation of the FLS2-flg-BAK1 PRR signaling complex, known as Pattern-Triggered Immunity (PTI) response.Seedling growth inhibition via flg-X responsewill be performed by monitoring seedling weight throughout their growth period.Ethylene production via flg-X responsewill be performed using gas chromatography for analysis. Briefly, 1 mm wide leaf strips of matured plants are floated on water, then multiple strips are chosen at random for each replicate, floated, and incubated with flg-X prior to gas-tight capping and evaluation of ethylene concentration after 5 hours.Callose deposition via flg-X responsewill be performed by measuring UV epifluorescence of flg-X treated seedlings.Train generative models on labeled data to find new-to-nature FLS2 sequences. While variational autoencoders (VAEs) are limited in that they must be trained exclusively on 'good' (functional) sequences, BART is able to accept both 'bad' (non-functional) and 'good' sequences as input. Recall that we will have already produced a large collection of both bad and good sequences during Objective 1 wet-lab experiments. BART has the ability to learn sequence epistasis signaturesthat are required among good sequences during pre-training, then trains the de-noising function. Finally, BART applies de-noising over the bad sequences to generate improved sequences that have a higher probability of being functional. BART enables a broader and more efficient search over protein sequence space.Rank and filter the generated sequences with computational methods. To validate sequence-to-structure consistency, the de novo FLS2 sequences will be modeled in AlphaFold, discarding any structure that contains any obvious misfolding. Then, embedding methods will be implemented and fine-tuned over natural FLS2 sequences in order to score the generated sequences. At this point, computationally validated 'good' sequences will be re-fed to the VAE and BART models with continued iteration yielding improved model performance.Characterization of chosen sequence candidates in wet-lab experiments as described in Objective 1. Engineered flg-Xs identified that were poorly recognized by FLS2 variants will be of particular interest in this activity for the discovery of novel detection. The expected outcome is that this evolutionary generative model will arrive at distinct solutions relative to the VAE or BART models independently.
Target Audience
The target audience for this research includes individuals fromacademic research groups and scientific research institutions. Within these larger organizations, individualsinterested inapplying protein engineering, machine learning, computational protein design, heterologous host expression systems, and high-throughputanalysisare of particular interest. Specifically, scientists and research personnelthatfocus on the improvementof plant defense proteins for sustainable agriculture, pest control, and future pathogen mitigation have the most to gain fromthe findings and products obtained from thesebroadly-describedapproaches. The outcomes of this research are beneficial to this group by providing a proof of concept and evidence-based experimental outcomes to guide future research.
Changes / Problems
Although our preliminary evidence indicated that we could utilize our high-throughput approach for detecting the protein's interactions with pathogenic molecules, multiple follow-up experiments determined that our yeast surface display system's native yeast biology will require further modifications for producing necessary plant-like post-translational modifications (e.g. glycosylation). This difference in glycosylation pattern greatly hindered protein interaction with pathogenic ligand, specifically that large quantities of pathogenic ligand were necessary to detect interactions using high-throughput approaches. To tackle these challenges, we explored a two-pronged approach where we either (1) changed our expression system or (2) identified plant immune proteins which are more suitable for our systems approach. In the first case, we developed a collaboration with a group at MSU that uses the algae Chlamydomonas reinhardtii for recombinant protein expression. These algae, like plants, are also photosynthetic organisms with similar post-translational modifications and glycosylation patterns. We are currently in the process of establishing recombinant algae with our original protein of interest (FLS2) that utilizes a recently described algal surface display system. In the second case, candidate plant defense proteins were identified that were less complex than our original protein of interest and their recombinant expression in yeast had been described in the literature. Importantly, these proteins are important players in plant defense with sequence diversity across crop species and have variable activity against pathogenic targets. We are currently in the process of evaluating these candidate proteins' activity against their respective pathogenic targets.
Training & Professional Development
This project has opened multiple avenues for graduate and undergraduate development in both experimental, communication, and professional areas. During the project's funding period, undergraduate students participating in achieving project goals learned a range of wet and dry lab skills. These include utilizing software for protein prediction and analysis, cloning and maintenance of heterologous hosts for proteins of interest, high-throughput analysis, and general scientific theory. With these skills as a foundation, undergraduate and graduate students used experimental data to produce multiple conference talks and poster presentations while catering to broad audiences. Both graduate and undergraduate students used their experimental and communication skills developed throughout the course of the funding period to engage in relevant professional student organizations or fellowship programs. These research outcomes also bolstered undergraduate and graduate resumes which helped students achieve offers and participate in internships at biotech companies such as Merck and Zoetis.
Dissemination Streams
Personnel involved in this project have produced multiple conference presentations and posters reaching broad audiences. Mainly, posters prepared by undergraduate students have been presented at MSU undergraduate research showcases such as UURAF, EnSURE, CMSE Symposium, and CHEMS Research Forums. Undergraduate students have also presented posters at conferences, including the 2024 BMES Annual Meeting. Additionally, our graduate personnel presented a talk for the 2023 PBHS Training Program Retreat. From the development of our machine learning architecture, two peer-reviewed publications were produced during the funding period and are freely accessible (open access). These avenues for communication of project results have reached audiences including plant biotechnologists, biomedical engineers, chemical engineers, MSU professors, and students.
Next Reporting Steps
To make use of machine learning architectures developed during the reporting period while taking advantage of our lessons learned from the limitations encountered, we are evaluating a panel of candidate plant defense proteins which would be suitable for high-throughput analysis in our yeast surface display system. Currently, we have confirmed that our candidate proteins are expressed in yeast and have produced pathogenic target proteins for binding analysis to be conducted by flow cytometry. Concurrently, we are developing variant libraries for our candidate proteins which we plan to deploy against the pathogenic targets using our approach described in the original project narrative. Next-generation sequencing data acquired from these analyses will be labelled based on activity against a target and used as input for developing machine learning models for improved versions of our candidate proteins. Importantly, once we have confirmed the activity of our candidate proteins against the pathogenic targets the project approach will continue as described in the original project narrative.