Proteins perform almost all essential tasks inside human cells, but a single gene does not always produce just one standard protein. Cells generate functional and structural diversity through processes like alternative splicing, where a single gene's RNA instructions are cut and reassembled in different ways, much like editing different scenes to create distinct versions of a movie, producing protein variants known as isoforms. Furthermore, proteins undergo chemical modifications after synthesis, such as prenylation, where a lipid "anchor" is attached to a protein to lock it onto cellular membranes. Understanding these modified protein variants is crucial because changes in their location, stability, or function can dramatically alter cell behavior and disease progression.
To identify proteins in complex biological samples, researchers rely on mass spectrometry, a technique that acts like a molecular scale by weighing protein fragments (peptides) and matching them against a digital reference database. However, standard reference proteome databases contain vast amounts of duplicate and overlapping sequence data across different protein variants. Searching experimental mass spectrometry data against these bloated databases creates a severe computational and statistical challenge. When the search space becomes very large, identification algorithms can miss the protein variants that we are looking for.
To overcome this limitation, we developed a targeted database designed to isolate protein isoforms and prenylation events without inflating the search space. First, we built a non-redundant canonical database by combining identical human protein entries into unified records. Second, we constructed an isoform-specific database containing unique protein fragments (like a fingerprint) for human protein isoforms. Finally, we created a specialized post-prenylation processing database to investigate whether proteins that are prenylated at internal sites undergo molecular trimming after lipid attachment. By restricting our search space strictly to unique, informative sequence regions, our custom databases can be integrated directly into standard mass spectrometry workflows without sacrificing statistical sensitivity.
We evaluated these databases using mass spectrometry data previously collected from human T helper 1 (Th1) immune cells. The isoform-specific database incorporated 25,711 unique peptides representing 14,144 human protein isoforms, enabling the potential identification of over 63% of all annotated human isoforms. Applying this database to the Th1 cell dataset successfully identified 170 unique protein isoforms across all samples that standard workflows could not cleanly isolate. In contrast, investigating internal prenylation sites provided limited evidence for post-prenylation processing, yielding only a single candidate peptide (ACAA1) that likely represents an unannotated splice variant rather than true post-prenylation cleavage. Overall, this project demonstrates that non-redundant, target-engineered databases significantly enhance our ability to detect hidden protein isoforms while maintaining a manageable search space, offering a strong foundation for future studies exploring the diversity of the human proteome.