Renal cell carcinoma (RCC) is driven by mutations in cancer driver proteins, the majority of which are encoded by tumor suppressor genes. These drivers cluster into functional categories including protein ubiquitination, chromatin remodeling, transcription, and the MTOR pathway, and they frequently act through protein complexes rather than in isolation.
A significant fraction of cancer-causing mutations occurs precisely at the interfaces where proteins bind one another. When a mutation disrupts a protein-protein interaction (PPI), it can inactivate an entire protein complex, short-circuiting essential cellular processes like protein degradation and transcriptional regulation. Mapping these mutations to structural interfaces is therefore critical for understanding RCC pathogenesis.
Traditional large-scale experimental studies of PPIs suffer from high false positive and false negative rates, making it difficult to confidently identify which interactions are real. Deep-learning structure prediction methods based on AlphaFold now offer a way to screen candidate interactions computationally and generate high-quality structural models of the resulting complexes.
The researchers compiled a comprehensive set of RCC driver proteins from two publicly available cancer gene databases, NCG and OncoVar. Candidate PPIs involving these drivers were drawn from the BioGRID database and filtered using publication-weighted confidence scores, yielding 3,595 candidate protein pairs for computational screening.
Three AlphaFold-based prediction methods were applied to all candidate pairs: AF-contact (an in-house method using inter-residue contact probabilities), AF2Complex (which scores interface quality via Predicted Aligned Error), and FoldDock (which combines interface pLDDT scores with contact counts). Each method independently ranked the candidate pairs by predicted interaction confidence.
The three methods showed strong agreement at the top of their ranked lists, with 39 of the top 100 predictions shared across all three programs. Somatic mutations from the COSMIC database and a dataset of 148 RCC patients were then mapped to the predicted and experimentally confirmed interfaces, and binding affinity changes were computed using BindProfX with FoldX to assess functional impact.
AF-contact identified 211 high-confidence protein pairs (represented by 299 domain-pair models) with contact probability above 0.9. Among these, 85 pairs were already supported by experimental structures in the Protein Data Bank, serving as validation that the method recovers known interactions with high accuracy.
Across the top 50 predictions, AF-contact, AF2Complex, and FoldDock recovered 30, 30, and 29 experimentally confirmed pairs respectively, demonstrating comparable precision at the highest confidence tier. AF-contact showed a slight advantage among the top 100 predictions, identifying 60 PDB-supported pairs compared to 49 and 48 for the other two methods.
More than 100 cancer somatic mutations were found to affect binding affinity at PPI interfaces involving key RCC drivers. The most affected proteins were VHL and TCEB1, which together participate in the ubiquitin ligase complex responsible for targeting HIF-alpha for degradation, a pathway whose disruption is a hallmark of clear cell RCC.
VHL, the most frequently mutated gene in clear cell RCC, forms complexes with TCEB1 (ELOC), ELOB, and CUL2. Multiple deleterious VHL mutations were mapped to the SOCS-box domain at the VHL-TCEB1-CUL2 interface, as well as to the interfaces with HIF1-alpha and HIF2-alpha, the transcription factors that drive tumor angiogenesis and glycolysis when VHL function is lost.
TCEB1 was revealed as a broader protein interaction hub than previously appreciated. AlphaFold predicted interactions between TCEB1 and eight distinct SOCS-box-containing proteins, including VHL, SOCS1, ELOA, ASB1, PRAME1, and WSB1. Hotspot RCC mutations in TCEB1 such as Y79C and A100P mapped to these interfaces, suggesting that TCEB1 mutations could impair the degradation of multiple targets beyond VHL.
For the COMPASS-related chromatin remodeling complexes involving KMT2C, KMT2D, and KDM6A, AlphaFold identified multiple short linear motif (SLiM)-mediated interactions in disordered protein regions. Few missense mutations were found at these interfaces, consistent with the observation that inactivating mutations (nonsense, frameshift) rather than missense mutations tend to be the functionally dominant alterations in these large chromatin remodeling subunits.
NRF2 (NFE2L2) is a transcription factor that drives stress response gene expression and is classified as an oncogene in RCC. It is normally degraded through its interaction with KEAP1, which bridges NRF2 to a CUL3-based E3 ubiquitin ligase complex. Mutations in NRF2 at residues Asp29, Leu30, and Gly31 were predicted to destabilize the NRF2-KEAP1 interface and prevent NRF2 from being degraded, enabling it to drive cancer progression.
AlphaFold predicted the NRF2-MAFK interaction with very high confidence (contact probability above 0.99). The structural model revealed that NRF2 and MAFK form a leucine zipper heterodimer through coiled-coil interactions, with the C-terminus of NRF2 contributing an additional beta-hairpin contact with MAFK's coiled coil. This beta-hairpin interface had not been previously characterized experimentally.
No destabilizing cancer mutations were found in the NRF2-MAFK interface, consistent with the idea that this interaction is actively maintained to support NRF2 transcriptional activity in cancer. The newly discovered NRF2 C-terminal beta-hairpin interface with MAFK represents a novel structural target for developing small-molecule inhibitors that could disrupt NRF2-driven transcription in RCC.
TRRAP, a large scaffolding protein involved in histone acetyltransferase complexes including SAGA and TIP60, was found to interact with multiple proteins that have established roles in chromatin regulation. AlphaFold predicted interactions between TRRAP and several known complex partners, providing structural models for interfaces that had not been previously resolved experimentally.
In the MTOR pathway, AlphaFold predicted interactions among MTOR, RPTOR, MLST8, and DEPTOR that were consistent with established experimental structures. An additional novel interaction was predicted between TSC1 and DYRK1A, a kinase known to regulate the mTOR pathway through the TSC complex. The structural model showed that TSC1's disordered region contributes two alpha helices that wrap around the C-lobe of the DYRK1A kinase domain.
A short sequence motif in TSC1 (residues 342-360) was predicted to interact with TSC2's HEAT repeat domain with high confidence, revealing a second TSC1-TSC2 interaction site beyond the previously known coiled-coil interface. This motif contains highly conserved residues W347, P349, and C353, suggesting it serves an important functional role in the TSC complex and could represent another interface of interest for therapeutic intervention.
By systematically mapping over 100 RCC somatic mutations to PPI interfaces, this study provides a structural framework for interpreting the functional consequences of specific mutations identified in patient tumors. Mutations predicted to destabilize complexes by more than 1.4 kcal/mol were highlighted as likely to have significant biological effects, giving clinicians and researchers a way to prioritize variants for further study.
The concentration of deleterious mutations at VHL and TCEB1 interfaces underscores the central importance of the HIF-alpha degradation pathway in RCC biology. Conversely, the low density of missense mutations at chromatin remodeling complex interfaces suggests a different mode of disruption for that functional category, with therapeutic implications for how each class of driver might best be targeted.
Targeting protein-protein interaction interfaces pharmacologically is an active area of drug discovery. The structural models generated in this study provide starting points for structure-based drug design campaigns, particularly for the NRF2-MAFK leucine zipper interface, which could be disrupted with peptide mimetics or small molecules to reduce NRF2 transcriptional activity in RCC.
A key limitation of this study is that AlphaFold does not incorporate post-translational modifications such as hydroxylated prolines, which are required for VHL-HIF interactions. This caused those interactions to fall below the confidence threshold despite being well-established, indicating that the method has blind spots for interactions that depend on chemically modified residues.
The candidate PPI list derived from BioGRID still contains false positives despite the publication-weighted filtering, and the computational predictions of mutation effects on binding affinity using BindProfX represent estimates rather than experimentally validated measurements. Future work requires direct biochemical and structural validation of the predicted complexes and mutation effects.
Future directions include extending the analysis to other cancer types and integrating these structural PPI models with functional genomics data to better understand how network-level disruptions of multiple protein complexes cooperate to drive RCC progression. Large-scale prospective mutation mapping across more patient cohorts would also refine understanding of which interface mutations are clinically most significant.