Skip to content

Latest commit

 

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

H²GNN: Hierarchical Hypergraph Neural Networks for Multi-Type Drug-Drug Interaction Prediction

Paper Project Page Python Hardware Colab

A hierarchical hypergraph deep learning framework that predicts whether two drugs interact and which of 86 interaction types they share, using only SMILES strings as input. The framework is built from two hypergraph neural networks: a Chemical HGNN that learns drug embeddings from molecular substructures, and an Interaction-Type HGNN that reuses those embeddings to classify the interaction mechanism. Both networks train and run on a two-core CPU without a GPU.

Scope of this README. This document covers the two neural networks only. The cascaded Clinical Decision Support System that chains them together, including the Clinical Reporting Module, is documented in its own directory.

📋 Table of Contents


Research Contributions

  1. A unified hierarchical hypergraph (New framework) that serves both binary interaction detection and 86-class type classification, instead of treating them as two unrelated tasks.
  2. An interaction-type hypergraph (Novel method) in which drugs are nodes and the 86 DrugBank interaction types are hyperedges, framing mechanism classification as hypergraph learning rather than flat multi-class prediction.
  3. An embedding-transfer mechanism (New discovery) that initializes the type hypergraph with the chemical representations learned in the binary stage, replacing random initialization or hand-crafted fingerprints. Its value is quantified in the first ablation.

The two networks are also deployed together as a cascaded clinical decision support system, documented separately from this directory.


Overview

Polypharmacy substantially raises the risk of drug-drug interactions (DDIs). Most graph-based DDI predictors treat interaction detection and interaction-type classification as separate problems, and many require heavy hardware and hand-engineered molecular features. H²GNN addresses both issues with a two-level hypergraph hierarchy:

  1. Low-level chemical plane (Chemical HGNN): each drug is a hyperedge connecting the SMILES substructures (k-mers) it contains. The network performs binary classification (interaction / no interaction) and learns a 128-dimensional embedding for every drug.
  2. High-level interaction plane (Interaction-Type HGNN): each interaction type is a hyperedge connecting all the drugs that participate in it. Drug nodes are initialized with the embeddings transferred from Stage 1, and the network performs multi-class classification over 86 interaction types.

Hyperedges let a single relation join many entities at once, capturing higher-order drug–substructure and drug–type relationships that pairwise graphs and fixed fingerprint vectors cannot encode.

Hierarchical hypergraph: interaction-type plane above, chemical-substructure plane below

Key Achievements

Achievement Value
🎯 Binary ROC-AUC (Chemical HGNN) 98.46%
🎯 Binary Accuracy (Chemical HGNN) 93.20%
🎯 Multi-class ROC-AUC (86 types) 99.44%
🎯 Top-1 / Top-3 Accuracy (86 types) 85.73% / 98.11%
🔁 5-Fold Cross-Validation Top-1 85.63 ± 0.16%
🔄 Recovery of Withheld Secondary Labels (Top-3) 94.7%
📈 Gain over Best Baseline +5 pts Accuracy (binary), +22 pts Macro F1 (multi-class)
⚡ Training Time 142 min (Stage 1), 18.6 min (Stage 2)
💾 RAM Increase During Training ≤ 0.46 GB
🖥️ Hardware CPU only (2 vCPUs), no GPU

Repository Structure

HHGNN/
│
├── Chemical/               # Stage 1: Chemical HGNN experiments (Colab notebooks + outputs)
├── Type-Interaction/       # Stage 2: Interaction-Type HGNN experiments (Colab notebooks + outputs)
├── DataSet/                # Merged DrugBank dataset and entity-level metadata files
├── DataSet-Partitions/     # Seeded train / validation / test splits
├── Hypergraph_Chemical/    # Chemical hypergraphs for k = 3, 6, 9, 12 and index maps
├── K-mer/                  # K-mer decomposition files (one per window size)
├── Images/                 # Figures used in this README
└── README.md

Main Data Artifacts

Artifact Content Used By
Merged interaction file Drug1_ID, Drug2_ID, Type-Interaction ID (191,877 pairs) Both networks
Drug–SMILES file DrugID, SMILES for all 1,709 drugs K-mer decomposition
Drug–Type of Interaction file DrugID, Interaction Types (semicolon-separated) Interaction-Type hypergraph
K-mer files (×4) Drug_ID, Segmented_SMILES for k = 3, 6, 9, 12 Chemical hypergraph
Hypergraph files Connection list + node_to_idx, drug_to_idx, type_to_idx Both networks
Drug embeddings E-1, E-2 1,709 × 128 matrices from the two best Chemical HGNN runs Interaction-Type HGNN
Metadata files Drug names and the 86 DrugBank interaction-type descriptions Result interpretation

Naming note: earlier versions of this project called the Interaction-Type HGNN the Metabolic network and the framework MLHGNN. These names refer to the same models.

Dataset and Preprocessing

Two DrugBank benchmark sources were combined:

Field DrugBank-KAIST DrugBank-TDC Merged Dataset
Interaction records 192,283 191,808 192,283
Unique drugs 1,709 1,706 1,709
Interaction types 86 86 86
SMILES coverage None 1,706 drugs 1,709 drugs

DrugBank-TDC is fully contained in DrugBank-KAIST. The three KAIST-only drugs (Colestipol DB00375, Butriptyline DB09016, Phenoxyethanol DB11304) have no SMILES in either source; their SMILES were obtained through collaboration with the DrugBank Foundation, giving 100% molecular coverage.

Cleaning. No missing values and no exact or inverse duplicates were found. 406 drug pairs carried more than one interaction type; the first type was kept and the others were stored in an auxiliary file (used later in the label-recovery ablation). The final dataset holds 1,709 drugs, 191,877 unique pairs, and 86 interaction types.

Class imbalance. The type distribution is strongly long-tailed and was deliberately left uncorrected to preserve its pharmacological character:

Frequency Band Types Records Share
Very high 3 118,870 61.95%
High 15 57,973 30.21%
Medium 35 13,760 7.17%
Low 31 1,261 < 0.7% combined
Very low 2 13

Splitting. An 80/10/10 entity-aware split guarantees that every drug (Chemical HGNN) and every interaction type (Interaction-Type HGNN) appears in the training set. A four-task leakage check returned zero overlapping pairs on every seed.

Network Seed Training Validation Test
Chemical HGNN 32 / 42 / 46 153,501 19,187 19,189
Interaction-Type HGNN 32 153,495 19,188 19,194
Interaction-Type HGNN 42 153,489 19,188 19,200

Model Architecture

Architecture of the Chemical HGNN and the Interaction-Type HGNN

Chemical HGNN (Stage 1). A single hypergraph attention layer operates on the chemical hypergraph, with drugs as hyperedges (one-hot features) and substructures as nodes (all-ones features). Attention flows at the hyperedge level and then the node level, producing a 128-dimensional embedding per drug. An MLP decoder combines the embeddings of the two drugs in a pair and passes them through two fully connected layers with ReLU and a sigmoid output (interaction: 0/1).

Interaction-Type HGNN (Stage 2). A single hypergraph attention layer operates on the interaction-type hypergraph, with drugs as nodes initialized from the transferred Stage 1 embeddings and interaction types as hyperedges (one-hot features). Attention flows hyperedge → node → hyperedge so that the type hyperedges, the prediction target, hold the final aggregated representation. The decoder outputs a softmax distribution over the 86 interaction types.

The drug_to_idx map from the chemical hypergraph is reused unchanged in Stage 2, so every drug keeps the same index in both hypergraphs and the embedding transfer stays aligned.

Experimental Setup

Component Specification
Platform Google Colab Free Tier
CPU Intel® Xeon® @ 2.20 GHz, 2 vCPUs
RAM 12 GiB
GPU None (all training and evaluation on CPU)
Python 3.12
Setting Chemical HGNN Interaction-Type HGNN
Runs 36 (4 k-mer sizes × 3 seeds × 3 configs) 12 (2 seeds × 2 embeddings × 3 configs)
Seeds 32, 42, 46 32, 42
Max epochs 500 500
Early-stopping patience 200 100
Checkpoint Lowest validation loss Lowest validation loss

The Colab Free Tier was chosen because it gives an identical environment for every run and a resource ceiling close to the modest hardware found in hospitals and pharmacies.


Stage 1: Chemical HGNN

Hypergraph Construction

Chemical hypergraph construction from SMILES k-mers

Each SMILES string is decomposed with a sliding window of size k. Each window size yields its own vocabulary and its own hypergraph:

k Unique k-mers (Nodes) Avg. k-mers per Drug Hyperedges (Drugs) Connections
3 1,298 62.51 1,709 106,834
6 11,861 59.52 1,709 101,731
9 29,462 56.54 1,709 96,656
12 43,655 53.59 1,709 91,615

Small windows give a dense hypergraph whose substructures are shared across many drugs; large windows give a sparser, more drug-specific hypergraph.

Results

Mean accuracy across three seeds (%)

k M1 M2 M3
3 87.53 84.54 84.95
6 92.42 89.76 91.83
9 93.00 90.78 92.34
12 92.94 90.53 92.33

M1 leads at every window size. All configurations gain sharply from k = 3 to k = 6 and plateau from k = 9 to k = 12, and PR-AUC and ROC-AUC follow the same pattern. The seed-to-seed standard deviation of accuracy averages only 0.36 points, so the rankings do not depend on a particular split.

Top five individual configurations (%)

Rank Configuration Accuracy Precision Recall F1 ROC-AUC PR-AUC
🥇 M1-S42-K9 (E-1) 93.20 92.36 94.49 93.41 98.46 98.47
🥈 M1-S32-K12 (E-2) 93.13 91.96 94.69 93.30 98.42 98.41
🥉 M1-S32-K9 93.01 92.05 94.58 93.30 98.41 98.40
4 M1-S42-K12 92.90 92.53 93.76 93.14 98.34 98.32
5 M1-S42-K6 92.88 91.38 94.91 93.12 98.24 98.21

M = model configuration, S = data-splitting seed, K = k-mer window size.

The drug embeddings of the two best runs were extracted as 1,709 × 128 matrices and designated E-1 and E-2, the two initialization sources evaluated in Stage 2.

Baseline Comparison

Model Accuracy F1 ROC-AUC PR-AUC
Random Forest + Morgan FP (4,096-d) 79% 80% 88% 88%
GCN (2-layer) + Morgan FP 88% 88% 95% 95%
Chemical HGNN (Ours) 93% 93% 98% 98%

✅ +5 points of accuracy and F1 and +3 points of ROC-AUC and PR-AUC over the stronger GCN baseline, using a single hypergraph layer and no engineered fingerprints. On the seed-42 split, the proposed network reached a best validation loss of 0.1706 (epoch 494) against 0.2812 (epoch 499) for the GCN.


Stage 2: Interaction-Type HGNN

Hypergraph Construction

Interaction-type hypergraph construction

Nodes (Drugs) Hyperedges (Types) Connections Avg. Types per Drug Avg. Drugs per Type
1,709 86 13,486 7.89 156.81

Hyperedge size ranges from 5 drugs (Type 42) to 1,308 drugs (Type 49), with a median of 63, reflecting the same long-tailed imbalance seen in the dataset.

Results

The three configurations differ only in the class-weighting strength α applied to the validation loss at model selection: M2 (α = 0.3), M1 (α = 0.5), and M3 (α = 0.7).

Mean performance across two seeds (%)

Model Embedding Top-1 Acc. Top-3 Acc. Macro F1 PR-AUC ROC-AUC
M1 E-1 85.47 97.97 80.5 86.92 99.43
M1 E-2 85.36 98.04 79.0 87.02 99.42
M2 E-1 85.53 98.02 81.0 87.00 99.44
M2 E-2 85.17 97.96 79.5 87.15 99.35
M3 E-1 76.01 94.58 43.0 56.73 98.10
M3 E-2 76.19 94.25 40.5 56.84 98.02

M1 and M2 are practically interchangeable, while the heaviest weighting (M3) loses about nine points of Top-1 accuracy and collapses on macro F1 and PR-AUC. The embedding source matters very little: E-1 and E-2 differ by at most 0.36 points of Top-1 accuracy.

Top three individual configurations (%)

Rank Configuration ROC-AUC PR-AUC Top-1 Acc. Top-3 Acc. Macro P–R–F1 Weighted P–R–F1
🥇 M2-S42-E-1 99.44 87.24 85.73 98.11 85–80–81 86–86–86
🥈 M1-S42-E-1 99.44 87.24 85.73 98.11 85–80–81 86–86–86
🥉 M2-S32-E-1 99.44 86.76 85.33 97.93 85–80–81 85–85–85

The first two runs tie on every metric; M2-S42-E-1 ranks first on its lower best-epoch validation loss (0.4900 vs. 0.5440).

Key findings

  • The ~12-point gap between Top-1 (85.73%) and Top-3 (98.11%) means the correct interaction type is almost always among the three highest-ranked predictions.
  • The confusion matrix is strongly diagonal. Most of the 86 types score above 0.8 F1, while a small group of the rarest types pulls macro F1 (81) below weighted F1 (86).
  • Even the ten hardest types keep per-class ROC-AUC between 0.825 and 0.990, so PR-AUC and macro F1, not the near-saturated ROC-AUC, are the discriminating measures under this imbalance.

Baseline Comparison

Model Top-1 Acc. Top-3 Acc. ROC-AUC PR-AUC Macro F1 Weighted F1
GCN (2-layer) + Morgan FP 45.9% 80.7% 97.89% 56.77% 42% 48%
XGBoost + Morgan FP 69.9% 92.4% 98.46% 70.92% 64% 69%
GraphSAGE (2-layer) + Morgan FP 80.9% 96.8% 99.41% 76.76% 59% 81%
Interaction-Type HGNN (Ours) 85.7% 98.1% 99.44% 87.24% 81% 86%

✅ Against the strongest baseline (GraphSAGE): +4.8 Top-1, +1.3 Top-3, +10.5 PR-AUC, +22.0 macro F1, and +5.0 weighted F1 points.

The margin is notable because the baselines had two advantages the proposed network did not: precomputed Morgan fingerprints as input features and inverse-frequency class weighting during training.

Model Best Validation Loss Epoch Behaviour
Interaction-Type HGNN (Ours) 0.4900 499 Monotonic descent across the full budget
GraphSAGE + MLP 0.6245 440 Stable
GCN + MLP 0.8342 137 Oscillatory
GAT + MLP 1.2339 79 Early plateau; excluded from the final comparison

Single-layer GCN and GraphSAGE variants failed to converge (validation error above 100), so both were deepened to two layers and moved to a T4 GPU.

Cross-Validation

A five-fold stratified cross-validation confirmed that performance does not depend on a single split. A fixed 10% validation set (19,188 pairs) was held out, and the remaining 172,689 pairs were divided into five folds (train 138,151 / test 34,538 per fold), with all 86 types present in every fold.

Fold ROC-AUC PR-AUC Top-1 Acc. Top-3 Acc. Macro F1 Weighted F1
1 99.27 85.38 85.64 98.17 81 85
2 99.58 86.36 85.71 98.16 82 85
3 99.26 86.75 85.39 98.03 80 85
4 99.49 86.67 85.82 97.97 80 86
5 99.27 87.01 85.61 98.19 82 85
Mean ± SD 99.37 ± 0.15 86.43 ± 0.63 85.63 ± 0.16 98.10 ± 0.09 81.0 ± 1.0 85.2 ± 0.45

The cross-validated means sit within about one point of the single-split results, corroborating them rather than revising them.


Ablation Studies

All ablations were run on the Interaction-Type HGNN against its strongest configuration (M2-S42-E-1).

1. Influence of the Transferred Chemical Embeddings

Drug nodes were re-initialized with an all-ones vector instead of the E-1 embeddings, with everything else fixed.

Initialization Top-1 Acc. Top-3 Acc. PR-AUC Weighted F1 Validation Loss
E-1 embeddings (default) 85.73 98.11 87.24 86 0.4900
All-ones vector 82.52 97.51 85.70 82 0.5442

✅ The 3.21-point Top-1 drop confirms the value of the hierarchical transfer from Stage 1 to Stage 2.

2. Direction of the Attention Mechanism

The attention flow was reversed to node → hyperedge → node (seed 32, E-1).

Configuration ROC-AUC PR-AUC Top-1 Acc. Top-3 Acc. Macro P–R–F1
Reversed, α = 0.5 98.64 63.54 67.42 91.63 57–47–48
Reversed, α = 0.3 98.72 64.92 69.45 92.19 61–48–50
Original direction (M1/M2, seed 32) 99.42–99.44 86.60–86.76 85.21–85.33 97.84–97.93 80–81 F1

✅ Aligning the attention direction with the prediction target (the type hyperedges) matters more than starting from the better-informed node side.

3. Placement of Class Weighting

Inverse-frequency class weights were placed in the training loss, the validation loss, both, or neither (α = 0.3, seed 42).

Config. Train Loss Val Loss Top-1 Acc. Top-3 Acc. Macro R Macro F1 PR-AUC ROC-AUC
NONE – – 85.96 98.17 81 82 87.5 99.56
VAL – ✓ 85.96 98.17 81 82 87.5 99.56
TRAIN ✓ – 83.84 97.91 88 83 89.5 99.18
BOTH ✓ ✓ 83.83 97.89 88 83 89.6 99.18

✅ Weighting the validation criterion changes nothing: the unweighted objective already selects the balance-aware checkpoint. Weighting the training loss only moves the precision–recall operating point (macro recall 81 → 88, Top-1 down about two points) and adds cost to every gradient update. The hypergraph network reaches a well-balanced optimum without any rebalancing, whereas the baselines needed training-stage weighting to stay stable.

4. Recovery of Withheld Interaction Labels

For the 406 multi-label pairs, only the first type was used in training. Of these pairs, 38 fell into the test split, and each withheld secondary label was checked against the model's Top-3 predictions.

Drug 1 Drug 2 Kept Label Withheld Label Top-1 (Score) Top-2 (Score) Recovered At
Rifapentine Ranolazine 4 75 4 (0.696) 75 (0.205) Top-2
Phenytoin Artemether 11 73 75 (0.548) 73 (0.217) Top-2
Phenytoin Dienogest 70 75 70 (0.925) 75 (0.062) Top-2
Vincristine Venlafaxine 47 73 73 (0.588) 47 (0.365) Top-1
Vincristine Mitomycin 49 73 49 (0.890) 73 (0.080) Top-2

✅ The model recovered 94.7% of the withheld labels within its Top-3 predictions, indicating that the hypergraph encodes the multi-mechanistic structure of drug interactions rather than collapsing each pair to a single label.


Computational Efficiency

Quantity Chemical HGNN Interaction-Type HGNN
Total training time 142.00 min 18.62 min
Average time per epoch 17.04 s 2.23 s
Epochs completed 500 500
Best epoch (validation loss) 494 (0.1706) 499 (0.4900)
RAM increase during training 0.46 GB 0.29 GB
Hardware CPU (2 vCPUs) CPU (2 vCPUs)

The Interaction-Type HGNN is about 7.6× cheaper per epoch than the Chemical HGNN because its hypergraph has only 1,795 entities (1,709 drugs + 86 types) and its nodes start from compact 128-dimensional embeddings instead of 1,709-dimensional one-hot vectors.

Training Time Against the Multi-Class Baselines

Model Training Time Hardware
Interaction-Type HGNN (Ours, 1 layer) 18.62 min CPU (2 vCPUs)
GCN (2 layers) 75.56 min GPU (T4)
GraphSAGE (2 layers) 148.52 min GPU (T4)
XGBoost 225.55 min CPU

✅ 4–12× faster than the baselines while achieving higher accuracy, with no GPU required.

Reproducibility

  • ✅ Fixed seeds for every data split (Chemical HGNN: 32, 42, 46; Interaction-Type HGNN: 32, 42)
  • ✅ Exact data partitions provided, verified for zero leakage
  • ✅ Pre-computed k-mer files, hypergraphs, and index maps provided
  • ✅ Google Colab notebooks with inline outputs for every experiment
  • ✅ CPU-only Free Tier environment, so no special hardware is needed
1. Open an experiment notebook from Chemical/ or Type-Interaction/ in Google Colab
2. Point the data paths to the DataSet/, DataSet-Partitions/, K-mer/, and Hypergraph_Chemical/ files
3. Run all cells sequentially to reproduce the reported results

Limitations and Future Directions

These are the boundaries of the current framework, and the most natural places to extend it.

# Limitation Why It Matters
1 Single input modality. Drugs are represented from SMILES alone, with no protein-target, pathway, or side-effect information. Chemical structure carries the entire representation. Additional modalities could be added as further hyperedge families.
2 No reaction-intensity prediction. The system reports the interaction type and its likelihood, but not the severity of the reaction. Severity is clinically decisive and could itself hint at the underlying mechanism.
3 Limited drug coverage. Training used 1,709 drugs, while DrugBank now holds over 4,000. Many approved drugs cannot yet be screened. Extending coverage mainly requires SMILES and interaction records for the missing drugs.
4 Pairwise interactions only. The system evaluates drug pairs, not the simultaneous multi-drug combinations that real polypharmacy involves. The hypergraph formulation can express higher-order co-administration directly, so this is a modelling extension rather than a structural barrier.

Contributing

Contributions are welcome. If you would like to extend the framework, the limitations above are the most useful starting points: adding a modality, predicting severity, widening drug coverage, or moving from pairs to true multi-drug sets.

How to take part:

  1. Open an issue to describe the idea or the problem before writing code, so the direction can be discussed first.
  2. Fork the repository and work on a branch named after the change, for example feature/target-hyperedges.
  3. Keep the experiment format. New experiments should be Google Colab notebooks with inline outputs, a fixed seed, and the split files from DataSet-Partitions/, so results stay comparable with those reported here.
  4. Report the full metric set used in this README (ROC-AUC, PR-AUC, Top-1 and Top-3 accuracy, macro and weighted F1) along with training time and hardware.
  5. Open a pull request describing what changed, which configuration you ran, and how the numbers compare with the tables above.

Questions, bug reports, and requests for clarification about the data or the notebooks are equally welcome through the issue tracker or by email.

Publication

Status ✅ Published
Conference ECAI 2026 — 18th International Conference on Electronics, Computers and Artificial Intelligence
Date & Location July 2–3, 2026, Bucharest, Romania
Publisher IEEE Xplore (Scopus-indexed)
Paper doi.org/10.1109/ECAI69016.2026.11613729
Project Page husseinmahdi.xyz/research/hhgnn

Citation

If you use this work in your research, please cite:

@inproceedings{alrubaie2026hhgnn,
  title     = {Hierarchical Hypergraph Neural Networks for Multi-Type Drug--Drug Interaction Prediction},
  author    = {Al-Rubaie, Hussein Mahdi and Al-Rashid, Sura},
  booktitle = {2026 18th International Conference on Electronics, Computers and Artificial Intelligence (ECAI)},
  address   = {Bucharest, Romania},
  publisher = {IEEE},
  year      = {2026},
  doi       = {10.1109/ECAI69016.2026.11613729}
}

Contact

For questions or collaborations:

Acknowledgements

The author gratefully acknowledges the DrugBank Foundation for providing access to the curated drug information resources that supported this study, including the SMILES representations that completed the molecular coverage of the dataset.


Repository Status: 🔓 Public

About

The repo for the H2GNN model published @ECAI2026 in IEEE-Xplore to predicting DDI and their mechanism from SMILES alone.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages