Skip to content

Repository files navigation

RGRGAD

Get Started

This is the source code of RGRGAD, a routing-guided graph anomaly detection framework for attributed networks.

CSV dataset download: Download csvdata.zip (27.5 MB)
Release page: dataset-v1.0

RGRGAD performs unsupervised node-level anomaly detection using an epoch-level curriculum that selects feature-similarity-based relation suppression or routing-guided context completion. The learned gate and matrix C directly modulate the completion branch.

Code Structure

File / Folder Description
Data Datasets in .mat format.
datasets/README.md CSV dataset availability and format documentation.
run.py Training and evaluation entry.
model.py GCN encoder, discriminator, and routing gate.
aug.py Redundancy pruning and neighbor completion.
utils.py Data loading, preprocessing, and subgraph sampling.

Datasets

Dataset # Nodes # Edges # Attributes # Anomalies
Cora 2,708 5,429 1,433 5.5%
Citeseer 3,327 4,732 3,703 4.5%
Pubmed 19,717 44,338 500 3.0%
ACM 16,484 71,980 8,337 3.6%
BlogCatalog 5,196 171,743 8,189 5.8%
Reddit 10,984 168,016 64 3.3%

Human- and Machine-Readable CSV Data

Human- and machine-readable CSV versions of all six benchmark datasets are provided in the GitHub Release:

The archive contains graph edge lists, node attribute matrices, node labels, metadata files, and README documentation. These CSV files correspond to the original MATLAB-format benchmark datasets and are provided for inspection, reuse, and reproducibility.

The datasets should be organized as follows for model execution:

RGRGAD/
└── Data/
    ├── acm.mat
    ├── blogcatalog.mat
    ├── citeseer.mat
    ├── cora.mat
    ├── pubmed.mat
    └── reddit.mat

The .mat files should contain node labels, node attributes, and the adjacency matrix. The loader supports the following field names:

Labels:     Label / gnd / label
Attributes: Attributes / X / attr
Adjacency:  Network / A / adj

Data Processing

The data processing follows the common benchmark setting used in CoLA for graph anomaly detection. The CoLA repository is available at: https://github.com/TrustAGI-Lab/CoLA. In this repository, the processed .mat files are directly placed in the Data/ folder.

During loading, the code applies the following operations:

  • The adjacency matrix is converted into a sparse CSR matrix and then into a DGL graph.
  • Node attributes are row-normalized before being fed into the model.
  • The adjacency matrix is symmetrically normalized and self-loops are added for graph convolution.
  • Local subgraphs are sampled around each target node for contrastive learning.
  • Labels are used only for final ROC-AUC and PR-metric evaluation, not for hyperparameter or checkpoint selection.

Main and sensitivity configuration

Setting Cora / CiteSeer ACM / BlogCatalog / PubMed / Reddit
Training epochs 200 100
Batch size 300 300
Local subgraph size 4 4
Inference rounds 256 256
Learning rate / dimension / degree threshold 0.001 / 64 / 8 0.001 / 64 / 8
Main alpha / tau 0.2 / 0.07 0.2 / 0.07
Seeds 42, 47, 58, 90, 100 42, 47, 58, 90, 100

Main comparisons and ablations use the dataset-specific budget. One-dimensional sensitivity varies only the investigated parameter. The revised Figure 6 reproduction protocol varies only alpha and tau; all other settings use the same dataset-specific main configuration. The earlier 128/2/100/196 launcher profile is superseded.

Usage

When train_epoch is omitted, run.py uses the dataset-specific default. Explicit CLI values override it.

python run.py --dataset cora --data_dir ./Data --seed 42 --train_epoch 200 --batch_size 300 --subgraph_size 4 --test_rounds 256
python run_experiments.py --profile main --dry-run
python run_experiments.py --profile main

For one joint-sensitivity cell:

python run_experiments.py --profile joint --datasets cora --seeds 42 47 58 90 100 --alpha 0.1 --tau 0.2 --dry-run
python run_experiments.py --profile joint --datasets cora --seeds 42 47 58 90 100 --alpha 0.1 --tau 0.2

The joint profile uses batch 300, subgraph size 4, 256 inference rounds and 200/100 training epochs by dataset. It does not infer a historical parameter grid. The command describes the revised reproduction protocol, not proof of how an existing heatmap was generated.

Routing gate and temporary views

The gate is Linear(2,16), ReLU, Linear(16,1), sigmoid, and its parameters are included in Adam. Detached pairwise weights are used for augmentation; differentiable gates are recomputed in every mini-batch and trained through the routing-consistency constraint. Categorical graph sampling itself does not propagate gradients.

Suppression uses feature-similarity weights restricted by the implementation adjacency and self-identity. Completion uses C; the main fusion is 0.6 S_feat + 0.2 S_ano + 0.2 R. Auxiliary nodes come from the full graph; a zero-total row uses uniform sampling over all nodes. Refinements are temporary training views; inference uses the original graph.

Model selection and metrics

The checkpoint minimizes training discrimination loss. Sensitivity results do not select hyperparameters or checkpoints. Five-run main comparisons report mean and sample standard deviation, with 95% confidence intervals in the supplementary table.

ROC-AUC and the reported PR metric are exported as auc and ap. The PR metric uses sklearn.metrics.average_precision_score, the AP formulation explicitly defined in the manuscript, rather than trapezoidal PR integration.

Software environment

The versions reported in manuscript Section 5.2 are:

Component Reported version
Python 3.7.1
PyTorch 1.10.2
CUDA 11.3
DGL 0.4.3post2
GPU NVIDIA GeForce RTX 4090, 24 GB

This is the single manuscript/README specification. The earlier Python >=3.8 and PyTorch >=1.12.1 requirements are superseded. requirements.txt pins the two package versions reported in the manuscript. Historical versions of torch-scatter, NumPy, SciPy, scikit-learn, NetworkX and tqdm are not inferred. Use mutually compatible CUDA/PyTorch/DGL/torch-scatter builds.

Capture installed versions without training:

python capture_environment.py
python -m pip freeze > environment_actual_requirements.txt

The reported specification does not claim an independently verified historical environment export or a tested GPU installation during this update.

Output

run_experiments.py isolates each dataset/seed/configuration, records its command and training log, and requests JSON metrics, scores and curve export. This code update does not run or overwrite experiments.

Cite

If this repository is useful for your research, please cite the corresponding paper when it becomes available.

About

Unsupervised Graph Anomaly Detection on Attributed Networks via Routing-Guided Dynamic Structural Refinement

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages