论文

G²MLP:基于曲率的 GNN 到 MLP 蒸馏框架

Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

精选理由

把 GNN 的几何信息蒸馏进普通 MLP,推理不用图也能保住精度,做图机器学习部署的可以看看

arXiv 论文提出 G²MLP,一个 GNN 到 MLP 的知识蒸馏框架。作者指出仅迁移节点预测会带来两种谱失败模式:稀疏图上的谱欠拟合和稠密图上的谱过拟合。G²MLP 用 Ollivier-Ricci 曲率定位错误集中位置,并据此在预测级和表示级对齐之间分配监督。在节点分类基准上,该框架优于无图蒸馏基线,还能迁移到 Graph Transformer 教师和链接预测任务,推理时仍是无需图访问的标准 MLP。

原文 · arXiv cs.LG

Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.