置换等变流匹配实现无对齐权重生成
Permutation-Equivariant Flow Matching for Alignment-Free Neural Weight Generation
这篇论文提出了一种新方法,能从独立训练的网络中生成新模型,无需昂贵的神经元对齐,还能跨不同架构工作。
研究人员提出置换等变流匹配方法,通过图元网络参数化流匹配速度场,无需对齐即可从独立训练网络学习。该方法在准确率、功能相似性和权重相似性联合统计上接近真实网络表现。单个条件模型可生成异构架构的任务特定网络,并泛化到未见隐藏宽度配置。在表格领域漂移任务中,中间条件生成的网络性能与logit集成相当。
Permutation-Equivariant Flow Matching for Alignment-Free Neural Weight Generation
A trained neural network can be represented by a parameter vector in high dimensions. Learning distributions over these vectors enables the generation of new models across various tasks and architectures. A central challenge is permutation symmetry: permuting hidden neurons can produce distant parameter vectors representing the same function. This introduces variations that a generative model must account for when learning from trained networks. Existing methods typically address this using networks derived from a common base model or costly approximate neuron alignment. We instead parameterize a flow-matching velocity field with a permutation-equivariant Graph Meta Network, enabling direct learning from independently trained networks without alignment. Extensive experiments show that our method closely reproduces the joint statistics of accuracy, functional similarity, and weight similarity of independently trained collections, providing evidence of generation beyond checkpoint memorization. A single conditional model also generates task-specific networks on heterogeneous architectures and generalizes to unseen hidden-width configurations. On a tabular domain-shift task, intermediate conditioning produces individual networks with performance comparable to logit ensembles across both domains. Taken together, our results show how permutation equivariance enables learning from diverse collections of independently trained networks without permutation alignment.