频谱对齐潜流匹配改进时间序列生成质量
Time series generation with spectrally aligned latent flow matching
做合成数据的朋友可以看看:给潜流模型加傅里叶和小波变换损失,补上频谱短板,合成的时间序列更像真实数据还更省算力。
潜流模型(latent flow matching)把时间序列压缩到潜空间时会引入频谱失真,导致合成样本难以用作训练替身。该方法在流匹配的潜空间训练中引入基于傅里叶变换、小波变换和 signature 变换的微调损失。这些变换具有可解释性,可保证合成信号在平滑度、目标频谱内容等特征上与真实数据对齐,而非仅依赖逐点重建损失。在真实世界长序列单变量与多变量基准数据集上,该方法在信号真实度和计算效率指标上均优于基础潜流模型和现有最优方法。
Time series generation with spectrally aligned latent flow matching
Latent flow models have proven to be a reliable and cost-effective method for time series generation. However, the latent compression induces unwanted artefacts, such as a spectral mismatch with respect to the underlying dataset, thus hindering their use as training surrogates. In this article, we propose a spectrally-aligned latent-flow time series generator, where the latent space for flow matching is trained to preserve dynamical properties that are relevant for the suitability of synthetic samples. We find that incorporating fine-tuning losses based on canonical signal representations such as the Fourier, wavelet and signature transforms helps overcome these issues. The interpretability of these transformations allows us to ensure that the synthetic signals are aligned with the true ones in terms of relevant features, such as smoothness or targeted spectral content, as opposed to relying on pointwise reconstruction losses only. We compare the proposed aligned models against a base latent-flow model and the state of the art over real-world long-range univariate and multivariate benchmark datasets. Our quantitative results validate the superiority of the proposed method in terms of its performance on metrics reflecting signal realness and computational efficiency, while being aligned to the training set with respect to its local structure.