数据与计算发展前沿 ›› 2026, Vol. 8 ›› Issue (4): 191-202.

CSTR: 32002.14.jfdc.CN10-1649/TP.2026.04.014

doi: 10.11871/jfdc.issn.2096-742X.2026.04.014

• • 上一篇    下一篇

融合无分类器引导与缓存机制的图像生成方法

王良君*(),钱诣   

  1. 江苏大学江苏 镇江 212013
  • 收稿日期:2026-02-08 出版日期:2026-08-20 发布日期:2026-08-21
  • 通讯作者: 王良君(E-mail: 2212308028@stmail.ujs.edu.cn
  • 作者简介:王良君,江苏大学计算机学院,硕士生导师,主要研究方向为图像压缩、压缩感知、图像生成。
    本文中主要负责方法研究与实验、论文撰写。
    WANG Liangjun is a a supervisor of master’s students at the School of Computer Science and Communication Engineering, Jiangsu University. His main research interests include image compression, compressive sensing, and image generation.
    In this paper, he is primarily responsible for methodology research, experiments, and manuscript writing.
    E-mail:2212308028@stmail.ujs.edu.cn

Image Generation Method Integrating Classifier-Free Guidance and Caching Mechanism

WANG Liangjun*(),QIAN Yi   

  1. Jiangsu University, Zhenjiang, Jiangsu 212013, China
  • Received:2026-02-08 Online:2026-08-20 Published:2026-08-21

摘要:

【目的】 扩散模型在生成图像时需多步去噪使得推理成本高。特征缓存机制可减少冗余计算,但影响生成质量。无分类器引导机制能提高生成质量却增加了计算量。为兼顾效率与质量,【方法】 本文改进了无分类器引导机制,使其只需少量计算即可与特征缓存机制协同工作。此外,设计了潜在空间中的蒸馏方式用于训练无分类器引导中所需的引导模型,使其引导效果更佳。设计了梯度平滑参数更新策略提高训练质量。【结果】 在U-ViT和DiT等架构上对不同分辨率的数据进行实验,通过FID、IS、Precision等指标评估生成质量、分布覆盖度和推理效率,【结论】 结果表明所提出的方法在不同分辨率和多种模型架构下均实现了更优的质量-速度权衡。

关键词: 扩散模型, 图像生成, 无分类器引导, 推理加速, 蒸馏

Abstract:

[Objective] Diffusion models require multi-step denoising during image generation, resulting in high inference cost. Feature caching can reduce redundant computation, but it degrades generation quality. Classifier-free guidance can improve generation quality but increases computational load. [Methods] To balance efficiency and quality, this paper improves classifier-free guidance so that it requires only a small amount of computation and can work synergistically with feature caching. In addition, we design a latent-space distillation method to train the guidance model needed for classifier-free guidance, achieving stronger guidance effects. A gradient-smoothing strategy is introduced to further improve training quality. [Results] Experiments are conducted on datasets of different resolutions using architectures such as U-ViT and DiT. Metrics such as FID, IS, precision are used to evaluate generation quality, distribution coverage, and inference efficiency. [Conclusion] The results show that the proposed method achieves a superior quality-speed trade-off across different resolutions and model architectures.

Key words: diffusion models, image generation, classifier-free guidance, inference acceleration, distillation