数据与计算发展前沿 ›› 2026, Vol. 8 ›› Issue (4): 246-256.

CSTR: 32002.14.jfdc.CN10-1649/TP.2026.04.019

doi: 10.11871/jfdc.issn.2096-742X.2026.04.019

• • 上一篇    

融合标签语义的农业科学数据层次多标签分类方法研究

张思扬(),胡林*()   

  1. 中国农业科学院农业信息研究所北京 100081
  • 收稿日期:2026-01-07 出版日期:2026-08-20 发布日期:2026-08-21
  • 通讯作者: 胡林(E-mail: hulin@caas.cn
  • 作者简介:张思扬,中国农业科学院农业信息研究所,硕士研究生,主要研究方向为数据科学与深度学习。
    本文中负责论文撰写和模型构建。
    ZHANG Siyang is a master’s student at the Agricultural Information Institute of the Chinese Academy of Agricultural Sciences. Her main research interests include data science and deep learning.
    In this paper, she is mainly responsible for manuscript writing and model building.
    E-mail: zhangsy0579@163.com|胡林,中国农业科学院农业信息研究所,研究员,硕士生导师,主要研究方向为数据科学、农业信息技术、智慧农业。
    本文中负责写作指导以及论文最终审定。
    HU Lin is a researcher and master's supervisor at the Agricultural Information Institute of the Chinese Academy of Agricultural Sciences. His main research interests include data science, agricultural information technology, and smart agriculture.
    In this paper, he is responsible for writing instruction and manuscript reviewing.
    E-mail: hulin@caas.cn
  • 基金资助:
    中央级公益性科研院所基本科研业务费专项(JBYW-AII-2026-46);中央级公益性科研院所基本科研业务费专项(JBYW-AII-2026-47 & Y2026JC11)

Fusing Label Semantics for Hierarchical Multi-Label Classification of Agricultural Science Data

ZHANG Siyang(),HU Lin*()   

  1. Agricultural Information Institute of CAAS, Beijing 100081, China
  • Received:2026-01-07 Online:2026-08-20 Published:2026-08-21

摘要:

【背景】 农业科学数据是支撑农业科技创新与现代农业发展的重要基础,其规模不断扩大、文本描述专业性强且学科标签具有层次结构,为分类带来较大挑战。【目的】 针对农业科学数据分类标签层次复杂、依赖人工标注导致效率低下的问题,构建一种融合标签语义的农业科学数据层次多标签分类模型Bert-BiGRU-HiGCN,实现农业科学数据的高效自动化分类,提升数据管理与检索效率。【方法】 基于国家农业科学数据中心农业科学数据集的元数据构建实验数据集,并在模型中引入GCN(Graph Convolutional Network)对农业学科标签的层次结构进行建模,提取包含层级信息的标签特征,然后通过交叉注意力机制融合文本与标签特征,获得融合标签语义的文本表示。最后,在构建的数据集上通过对比实验评估改进后的模型效果。【结果】 实验结果表明,Bert-BiGRU-HiGCN模型在农业科学数据层次多标签分类任务中精确率为78.1%、召回率为76.3%、F1值为77.2%,整体性能优于对比模型,有效提高了分类准确性及效率。

关键词: 农业科学数据, 文本分类, 层次多标签分类, 深度学习

Abstract:

[Background] Agricultural science data serve as a key foundation for supporting agricultural technological innovation and modern agricultural development. These datasets are continuously expanding, feature highly specialized textual descriptions, and have hierarchical subject labels, which pose significant challenges for classification. [Objective] To address the complexity of hierarchical labels in agricultural science data and the low efficiency caused by manual annotation, this study develops a label-semantic-aware hierarchical multi-label classification model, Bert-BiGRU-HiGCN, aiming to achieve efficient automated classification and improve data management and retrieval. [Methods] An experimental dataset is constructed based on metadata from the National Agricultural Science Data Center. In the model, a Graph Convolutional Network was used to capture the hierarchical structure of agricultural subject labels and extract label features containing hierarchical information. A cross-attention mechanism was then applied to fuse textual and label features, obtaining text representations enriched with label semantics. The improved model was evaluated on the constructed dataset through comparative experiments. [Results] Experimental results indicate that the proposed Bert-BiGRU-HiGCN model achieves a precision of 78.1%, recall of 76.3%, and F1-score of 77.2% in the hierarchical multi-label classification task of agricultural science data, outperforming baseline models and effectively improving classification accuracy and efficiency.

Key words: agricultural science data, text classification, hierarchical multi-label classification, deep learning