数据与计算发展前沿 ›› 2026, Vol. 8 ›› Issue (4): 42-71.

CSTR: 32002.14.jfdc.CN10-1649/TP.2026.04.004

doi: 10.11871/jfdc.issn.2096-742X.2026.04.004

• • 上一篇    下一篇

向量数据库:存储与检索技术综述

马乐1(),张然2,5,韩颐堃3,于诗睿4,5,王在田2,5,宁致远2,5,张静涵6,许萍2,5,李鹏江2,5,乔子越7,琚玮8,陈冲9,王东杰10,刘鲲鹏6,汪澎洋11,王鹏飞2,5,傅衍杰12,刘春江4,5,*(),吕长天13   

  1. 1 四川大学公共管理学院四川 成都 610065
    2 中国科学院计算机网络信息中心北京 100083
    3 伊利诺伊大学厄巴纳-香槟分校信息科学学院美国 香槟 61801
    4 中国科学院成都文献情报中心四川 成都 610299
    5 中国科学院大学北京 100190
    6 波特兰州立大学计算机科学系美国 波特兰 97207-0751
    7 东莞大湾区大学广东 东莞 523808
    8 四川大学计算机学院四川 成都 610065
    9 特斯联科技集团未来城市实验室北京 100027
    10 堪萨斯大学电气工程与计算机科学系美国 劳伦斯 66045
    11 澳门大学计算机与信息科学系澳门 999078
    12 亚利桑那州立大学美国 坦佩 85287
    13 弗吉尼亚理工大学美国 布莱克斯堡 24061
  • 收稿日期:2026-01-22 出版日期:2026-08-20 发布日期:2026-08-21
  • 通讯作者: 刘春江(E-mail: liucj@clas.ac.cn
  • 作者简介:马乐,四川大学公共管理学院,硕士,研究方向为推荐系统、搜索澄清。
    本文中负责编程、论文写作等。
    MA Le holds a master’s degree from College of Public Administration, Sichuan University. His research interests include recommendation systems and search clarification.
    In this paper, he is responsible for programming and manuscript writing.
    E-mail:2713347318@qq.com|刘春江,中国科学院成都文献情报中心,高级工程师,博士,中国科学院大学,硕士研究生导师,研究方向为科技文献挖掘与知识发现。
    本文中负责论文写作、修改等。
    LIU Chunjiang, Ph.D., is a senior engineer of National Science Library (Chengdu), Chinese Academy of Sciences, and a supervisor of master’s students at the University of Chinese Academy of Science. His research interests include S&T literature mining and scientific knowledge discovery.
    In this paper, he is responsible for writing and revising the manuscript.
    E-mail: liucj@clas.ac.cn

A Comprehensive Survey on Vector Database: Storage and Retrieval Techniques

MA Le1(),ZHANG Ran2,5,HAN Yikun3,YU Shirui4,5,WANG Zaitian2,5,NING Zhiyuan2,5,ZHANG Jinghan6,XU Ping2,5,LI Pengjiang2,5,QIAO Ziyue7,JU Wei8,CHEN Chong9,WANG Dongjie10,LIU Kunpeng6,WANG Pengyang11,WANG Pengfei2,5,FU Yanjie12,LIU Chunjiang4,5,*(),LYU Changtian13   

  1. 1 College of Public Administration, Sichuan University, Chengdu, Sichuan 610065, China
    2 Computer Network Information Center, Chinese Academy of Sciences, Beijing 100083, China
    3 University of Illinois Urbana-Champaign, Champaign 61801, USA
    4 National Science Library (Chengdu), Chinese Academy of Sciences, Chengdu, Sichuan 610299, China
    5 University of Chinese Academy of Sciences, Beijing 100190, China
    6 Portland State University, Portland 97207-0751, USA
    7 Great Bay University, Dongguan, Guangdong 523808, China
    8 College of Computer Science, Sichuan University, Chengdu, Sichuan 610065, China
    9 Future City Lab, Terminus Group, Beijing 100027, China
    10 University of Kansas, Lawrence 66045, USA
    11 University of Macau, Macau 999078, China
    12 Arizona State University, Tempe 85287, USA
    13 Virginia Tech, Blacksburg 24061, USA
  • Received:2026-01-22 Online:2026-08-20 Published:2026-08-21

摘要:

【目的】 高维向量数据的涌现已超越传统数据库的处理能力,推动向量数据库(VDBs)快速发展并与大语言模型深度融合,广泛支持现代人工智能系统。然而,现有研究多集中于近似最近邻搜索等技术细节,缺乏从系统架构层面出发的整体性综述,也未深入探讨核心技术如何协同构建VDBs的综合能力。本文旨在系统梳理向量数据库的核心设计、关键算法与架构,为其发展提供完整认知框架。【方法】 首先,围绕存储与检索两个核心维度,系统回顾VDBs的关键技术与设计理念,梳理其技术演进路径;其次,对比分析多款主流向量数据库的架构特性,总结其优势、局限和适用场景;最后,探讨向量数据库与大语言模型融合的前沿趋势,包括新型索引策略等开放问题与发展方向。【结论】 本综述可为研究人员与从业者提供系统性参考,帮助读者把握该领域技术全景与发展动态,促进向量数据库在理论与应用层面的进一步创新。

关键词: 向量数据库, 相似性搜索, 索引技术, 机器学习嵌入, 大语言模型, 人工智能基础设施

Abstract:

[Objective] As high-dimensional vector data increasingly surpass the processing capabilities of traditional database management systems, Vector Databases (VDBs) have emerged and become tightly integrated with large language models, being widely applied in modern artificial intelligence systems. However, existing research has primarily focused on underlying technologies such as approximate nearest neighbor search, with relatively few studies providing a systematic architectural-level review of VDBs or analyzing how these core technologies collectively support the overall capacity of VDBs. This survey aims to offer a comprehensive overview of the core designs and algorithms of VDBs, establishing a holistic understanding of this rapidly evolving field. [Methods] First, we systematically review the key technologies and design principles of VDBs from the two core dimensions of storage and retrieval, tracing their technological evolution. Next, we conduct an in-depth comparison of several mainstream VDB architectures, summarizing their strengths, limitations, and typical application scenarios. Finally, we explore emerging directions for integrating VDBs with large language models, including open research challenges and trends such as novel indexing strategies. [Conclusions] This survey serves as a systematic reference guide for researchers and practitioners, helping readers quickly grasp the technological landscape and development trends in the field of vector databases, and promoting further innovation in both theoretical and applied aspects.

Key words: vector database, similarity search, indexing techniques, machine learning embeddings, large language models, AI infrastructure