Frontiers of Data and Computing ›› 2026, Vol. 8 ›› Issue (4): 42-71.

CSTR: 32002.14.jfdc.CN10-1649/TP.2026.04.004

doi: 10.11871/jfdc.issn.2096-742X.2026.04.004

Previous Articles     Next Articles

A Comprehensive Survey on Vector Database: Storage and Retrieval Techniques

MA Le1(),ZHANG Ran2,5,HAN Yikun3,YU Shirui4,5,WANG Zaitian2,5,NING Zhiyuan2,5,ZHANG Jinghan6,XU Ping2,5,LI Pengjiang2,5,QIAO Ziyue7,JU Wei8,CHEN Chong9,WANG Dongjie10,LIU Kunpeng6,WANG Pengyang11,WANG Pengfei2,5,FU Yanjie12,LIU Chunjiang4,5,*(),LYU Changtian13   

  1. 1 College of Public Administration, Sichuan University, Chengdu, Sichuan 610065, China
    2 Computer Network Information Center, Chinese Academy of Sciences, Beijing 100083, China
    3 University of Illinois Urbana-Champaign, Champaign 61801, USA
    4 National Science Library (Chengdu), Chinese Academy of Sciences, Chengdu, Sichuan 610299, China
    5 University of Chinese Academy of Sciences, Beijing 100190, China
    6 Portland State University, Portland 97207-0751, USA
    7 Great Bay University, Dongguan, Guangdong 523808, China
    8 College of Computer Science, Sichuan University, Chengdu, Sichuan 610065, China
    9 Future City Lab, Terminus Group, Beijing 100027, China
    10 University of Kansas, Lawrence 66045, USA
    11 University of Macau, Macau 999078, China
    12 Arizona State University, Tempe 85287, USA
    13 Virginia Tech, Blacksburg 24061, USA
  • Received:2026-01-22 Online:2026-08-20 Published:2026-08-21

Abstract:

[Objective] As high-dimensional vector data increasingly surpass the processing capabilities of traditional database management systems, Vector Databases (VDBs) have emerged and become tightly integrated with large language models, being widely applied in modern artificial intelligence systems. However, existing research has primarily focused on underlying technologies such as approximate nearest neighbor search, with relatively few studies providing a systematic architectural-level review of VDBs or analyzing how these core technologies collectively support the overall capacity of VDBs. This survey aims to offer a comprehensive overview of the core designs and algorithms of VDBs, establishing a holistic understanding of this rapidly evolving field. [Methods] First, we systematically review the key technologies and design principles of VDBs from the two core dimensions of storage and retrieval, tracing their technological evolution. Next, we conduct an in-depth comparison of several mainstream VDB architectures, summarizing their strengths, limitations, and typical application scenarios. Finally, we explore emerging directions for integrating VDBs with large language models, including open research challenges and trends such as novel indexing strategies. [Conclusions] This survey serves as a systematic reference guide for researchers and practitioners, helping readers quickly grasp the technological landscape and development trends in the field of vector databases, and promoting further innovation in both theoretical and applied aspects.

Key words: vector database, similarity search, indexing techniques, machine learning embeddings, large language models, AI infrastructure