Research on Medical Information Retrieval Methods Integrating Query Intent and Large Language Models
Han Pu1,2, Liu Senling1, Li Xiong1, Wang Wei1
1.School of Management, Nanjing University of Posts and Telecommunications, Nanjing 210003 2.Key Laboratory of Data Engineering and Knowledge Services in Provincial Universities (Nanjing University), Nanjing 210023
韩普, 刘森嶺, 李雄, 王伟. 融合查询意图与大语言模型的医学信息检索方法研究[J]. 情报学报, 2026, 45(8): 1154-1165.
Han Pu, Liu Senling, Li Xiong, Wang Wei. Research on Medical Information Retrieval Methods Integrating Query Intent and Large Language Models. 情报学报, 2026, 45(8): 1154-1165.
1 马费成, 宋恩梅, 赵一鸣. 信息管理学基础[M]. 3版. 武汉: 武汉大学出版社, 2018. 2 Nakpih C I. A modified vector space model for semantic information retrieval[J]. Natural Language Processing Journal, 2024, 8: 100081. 3 李纲, 毛进, 芦昆. 医学信息检索中一种基于概念的查询相关模型[J]. 情报学报, 2014, 33(3): 239-249. 4 杨善林, 丁帅, 顾东晓, 等. 医疗健康大数据驱动的知识发现与知识服务方法[J]. 管理世界, 2022, 38(1): 219-229. 5 梁少博, 吴丹, 徐惟佳. 面向数字图书馆和档案馆的信息基础设施与机器学习: 数据管理、分析与出版的融合[J]. 图书情报知识, 2018(5): 72-80. 6 Zakaria Y, Ishiyama R, Ishidera E, et al. Fast retrieval of pharmaceutical packaging images using keypoint matching with angle and scale voting for outlier rejection[C]// Proceedings of the 2024 IEEE International Conference on Visual Communications and Image Processing. Piscataway: IEEE, 2024: 1-5. 7 赵宇翔, 付振康, 曾鹏翔, 等. 生成式人工智能驱动下智慧图书馆信息搜索的技术框架及服务模式研究[J]. 中国图书馆学报, 2025, 51(5): 54-66. 8 张贞港, 余传明. 基于知识增强的文本语义匹配模型研究[J]. 情报学报, 2024, 43(4): 416-429. 9 赵广宇, 段永康, 耿骞, 等. 基于大语言模型数据增强和对比学习的政务相似问题检索研究[J]. 数据分析与知识发现, 2026, 10(4): 116-129. 10 Cui Y M, Che W X, Liu T, et al. Revisiting pre-trained models for Chinese natural language processing[C]// Findings of the Association for Computational Linguistics: EMNLP 2020. Stroudsburg: Association for Computational Linguistics, 2020: 657-668. 11 Song M Y, Zheng M. A survey of query optimization in large language models[PP/OL]. V3. arXiv (2026-03-03). https://doi.org/10.48550/arXiv.2412.17558. 12 李澳, 涂新辉, 姚彪, 等. 基于ChatGPT查询改写的文档检索方法[J]. 中文信息学报, 2025, 39(8): 107-116, 138. 13 Wang L, Yang N, Wei F R. query2doc: query expansion with large language models[C]// Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: Association for Computational Linguistics, 2023: 9414-9423. 14 Gao L Y, Ma X G, Lin J, et al. Precise zero-shot dense retrieval without relevance labels[C]// Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2023: 1762-1777. 15 Dai A J, Zhu Z Y, Hu H Q, et al. Enhancing E-commerce query rewriting: a large language model approach with domain-specific pre-training and reinforcement learning[C]// Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. New York: ACM Press, 2024: 4439-4445. 16 Ma X B, Gong Y Y, He P C, et al. Query rewriting in retrieval-augmented large language models[C]// Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: Association for Computational Linguistics, 2023: 5303-5315. 17 Seonwoo Y, Wang G Y, Seo C, et al. Ranking-enhanced unsupervised sentence representation learning[C]// Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2023: 15783-15798. 18 夏立新, 龙存钰, 胡畔, 等. 基于查询意图的细粒度图书分面检索研究[J]. 图书情报工作, 2024, 68(8): 122-132. 19 陆伟, 周红霞, 张晓娟. 查询意图研究综述[J]. 中国图书馆学报, 2013, 39(1): 100-111. 20 赵一鸣, 潘沛, 毛进. 基于任务知识融合与文本数据增强的医学信息查询意图强度识别研究[J]. 数据分析与知识发现, 2023, 7(2): 38-47. 21 Duan Q. Method of short text classification based on TF-IWF feature selection[J]. International Journal of Social Science and Education Research, 2021, 4(4): 367-375. 22 Cai R C, Zhu B J, Ji L, et al. A CNN-LSTM attention approach to understanding user query intent from online health communities[C]// Proceedings of the 17th IEEE International Conference on Data Mining Workshops. Piscataway: IEEE, 2017: 430-437. 23 Luo Y Y, Xie Y, Yan E L, et al. A user intent recognition model for medical queries based on attentional interaction and focal loss boost[C]// Proceedings of the International Conference on Neural Computing for Advanced Applications. Singapore: Springer, 2023: 245-259. 24 龚建全, 王德军, 孟博. 基于样本构造和孪生胶囊网络的医学意图识别[J]. 中南民族大学学报(自然科学版), 2022, 41(5): 606-612. 25 Wang Y Q, Wang S, Li Y Y, et al. Recognizing medical search query intent by few-shot learning[C]// Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM Press, 2022: 502-512. 26 Liang Z J, Zhao Y X, Xu H W, et al. A hybrid model integrating RoBERTa, TF-IDF, and attention mechanism for medical query intent classification[J]. Scientific Reports, 2025, 15: Article No.42847. 27 韩普, 李雄, 陈文祺, 等. AIGC驱动的医疗健康领域知识服务模式研究[J]. 现代情报, 2026, 46(4): 68-79. 28 Shah C, White R, Andersen R, et al. Using large language models to generate, validate, and apply user intent taxonomies[J]. ACM Transactions on the Web, 2025, 19(3): Article No.34. 29 Madrue?o N, Fernández-Isabel A, Fernández R R, et al. Exploring new methods of Data augmentation for Intent classification through large language models[C]// Proceedings of the Computational Science and Computational Intelligence. Cham: Springer, 2025: 16-29. 30 Yang D K, Wei J J, Li M C, et al. MedAide: information fusion and anatomy of medical intents via LLM-based agent collaboration[J]. Information Fusion, 2026, 127: 103743. 31 谌文佳, 杨琳, 李金林. 嵌入意图识别的医疗健康问答文本语义分类模型[J]. 数据分析与知识发现, 2025, 9(2): 26-38. 32 Tamine L, Goeuriot L. Semantic information retrieval on medical texts: research challenges, survey, and open issues[J]. ACM Computing Surveys, 2021, 54(7): Article No.146. 33 Seo W, Zhang H J, Zhang Y Y, et al. GenCRF: generative clustering and reformulation framework for enhanced intent-driven information retrieval[PP/OL]. V1. arXiv (2024-09-17). https://doi.org/10.48550/arXiv.2409.10909. 34 Li L, Zhang X X, Zhou X, et al. AutoMIR: effective zero-shot medical information retrieval without relevance labels[C]// Findings of the Association for Computational Linguistics: EMNLP 2025. Stroudsburg: Association for Computational Linguistics, 2025: 24028-24047. 35 Zhang N Y, Chen M S, Bi Z, et al. CBLUE: a Chinese biomedical language understanding evaluation benchmark[C]// Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2022: 7888-7915. 36 Pokrywka J. Passage retrieval of Polish texts using OKAPI BM25 and an ensemble of cross encoders[C]// Proceedings of the 18th Conference on Computer Science and Intelligence Systems. Piscataway: IEEE, 2023: 1265-1269. 37 Izacard G, Caron M, Hosseini L, et al. Unsupervised dense information retrieval with contrastive learning[PP/OL]. V4. arXiv (2022-08-29). https://doi.org/10.48550/arXiv.2112.09118. 38 Wang Y, Sun Q, He S. M3E: Moka massive mixed embedding model[EB/OL]. (2023-06-07) [2025-08-12]. https://huggingface.co/moka-ai/m3e-base. 39 Xiao S T, Liu Z, Zhang P T, et al. C-pack: packed resources for general Chinese embeddings[C]// Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM Press, 2024: 641-649. 40 Li Z H, Zhang X, Zhang Y Z, et al. Towards general text embeddings with multi-stage contrastive learning[PP/OL]. V1. arXiv (2023-08-07). https://doi.org/10.48550/arXiv.2308.03281. 41 Wu T, Qin Y L, Zhang E W, et al. Towards robust text retrieval with progressive learning[PP/OL]. V1. arXiv (2023-11-20). https://doi.org/10.48550/arXiv.2311.11691. 42 Wang L, Yang N, Huang X L, et al. Multilingual E5 text embeddings: a technical report[PP/OL]. V1. arXiv (2024-02-08). https://doi.org/10.48550/arXiv.2402.05672. 43 Zhang X X, Li L, Zhou X, et al. R2MED: a benchmark for reasoning-driven medical retrieval[PP/OL]. V2. arXiv (2026-04-03). https://doi.org/10.48550/arXiv.2505.14558.