SmartRGPT: An NLP-Based Conversational Research Support System Designed Using Open- Source RAG and LLM

Authors

DOI:

https://doi.org/10.17821/srels/2026/v63i3/172070

Keywords:

Conversational AI, Generative AI, LangChain, Large Language Models (LLM), Llama-3, Natural Language Processing (NLP), Retrieval Augmented Generation (RAG), SmartRGPT, SmartRLibrary

Abstract

This case study research aims to focus on designing a conversational Research Support System (RSS) named SmartRGPT using open-source Retrieval Augmented Generation (RAG) and a Large Language Model (LLM) to enhance and address the challenges faced by researchers while using the existing Research Support Services called SmartRLibrary (initiated by and applied in B C Roy Memorial Library, alternatively, IIM Calcutta Library). It also addresses the limitations of traditional keyword-based search and the hallucination issues of standalone LLM. The prototype has been designed using several open-source software components, including the RAG pipeline, LangChain, the ChromaDB vector database, and the Llama-3 (70-billion-parameter model). A curated set of over 250 datasets was collected, preprocessed, and ingested using Wget (WarcGPT framework) for preparing the knowledge base. The prototype was tested and evaluated using real-world queries. Based on internal review and initial observations of the authors on the generated responses, in the majority of tested cases, the findings demonstrate that the proposed system generated accurate, context‑aware responses without hallucinations. It has responded to short and long-range queries based on its ingested knowledge bases, citing the sources as references. The findings further indicate that the proposed system has the potential to provide 24/7 personalised research assistance, reduce repetitive library workload, and enable the library to provide more advanced services if applied after rigorous evaluation in larger populations. Its cost-effective open-source architecture also offers libraries with limited budgets an independent and customisable alternative to vendor-dependent solutions, thereby contributing to the advancement of the Library and Information Science (LIS) domain.

Downloads

Download data is not yet available.

Published

2026-07-31

How to Cite

Mazumder, J., Datta, A., Roy, S., & Ghosh, N. C. (2026). SmartRGPT: An NLP-Based Conversational Research Support System Designed Using Open- Source RAG and LLM. Journal of Information and Knowledge, 63(3), 147–156. https://doi.org/10.17821/srels/2026/v63i3/172070

Issue

Section

Articles

References

Agrawal, G., Kumarage, T., Alghamdi, Z., & Liu, H. (2024). Can knowledge graphs reduce hallucinations in LLMs? A survey. arXiv. https://doi.org/10.18653/v1/2024.naacllong. 219

Balakrishnan, G., & Purwar, A. (2024). Evaluating the efficacy of open-source LLMs in enterprise-specific RAG systems: A comparative study of performance and scalability. 2024 IEEE 21st India Council International Conference, 1-9. https://doi.org/10.1109/INDICON63790.2024.10958508

Bevara, R. V. K., Lund, B. D., Mannuru, N. R., Karedla, S. P., Mohammed, Y., Kolapudi, S. T., & Mannuru, A. (2025). Prospects of retrieval-augmented generation (RAG) for academic library search and retrieval. Information Technology and Libraries, 44(2). https://doi.org/10.5860/ ital.v44i2.17361

Boateng, F. (2025). The transformative potential of generative AI in academic library access services: Opportunities and challenges. Information Services and Use, 45(1-2), 140-147. https://doi.org/10.1177/18758789251332800

Chen, J., Lin, H., Han, X., & Sun, L. (2024). Benchmarking large language models in retrieval-augmented generation. Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence, 38, 17754-17762. https://doi. org/10.1609/aaai.v38i16.29728

Cheung, H. C., Lo, Y. Y. M., Chiu, D. K. W., & Kong, E. W. S. (2023). Development of smart academic library services with internet of things technology: A qualitative study in Hong Kong. Library Hi Tech, 43(1), 398-422. https://doi. org/10.1108/LHT-06-2023-0219

Cox, A. (2022). The ethics of AI for information professionals: Eight scenarios. Journal of the Australian Library and Information Association, 71(3), 201-214. https://doi.org/10 .1080/24750158.2022.2084885

Cox, A. M., Pinfield, S., & Rutter, S. (2018). The intelligent library: Thought leaders’ views on the likely impact of artificial intelligence on academic libraries. Library Hi Tech, 37(3), 418-435. https://doi.org/10.1108/LHT-08-2018-0105

Dwivedi, Y. K., Kshetri, N., Hughes, L., Slade, E. L., Jeyaraj, A., Kar, A. K., Baabdullah, A. M., Koohang, A., Raghavan, V., Ahuja, M., Albanna, H., Albashrawi, M. A., Al-Busaidi, A. S., Balakrishnan, J., Barlette, Y., Basu, S., Bose, I., Brooks, L., Buhalis, D., … Wright, R. (2023). Opinion paper: “So what if ChatGPT wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. International Journal of Information Management, 71, 102642. https://doi.org/10.1016/j.ijinfomgt.2023.102642

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv. https://doi.org/10.48550/arXiv.2312.10997

Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Computation and Language. https://doi.org/10.48550/arXiv.2005.11401

Li, L., & Coates, K. (2024). Academic library online chat services under the impact of artificial intelligence. Information Discovery and Delivery, 53(2), 192-205. https:// doi.org/10.1108/IDD-11-2023-0143

Mazumder, J., & Mukhopadhyay, P. (2024). Designing questionanswer- based search system in libraries: Application of open source retrieval augmented generation (RAG) pipeline. Journal of Information and Knowledge, 255-260. https://doi.org/10.17821/srels/2024/v61i5/171583

Mukhopadhyay, P. (2026). Optimizing retrieval in libraries through RAG: A framework. Indian Journal of Information Library and Society, 37(1-2), 6-22.

Ovadia, O., Brief, M., Mishaeli, M., & Elisha, O. (2024). Fine-tuning or retrieval? Comparing knowledge injection in LLMs. arXiv. https://doi.org/10.18653/v1/2024.emnlp-main.15

Safdar, M., Siddique, N., Gulzar, A., Yasin, H., & Khan, M. A. (2024). Does ChatGPT generate fake results? Challenges in retrieving content through ChatGPT. Digital Library Perspectives, 40(4), 668-680. https://doi.org/10.1108/DLP-01-2024-0006

Vakilzadeh, H., & Wood, D. A. (2025). The development of a RAGbased artificial intelligence research assistant. Social Science Research Network. https://doi.org/10.2139/ssrn.5283702

Yan, S.-Q., Gu, J.-C., Zhu, Y., & Ling, Z.-H. (2024). Corrective retrieval augmented generation. arXiv. https://doi. org/10.2139/ssrn.5267341

Zheng, X., Li, Z., Chen, Q., & Zhang, Y. (2025). Beyond decomposition: Hierarchical dependency management in multi-document question answering. Journal of the Association for Information Science and Technology, 76(5), 770-789. https://doi.org/10.1002/asi.24971