Toward effective Text-to-MongoDB query translation

(1) * Aicha Aggoune Mail (LabSTIC, Computer Science Department, University of 8th May 1945, Algeria)
*corresponding author

Abstract


Translating natural language questions into MongoDB queries is critical for flexible data access in current NoSQL systems. However, semantic ambiguity in user questions and MongoDB's dynamic schema make this work difficult. This study presents QMQL (Question to Mongo Query Language), a hybrid approach meant to address these challenges. QMQL combines a Graph Attention Network (GAT) for refining schema elements with a Retrieval-Augmented Generation (RAG) mechanism that employs BERT embeddings to retrieve relevant schema and resolve semantic ambiguity. A T5-base model is used to generate a MongoDB query corresponding to the user’s question. An experimental evaluation on an extended dataset encompassing various real-world domains demonstrates the effectiveness of the proposed approach. QMQL achieves excellent performance with an EMA of 0.89, an EM of 0.91, and a BLEU score of 0.95, exceeding previous approaches, particularly for semantically ambiguous questions and sophisticated queries across flexible MongoDB schemas,

Keywords


Flexible schema; MongoDB querying; Query translation; RAG-SBERT-T5-base; Semantic ambiguity.

   

DOI

https://doi.org/10.26555/ijain.v12i2.2360
      

Article metrics

Abstract views : 195 | PDF views : 3

   

Cite

   

Full Text

Download

References


[1] M. Rathore and S. S. Bagui, “MongoDB: Meeting the Dynamic Needs of Modern Applications,” Encyclopedia, vol. 4, no. 4, pp. 1433–1453, Sep. 2024, doi: 10.3390/encyclopedia4040093.

[2] React Query Builder, “React Query Builder Documentation.” [Online]. Available: https://react-querybuilder.js.org/.

[3] M. Al Khatib, A. Alshammari, and M. Al-Rakhami, “Rethinking Concept Drift Detection in Data Streams: A Large Language Model-Based Adaptive Framework,” in Proceedings of the LSGDA Workshop at VLDB 2024, 2024, pp. 81–90. [Online]. Available: https://vldb.org/workshops/2024/proceedings/LSGDA/LSGDA24.09.

[4] J. Qi et al., “RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, pp. 3215–3229. doi: 10.18653/v1/2022.emnlp-main.211.

[5] L. Shi, Z. Tang, N. Zhang, X. Zhang, and Z. Yang, “A Survey on Employing Large Language Models for Text-to-SQL Tasks,” ACM Comput. Surv., vol. 58, no. 2, pp. 1–37, Jan. 2026, doi: 10.1145/3737873.

[6] M. Pourreza et al., “CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL,” Oct. 2024, pp. 1–30. Accessed: May 31, 2026. [Online]. Available: http://arxiv.org/abs/2410.01943.

[7] M. Pourreza and D. Rafiei, “DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction,” in Advances in Neural Information Processing Systems 36, San Diego, California, USA: Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2023, pp. 36339–36348. doi: 10.52202/075280-1577.

[8] Q. Li, T. You, J. Chen, Y. Zhang, and C. Du, “LI-EMRSQL: Linking Information Enhanced Text2SQL Parsing on Complex Electronic Medical Records,” IEEE Trans. Reliab., vol. 73, no. 2, pp. 1280–1290, Jun. 2024, doi: 10.1109/TR.2023.3336330.

[9] G. Jeong et al., “Improving Text-to-SQL with a Hybrid Decoding Method,” Entropy, vol. 25, no. 3, p. 513, Mar. 2023, doi: 10.3390/e25030513.

[10] W. Zhang, K. Zeng, X. Yang, T. Shi, and P. Wang, “Text-to-ESQ: A Two-Stage Controllable Approach for Efficient Retrieval of Vaccine Adverse Events from NoSQL Database,” in Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics, New York, NY, USA: ACM, Sep. 2023, pp. 1–10. doi: 10.1145/3584371.3613008.

[11] S. Benhur, N. Nayan, P. Manoharan, A. Venugopalan, K. S. Rajabhupati, and N. Sundaramahalingam, “NL2EQ: Generating Elasticsearch Query DSL from Natural Language Text Using Large Language Models,” in Lecture Notes in Networks and Systems, vol. 1424 LNNS, Springer Science and Business Media Deutschland GmbH, 2025, pp. 223–242. doi: 10.1007/978-3-031-92605-1_15.

[12] K. M. Hossen, M. N. Uddin, M. Arefin, and M. A. Uddin, “BERT Model-based Natural Language to NoSQL Query Conversion using Deep Learning Approach,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 2, pp. 810–821, Feb. 2023, doi: 10.14569/IJACSA.2023.0140293.

[13] M. Hornsteiner, M. Kreussel, C. Steindl, F. Ebner, P. Empl, and S. Schönig, “Real-Time Text-to-Cypher Query Generation with Large Language Models for Graph Databases,” Futur. Internet, vol. 16, no. 12, p. 438, Nov. 2024, doi: 10.3390/fi16120438.

[14] M. Kobeissi, N. Assy, W. Gaaloul, B. Defude, B. Benatallah, and B. Haidar, “Natural language querying of process execution data,” Inf. Syst., vol. 116, no. June, p. 102227, Jun. 2023, doi: 10.1016/j.is.2023.102227.

[15] J. Lu, Y. Song, Z. Qin, H. Zhang, C. Zhang, and R. C.-W. Wong, “Bridging the Gap: Enabling Natural Language Queries for NoSQL Databases through Text-to-NoSQL Translation,” pp. 1–16, Feb. 2025, Accessed: May 31, 2026. [Online]. Available: http://arxiv.org/abs/2502.11201.

[16] Z. Qin, Y. Song, J. Lu, Y. Song, S. Li, and C. J. Zhang, “MultiTEND: A Multilingual Benchmark for Natural Language to NoSQL Query Translation,” in Findings of the Association for Computational Linguistics: ACL 2025, Stroudsburg, PA, USA: Association for Computational Linguistics, Aug. 2025, pp. 24632–24657. doi: 10.18653/v1/2025.findings-acl.1265.

[17] P. Do, T. H. V. Phan, and B. B. Gupta, “Developing a Vietnamese Tourism Question Answering System Using Knowledge Graph and Deep Learning,” ACM Trans. Asian Low-Resource Lang. Inf. Process., vol. 20, no. 5, pp. 1–18, Sep. 2021, doi: 10.1145/3453651.

[18] Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and P. S. Yu, “A Comprehensive Survey on Graph Neural Networks,” IEEE Trans. Neural Networks Learn. Syst., vol. 32, no. 1, pp. 4–24, Jan. 2021, doi: 10.1109/TNNLS.2020.2978386.

[19] P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Advances in Neural Information Processing Systems, Neural information processing systems foundation, May 2020, pp. 9459–9474. Accessed: May 31, 2026. [Online]. Available: https://arxiv.org/pdf/2005.11401

[20] C. Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, Sep. 2023, Accessed: May 31, 2026. [Online]. Available: http://arxiv.org/abs/1910.10683

[21] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 3980–3990. doi: 10.18653/v1/D19-1410.

[22] A. Aggoune and Z. Mihoubi, “Towards Efficient Dataset Development: A Case Study of M2Q2+ in Movie QA Systems,” in 2024 International Conference on Advanced Aspects of Software Engineering (ICAASE), IEEE, Nov. 2024, pp. 1–5. doi: 10.1109/ICAASE64542.2024.10850962.

[23] Mukesh Reddy Dhanagari, “MongoDB and Data Consistency: Bridging the Gap between Performance and Reliability,” J. Comput. Sci. Technol. Stud., vol. 6, no. 2, pp. 183–198, Jun. 2024, doi: 10.32996/jcsts.2024.6.2.21.

[24] C. M. Sowandi, I. Leonita, S. Hartanto, P. Arisaputra, and D. David, “Exploring the Effectiveness, Features, and Compatibility of MongoDB and MySQL: A Comprehensive Comparison of NoSQL and Relational Databases,” MIND (Multimedia Artif. Intell. Netw. Database) J., vol. 8, no. 2, pp. 217–229, Dec. 2023, Accessed: May 31, 2026. [Online]. Available: https://ejurnal.itenas.ac.id/index.php/mindjournal/article/view/9145.

[25] M. Nuriev, R. Zaripova, O. Yanova, I. Koshkina, and A. Chupaev, “Enhancing MongoDB query performance through index optimization,” E3S Web Conf., vol. 531, p. 03022, Jun. 2024, doi: 10.1051/e3sconf/202453103022.

[26] MongoDB Inc., “Sample Mflix Dataset.” Accessed: Jun. 04, 2026. [Online]. Available: https://www.mongodb.com/docs/atlas/sample-data/sample-mflix/.

[27] MongoDB Inc., “Atlas Architecture Center.” Accessed: May 31, 2026. [Online]. Available: https://www.mongodb.com/docs/atlas/architecture/current/.

[28] N. Bansal, S. Sachdeva, and L. K. Awasthi, “Query-based denormalization using hypergraph (QBDNH): a schema transformation model for migrating relational to NoSQL databases,” Knowl. Inf. Syst., vol. 66, no. 1, pp. 681–722, Jan. 2024, doi: 10.1007/s10115-023-02017-y.

[29] S. Bharany, K. Kaur, S. E. M. Eltaher, A. O. Ibrahim, S. Sharma, and M. M. M. A. Elsalam, “A Comparative Study of Cloud Data Portability Frameworks for Analyzing Object to NoSQL Database Mapping from ONDM’s Perspective,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 10, pp. 805–814, Oct. 2023, doi: 10.14569/IJACSA.2023.0141086.

[30] A. Samydurai, K. Revathi, L. Karthikeyan, B. Vanathi, and K. Devi, “An Enhanced Entity Model for Converting Relational to Non-Relational Documents in Hospital Management System Based on Cloud Computing,” IETE Tech. Rev., vol. 39, no. 6, pp. 1449–1462, Nov. 2022, doi: 10.1080/02564602.2021.2016075.

[31] A. Aggoune and M. S. Namoune, “P3 Process for Object-Relational Data Migration to NoSQL Document-Oriented Datastore,” Int. J. Softw. Sci. Comput. Intell., vol. 14, no. 1, pp. 1–20, Sep. 2022, doi: 10.4018/IJSSCI.309994.

[32] A. Aggoune and M. S. Namoune, “Metadata-driven Data Migration from Object-relational Database to NoSQL Document-oriented Database,” Comput. Sci., vol. 23, no. 4, pp. 495–519, Nov. 2022, doi: 10.7494/csci.2022.23.4.4375.

[33] MongoDB Inc., “Sample Datasets.” [Online]. Available: https://www.mongodb.com/docs/atlas/sample-data/.

[34] A. G. Vrahatis, K. Lazaros, and S. Kotsiantis, “Graph Attention Networks: A Comprehensive Review of Methods and Applications,” Futur. Internet, vol. 16, no. 9, p. 318, Sep. 2024, doi: 10.3390/fi16090318.

[35] J. Li et al., “Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing,” Proc. AAAI Conf. Artif. Intell., vol. 37, no. 11, pp. 13076–13084, Jun. 2023, doi: 10.1609/aaai.v37i11.26536.

[36] Y. Yang, Z. Peng, F. Zhou, X. Yao, and Y. Zhang, “Conversational Text-to-SQL: A Comprehensive Survey of Paradigms, Challenges, and Future Directions,” in Communications in Computer and Information Science, vol. 2700 CCIS, Springer Science and Business Media Deutschland GmbH, 2025, pp. 1–19. doi: 10.1007/978-3-032-08049-3_1.

[37] [W. Zhao, L. Zhao, F. Wu, Z. Zheng, H. Jin, and B. Gu, “Enhancing Interaction Graph of Data Schema and Syntactic Structure with Pre-trained Language Model for Text-to-SQL,” in Communications in Computer and Information Science, vol. 2301, Springer Science and Business Media Deutschland GmbH, 2025, pp. 159–173. doi: 10.1007/978-981-96-1024-2_12.

[38] P. Lekheshwar Balley and S. V. Sonekar, “Design of an Integrated Model Using R-GCN, TPOT, and Transformers for Efficient NoSQL Data Processing and Analysis,” J. Neonatal Surg., vol. 14, no. 6S, pp. 298–314, Mar. 2025, doi: 10.52783/jns.v14.2237.

[39] D. Thakur, J. K. Saini, and S. Srinivasan, “DeepThink IoT: The Strength of Deep Learning in Internet of Things,” Artif. Intell. Rev., vol. 56, no. 12, pp. 14663–14730, Dec. 2023, doi: 10.1007/s10462-023-10513-4.

[40] H. Han et al., “Retrieval-Augmented Generation with Graphs (GraphRAG),” pp. 1–88, Jan. 2025, Accessed: Jun. 04, 2026. [Online]. Available: https://arxiv.org/pdf/2501.00309.

[41] Y. Gao, Y. Xiong, M. Wang, and H. Wang, “Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks,” pp. 1–17, Jul. 2024, Accessed: Jun. 04, 2026. [Online]. Available: http://arxiv.org/abs/2407.21059.

[42] W. X. Zhao, J. Liu, R. Ren, and J.-R. Wen, “Dense Text Retrieval Based on Pretrained Language Models: A Survey,” ACM Trans. Inf. Syst., vol. 42, no. 4, pp. 1–60, Jul. 2024, doi: 10.1145/3637870.

[43] E. Eldele et al., “Self-Supervised Contrastive Representation Learning for Semi-Supervised Time-Series Classification,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 12, pp. 15604–15618, Dec. 2023, doi: 10.1109/TPAMI.2023.3308189.

[44] E. Daraghmi, L. Atwe, and A. Jaber, “A Comparative Study of PEGASUS, BART, and T5 for Text Summarization Across Diverse Datasets,” Futur. Internet, vol. 17, no. 9, p. 389, Aug. 2025, doi: 10.3390/fi17090389.

[45] S. S. Gu, Y. Iwasawa, T. Kojima, Y. Matsuo, and M. Reid, “Large Language Models Are Zero-Shot Reasoners,” in Advances in Neural Information Processing Systems 35, San Diego, California, USA: Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2022, pp. 22199–22213. doi: 10.52202/068431-1613.

[46] A. Aggoune, “A Survey on Dataset Development Techniques for QA Systems ⋆,” in A Comparative Study of Retrieval-Augmented Generation Methods in Enterprise Knowledge Management, 2024, pp. 1–12. Accessed: Jun. 04, 2026. [Online]. Available: https://ceur-ws.org/Vol-3922/paper1.

[47] A. Aggoune and Z. Mihoubi, “M2Q2: A Text-to-MQL Dataset for Movie QA Systems,” in Lecture Notes in Networks and Systems, vol. 1393 LNNS, Springer Science and Business Media Deutschland GmbH, 2026, pp. 441–453. doi: 10.1007/978-3-031-90893-4_29.

[48] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph Attention Networks,” in International Conference on Learning Representations (ICLR), International Conference on Learning Representations, ICLR, Feb. 2018, pp. 1–12. Accessed: Jun. 04, 2026. [Online]. Available: http://arxiv.org/abs/1710.10903.

[49] G. Datta, N. Joshi, and K. Gupta, “Analysis of Automatic Evaluation Metric on Low-Resourced Language: BERTScore vs BLEU Score,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 13721 LNAI, Springer Science and Business Media Deutschland GmbH, 2022, pp. 155–162. doi: 10.1007/978-3-031-20980-2_14.

[50] A. Grattafiori et al., “The Llama 3 Herd of Models,” pp. 1–92, Nov. 2024, Accessed: Jun. 04, 2026. [Online]. Available: http://arxiv.org/abs/2407.21783.

[51] A. Taloni et al., “Comparative performance of humans versus GPT-4.0 and GPT-3.5 in the self-assessment program of American Academy of Ophthalmology,” Sci. Rep., vol. 13, no. 1, p. 18562, Oct. 2023, doi: 10.1038/s41598-023-45837-2.

[52] M. Shen, Y. Li, L. Chen, Z. Fan, Y. Li, and Q. Yang, “From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent,” pp. 1–21, Mar. 2026, Accessed: Jun. 04, 2026. [Online]. Available: http://arxiv.org/abs/2505.02024.

[53] Y. Wang, W. Wang, S. Joty, and S. C. H. Hoi, “CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 8696–8708. doi: 10.18653/v1/2021.emnlp-main.685.




Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571  (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
 andri.pranolo.id@ieee.org (publication issues)

View IJAIN Stats

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0