IndoBERTSkill: pretrained domain-specific language model for recognition Indonesian skill

(1) Meilany Nonsi Tentua Mail (Universitas PGRI Yogyakarta, Indonesia)
(2) * Suprapto Suprapto Mail (Universitas Gadjah Mada, Indonesia)
(3) Afiahayati Afiahayati Mail (Universitas Gadjah Mada, Indonesia)
*corresponding author

Abstract


The pretrained language model in Indonesian is already available for natural language processing tasks. However, this pre-trained model has been trained on Indonesian text, which has a different structure from the job description. Due to this, the pre-trained language model is less effective for skill recognition purposes. IndoBERTSkill is a novel pre-trained domain-specific language model that recognizes Indonesian language skills. It is built on the Bidirectional Encoder Representations from Transformers (BERT) architecture. IndoBERTSkill was trained on an extensive collection of Indonesian language texts from the Indonesian Wikipedia, the English Wikipedia, and Indonesian job descriptions from the job portal. IndoBERTSkill's performance was evaluated through two main approaches: (1) language modeling via Masked Language Model (MLM) prediction, and (2) fine-tuning on a custom annotated dataset (NERSkill) for Named Entity Recognition (NER) tasks. The fine-tuning process involved training a classification layer on top of the IndoBERTSkill model using BIO tagging to identify hard skills, soft skills, and technology entities. Similarly, the skill recognition model derived from IndoBERTSkill exhibits the highest F1-Score among various pre-trained language models, precisely at 87%, thus demonstrating robustness and strong generalizability for skill entity recognition in Indonesian job descriptions. IndoBERTSkill provides valuable resources for developing Indonesian natural language processing applications that require skills introduction. This could increase the accuracy and efficiency of skills recognition across various domains, including job matching, education, and training.

Keywords


BERT, Skill Recognition, Named Entity Recognition, domain-specific, Pretrained Language Model

   

DOI

https://doi.org/10.26555/ijain.v12i2.1287
      

Article metrics

Abstract views : 188 | PDF views : 8

   

Cite

   

Full Text

Download

References


[1] N. M. Gardazi, A. Daud, M. K. Malik, A. Bukhari, T. Alsahfi, and B. Alshemaimri, “BERT applications in natural language processing: a review,” Artif. Intell. Rev., vol. 58, no. 6, p. 166, Mar. 2025, doi: 10.1007/s10462-025-11162-5.

[2] J. Golec and T. Hachaj, “Ten Natural Language Processing Tasks with Generative Artificial Intelligence,” Appl. Sci., vol. 15, no. 16, p. 9057, Aug. 2025, doi: 10.3390/app15169057.

[3] A. Vaswani et al., “Attention is All you Need,” in Advances in Neural Information Processing Systems, 2017, p. 11. doi: 10.48550/arXiv.1706.03762.

[4] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.

[5] Y. Zhang et al., “A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2024, pp. 8783–8817. doi: 10.18653/v1/2024.emnlp-main.498.

[6] I. D. Mienye, N. Jere, G. Obaido, O. O. Ogunruku, E. Esenogho, and C. Modisane, “Large language models: an overview of foundational architectures, recent trends, and a new taxonomy,” Discov. Appl. Sci., vol. 7, no. 9, p. 1027, Sep. 2025, doi: 10.1007/s42452-025-07668-w.

[7] X. Jiang, W. Wang, S. Tian, H. Wang, T. Lookman, and Y. Su, “Applications of natural language processing and large language models in materials discovery,” npj Comput. Mater., vol. 11, no. 1, p. 79, Mar. 2025, doi: 10.1038/s41524-025-01554-0.

[8] N. Abdel Samee, M. Alabdulhafith, S. Muhammad Ahmed Hassan Shah, and A. Rizwan, “JusticeAI: A Large Language Models Inspired Collaborative and Cross-Domain Multimodal System for Automatic Judicial Rulings in Smart Courts,” IEEE Access, vol. 12, pp. 173091–173107, 2024, doi: 10.1109/ACCESS.2024.3491775.

[9] Y. Mao, Q. Liu, and Y. Zhang, “Sentiment analysis methods, applications, and challenges: A systematic literature review,” J. King Saud Univ. - Comput. Inf. Sci., vol. 36, no. 4, p. 102048, Apr. 2024, doi: 10.1016/j.jksuci.2024.102048.

[10] M. R. R. Rana, A. Nawaz, S. U. Rehman, M. A. Abid, M. Garayevi, and J. Kajanová, “BERT-BiGRU-Senti-GCN: An Advanced NLP Framework for Analyzing Customer Sentiments in E-Commerce,” Int. J. Comput. Intell. Syst., vol. 18, no. 1, p. 21, Feb. 2025, doi: 10.1007/s44196-025-00747-1.

[11] D. Herhausen, S. Ludwig, E. Abedin, N. U. Haque, and D. de Jong, “From words to insights: Text analysis in business research,” J. Bus. Res., vol. 198, p. 115491, Sep. 2025, doi: 10.1016/j.jbusres.2025.115491.

[12] X. Xu, Z. Li, H. Zhang, and K. Ma, “Named entity recognition for Chinese electronic medical records by integrating knowledge graph and ClinicalBERT,” Front. Artif. Intell., vol. 8, p. 1634774, Sep. 2025, doi: 10.3389/frai.2025.1634774.

[13] D. Khurana, A. Koli, K. Khatter, and S. Singh, “Natural language processing: state of the art, current trends and challenges,” Multimed. Tools Appl., vol. 82, no. 3, pp. 3713–3744, Jan. 2023, doi: 10.1007/s11042-022-13428-4.

[14] T. Aggarwal, A. Salatino, F. Osborne, and E. Motta, “Large language models for scholarly ontology generation: An extensive analysis in the engineering field,” Inf. Process. Manag., vol. 63, no. 1, p. 104262, Jan. 2026, doi: 10.1016/j.ipm.2025.104262.

[15] K. Arai, “Design of On-Premises Version of RAG with AI Agent for Framework Selection Together with Dify and DSL as Well as Ollama for LLM,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 12, pp. 117–124, Dec. 2024, doi: 10.14569/IJACSA.2024.0151212.

[16] D. Bzdok, A. Thieme, O. Levkovskyy, P. Wren, T. Ray, and S. Reddy, “Data science opportunities of large language models for neuroscience and biomedicine,” Neuron, vol. 112, no. 5, pp. 698–717, Mar. 2024, doi: 10.1016/j.neuron.2024.01.016.

[17] B. Chen, Z. Zhang, N. Langrené, and S. Zhu, “Unleashing the potential of prompt engineering for large language models,” Patterns, vol. 6, no. 6, p. 101260, Jun. 2025, doi: 10.1016/j.patter.2025.101260.

[18] A. Miftaroski, R. Zowalla, M. Wiesner, and M. Pobiruchin, “Leveraging Large Language Models to Improve the Readability of German Online Medical Texts: Evaluation Study,” JMIR AI, vol. 5, no. 1, pp. e77149–e77149, Jan. 2026, doi: 10.2196/77149.

[19] Y. Wang, M. Lin, Q. Hu, S. Bai, Y. Li, and L. Bao, “A domain-specific cross-lingual semantic alignment learning model for low-resource languages,” Neural Networks, vol. 194, no. Februari, p. 108114, Feb. 2026, doi: 10.1016/j.neunet.2025.108114.

[20] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Stroudsburg, PA, USA: International Committee on Computational Linguistics, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.

[21] B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 843–857. doi: 10.18653/v1/2020.aacl-main.85.

[22] F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 10660–10668. doi: 10.18653/v1/2021.emnlp-main.833.

[23] S. M. Octafia, “IndoBERT-SupCon: A Supervised Contrastive Learning Model for Analyzing Public Perception on Halal Tourism,” J. Appl. Data Sci., vol. 7, no. 1, pp. 218–231, Jan. 2026, doi: 10.47738/jads.v7i1.1045.

[24] N. C. Mei, S. Tiun, and G. Sastria, “Multi-Label Aspect-Sentiment Classification on Indonesian Cosmetic Product Reviews with IndoBERT Model,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 11, pp. 712–720, Nov. 2024, doi: 10.14569/IJACSA.2024.0151168.

[25] A. Riyadi, M. Kovacs, U. Serdült, and V. Kryssanov, “IndoGovBERT: A Domain-Specific Language Model for Processing Indonesian Government SDG Documents,” Big Data Cogn. Comput., vol. 8, no. 11, p. 153, Nov. 2024, doi: 10.3390/bdcc8110153.

[26] S. O. Khairunnisa, Z. Chen, and M. Komachi, “Improving Domain-Specific NER in the Indonesian Language Through Domain Transfer and Data Augmentation,” J. Adv. Comput. Intell. Intell. Informatics, vol. 28, no. 6, pp. 1299–1312, Nov. 2024, doi: 10.20965/jaciii.2024.p1299.

[27] E. Yulianti, N. Bhary, J. Abdurrohman, F. W. Dwitilas, E. Q. Nuranti, and H. S. Husin, “Named entity recognition on Indonesian legal documents: a dataset and study using transformer-based models,” Int. J. Electr. Comput. Eng., vol. 14, no. 5, p. 5489, Oct. 2024, doi: 10.11591/ijece.v14i5.pp5489-5501.

[28] M. Laugiwa and E. Yulianti, “Question Answering through Transfer Learning on Closed-Domain Educational Websites,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 9, no. 1, pp. 104–110, Feb. 2025, doi: 10.29207/resti.v9i1.6163.

[29] S. Thavasi and T. Revathi, “A personalized machine learning–based system to evaluate students’ skillset and analyze the gap between academia and industry for engineering students,” Kybernetes, vol. 54, no. 9, pp. 4850–4865, Oct. 2025, doi: 10.1108/K-01-2024-0053.

[30] N. P, M. V, K. K. Ram, P. Vamsi, V. V. Vardhan Reddy, and D. Sridhar, “Sentimental Analysis of Job Classification using Machine Learning Algorithms,” in 2023 8th International Conference on Communication and Electronics Systems (ICCES), IEEE, Jun. 2023, pp. 1279–1284. doi: 10.1109/ICCES57224.2023.10192717.

[31] M. A. Berawi, M. Sari, N. H. A. Amiri, suci indah Susilowati, S. R. Utami, and A. Kulachinskaya, “Developing a Machine Learning Model to Improve the Accuracy of Owner Estimate Cost in the Capital Expenditure Procurement Process,” Int. J. Technol., vol. 16, no. 4, p. 1179, Jul. 2025, doi: 10.14716/ijtech.v16i4.7409.

[32] K. Jian and M. A. A. D. Dizon, “Analysis of employment trends and prediction model for Chinese college graduates based on natural language processing,” Edelweiss Appl. Sci. Technol., vol. 9, no. 4, pp. 1145–1159, Apr. 2025, doi: 10.55214/25768484.v9i4.6188.

[33] W. Fang, Y. Xu, T. Zhang, Y. Wang, and L. Zheng, “Towards large language model for cognitive industrial mixed reality: A survey,” Adv. Eng. Informatics, vol. 71, no. April, p. 104276, Apr. 2026, doi: 10.1016/j.aei.2025.104276.

[34] S. Fareri, R. Apreda, V. Mulas, and R. Alonso, “The worker profiler: Assessing the digital skill gaps for enhancing energy efficiency in manufacturing,” Technol. Forecast. Soc. Change, vol. 196, no. November, p. 122844, Nov. 2023, doi: 10.1016/j.techfore.2023.122844.

[35] X. Q. Ong and K. Hui Lim, “SkillRec: A Data-Driven Approach to Job Skill Recommendation for Career Insights,” in 2023 15th International Conference on Computer and Automation Engineering (ICCAE), IEEE, Mar. 2023, pp. 40–44. doi: 10.1109/ICCAE56788.2023.10111438.

[36] P. Achananuparp, Y. Xu, Y. Lu, X. J. S. Ashok, and E.-P. Lim, “Leveraging large language models for career mobility analysis: a study of gender, race, and job change using U.S. online resume profiles,” EPJ Data Sci., vol. 15, no. 1, p. 4, Jan. 2026, doi: 10.1140/epjds/s13688-025-00607-0.

[37] R. Alonso, D. Dessí, A. Meloni, and D. Reforgiato Recupero, “A novel approach for job matching and skill recommendation using transformers and the O*NET database,” Big Data Res., vol. 39, no. February, p. 100509, Feb. 2025, doi: 10.1016/j.bdr.2025.100509.

[38] N. Matkin et al., “Correction to: Comparative Analysis of Encoder-Based NER and Large Language Models for Skill Extraction from Russian Job Vacancies,” in Communications in Computer and Information Science, vol. 2364 CCIS, Springer Science and Business Media Deutschland GmbH, 2026, pp. C1–C1. doi: 10.1007/978-3-031-97019-1_13.

[39] R. Wang, Q. Chen, Y. Wang, L. Xiong, and B. Shen, “JobViz: Skill-driven visual exploration of job advertisements,” Vis. Informatics, vol. 8, no. 3, pp. 18–28, Sep. 2024, doi: 10.1016/j.visinf.2024.07.001.

[40] K. Nguyen, M. Zhang, S. Montariol, and A. Bosselut, “Rethinking Skill Extraction in the Job Market Domain using Large Language Models,” in Proceedings of the First Workshop on Natural Language Processing for Human Resources (NLP4HR 2024), Stroudsburg, PA, USA: Association for Computational Linguistics, 2024, pp. 27–42. doi: 10.18653/v1/2024.nlp4hr-1.3.

[41] M. Hasan Nizami, A. Uddin, M. T. Salani, A. Saeed, F. Alvi, and A. Samad, “Enhancing Human Capital Management: AI Techniques for Candidate Matching and Skill Extraction Notebook for the TalentCLEF Lab at CLEF 2025,” 2025, p. 11. Accessed: May 24, 2026. [Online]. Available: https://www.semanticscholar.org/paper/Enhancing-Human-Capital-Management%3A-AI-Techniques-Nizami-Uddin/28430020b7ad812bc59ff8555b036edee1bbd208.

[42] M. Lukauskas, V. Šarkauskaitė, V. Pilinkienė, A. Stundžienė, A. Grybauskas, and J. Bruneckienė, “Enhancing Skills Demand Understanding through Job Ad Segmentation Using NLP and Clustering Techniques,” Appl. Sci., vol. 13, no. 10, p. 6119, May 2023, doi: 10.3390/app13106119.

[43] M. N. Tentua, Suprapto, and Afiahayati, “NERSkill.Id: Annotated dataset of Indonesian’s skill entity recognition,” Data Br., vol. 53, no. April, p. 110192, Apr. 2024, doi: 10.1016/j.dib.2024.110192.

[44] H. Kavas, M. Serra-Vidal, and L. Wanner, “Multilingual Skill Extraction for Job Vacancy–Job Seeker Matching in Knowledge Graphs,” 2025. Accessed: May 24, 2026. [Online]. Available: https://aclanthology.org/2025.genaik-1.15/.




Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571  (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
 andri.pranolo.id@ieee.org (publication issues)

View IJAIN Stats

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0