APPLICATION OF RULE-BASED METHOD FOR AUTOMATIC EXTRACTION OF TAGS FROM COLUMN-STYLE PDF-DOCUMENTS

Автор(и)

DOI:

https://doi.org/10.26577/jpcsit2025336

Ключові слова:

PDF parsing, Rule-based extraction, Metadata extraction, Document structure recognition, Text mining, Low-resource language processing, Knowledge base for LLMs

Анотація

Бұл зерттеуде қазақ тіліндегі газет мақалаларының PDF-нұсқаларынан құрылымдалған метадеректерді автоматты түрде алу үшін ережеге негізделген (rule-based) гибридті пайплайн ұсынылады. Зерттеу негізінен ұлттық «Egemen Qazaqstan» газетіне бағытталған. Негізгі мақсат – болашақта үлкен тілдік модельдерді (LLM) оқыту және Қазақстанда деректер журналистикасы үшін ЖИ-негізделген көмекші құруға арналған машинамен оқылатын білім базасын әзірлеу. Ұсынылған пайплайн үш ашық бастапқы кітапхананы — pdfminer.six, PyMuPDF және pdfplumber — біріктіреді және оларды келесі негізгі элементтерді алу үшін қолданады: тақырып, автор, күні, аннотация, мәтін, басылым атауы және категория. Алу сапасын бағалау үшін автоматты түрде алынған нәтижелер нақты газет шығарылымдарынан қолмен таңбаланған эталондық файлдармен салыстырылды. Бағалау үш толықтырушы метриканы қолданды: Precision, Textual Semantic Similarity (TSS) және Holistic Precision, олар дәл сәйкестіктермен қатар мағыналық ұқсастықтарды да ескеруге мүмкіндік береді. Эксперимент нәтижелері тегтердің көбісі – әсіресе құрылымдалған өрістер (күн, басылым атауы, категория) – Holistic Precision бойынша мінсіз нәтижеге (1.00) жеткенін көрсетті, ал ауыспалы өрістер, мысалы тақырып, 0.85-тен жоғары көрсеткіш көрсетті. Тестілеу жүргізілгеннен кейін бұл пайплайн 2017 жылдан 2025 жылғы наурызға дейін жарияланған 2 140 газет PDF файлына қолданылып, 159 135 мақала JSON құрылымына сәтті түрлендірілді. Бұл толықтырылған корпус қазақ тілді ЖИ жүйелерін дамытуда, журналистика мен медиа-талдауда негіз бола алады.

Завантажити

Дані для завантаження поки недоступні.

Біографії авторів

  • автор Assel Ospan, афіліація Al Farabi Kazakh National University, Almaty, Kazakhstan

    Assel Ospan is a senior lecturer at the Department of Artificial Intelligence and Big Data, al-Farabi Kazakh National University (Almaty, Kazakhstan, assel.ospan@kaznu.edu.kz). Her research focuses on the development of large language models for the Kazakh language, intelligent information extraction, and knowledge base construction. She actively participates in national AI research initiatives and has authored several publications on NLP and data journalism.
    ORCID iD: 0000-0002-1860-6997.

  • автор Madina Mansurova, афіліація Al Farabi Kazakh National University, Almaty, Kazakhstan

    Madina Mansurova is the head of the Department of Artificial Intelligence and Big Data, Professor, al-Farabi Kazakh National University (Almaty, Kazakhstan, madina.mansurova@kaznu.edu.kz). She has been successfully working in higher education and actively contributing to the advancement of new technologies. Prof.Mansurova is the author of more than 100 scientific articles, 10 monographs, 2 textbooks approved by the Ministry of Education and Science of the Republic of Kazakhstan, 5 patents for useful models in the field of automation and control, and over 40 copyrights on intellectual property. Since 2012, she has been the scientific supervisor of grant and program-targeted funding projects of the Ministry of Education and Science of the Republic of Kazakhstan. She has published 108 articles indexed in Scopus and Web of Science, with a Hirsch index of 7 in Scopus and 205 citations..
    ORCID iD: 0000-0002-9680-2758

  • автор Kanat Auyesbay, афіліація Al Farabi Kazakh National University, Almaty, Kazakhstan

    Kanat Auyesbay is the Dean of the Faculty of Journalism at Al-Farabi Kazakh National University (Almaty, Kazakhstan, kanat.auyesbay@kaznu.edu.kz). He is a journalist-educator who bridges the fields of media and higher education. Dr. Auesbay holds a Candidate of Philological Sciences degree (equivalent to PhD) and has extensive experience in both media production and academic leadership. As a recipient of the Bolashak International Scholarship, Kanat Auesbay completed a research and teaching internship at the University of East Anglia, UK (Norwich, 2013–2014). He served as Chairman of the State Attestation Commission at the Faculty of Journalism and Political Science of L.N. Gumilyov Eurasian National University (2023–2024). Since 2018, he has been a corresponding member of the Kazakhstan Academy of Pedagogical Sciences and a member of the Educational-Methodical Association under the Republican Educational-Methodical Council (ROƏK) for Journalism and Information (2019–2021). He has also served on the expert commission for training specialists abroad under the Bolashak program and has supervised and reviewed numerous theses and doctoral dissertations in media studies.  ORCID iD: 0009-0001-3529-9888

  • автор Talshyn Sarsembayeva, афіліація Al Farabi Kazakh National University, Almaty, Kazakhstan

    Talshyn Sarsembayeva is a a senior lecturer at the Department of Artificial Intelligence and Big Data, al-Farabi Kazakh National University (Almaty, Kazakhstan, talshyn.sagdatbek@kaznu.edu.kz). Her work focuses on the integration of artificial intelligence and data processing tools in journalistic practice. She has contributed to projects involving the structuring of large-scale media archives and the development of AI-assisted systems for Kazakh-language content. ORCID iD: 0000-0001-7668-2640.

  • автор Aman Mussa, афіліація Al Farabi Kazakh National University, Almaty, Kazakhstan

    Aman Mussa is a research assistant at the Department of Artificial Intelligence and Big Data, al-Farabi Kazakh National University (Almaty, Kazakhstan, mussa.aman0519@gmail.com). He is engaged in the development of rule-based and hybrid NLP pipelines, with a focus on Kazakh-language PDF processing. His work supports large-scale knowledge base generation for intelligent assistants in data journalism. ORCID iD: 0009-0001-9972-7677.

Завантаження

Опубліковано

2025-10-02