تحلیل موضوعی و روند تولیدات علمی سلامت در حوزه سواد اطلاعاتی با استفاده از فنون متن کاوی (مقاله علمی وزارت علوم)
درجه علمی: نشریه علمی (وزارت علوم)
آرشیو
چکیده
هدف: وجود پژوهش های متنوع درزمینه سواد اطلاعاتی، نیاز به تحلیل موضوعات این مطالعات را برای داشتن چشم اندازی روشن و جامع از این حوزه ضروری می سازد. پژوهش حاضر با هدف مدل سازی موضوعی مقالات منتشرشده درزمینه سواد اطلاعاتی متون سلامت با استفاده از پایگاه PubMed انجام شده است. روش: این پژوهش ازلحاظ رویکرد کمی و ازنظر نوع کاربردی است که با استفاده از روش های متن کاوی انجام گرفت. تولیدات علمی درزمینهه سواد اطلاعاتی مبتنی بر سرعنوان موضوعی مش و با استفاده از فرمول جستجوی "information literacy" [Majr] بدون محدودیت زمانی از پایگاه PubMed استخراج شدند. جستجو در تاریخ 15 مرداد 1403 انجام شد که تعداد 8407 رکورد از 1519 عنوان نشریه و کتاب بازیابی شد. سپس، چکیده و عنوان مقالات به فرمت تکست ذخیره و با هدف تحلیل، به صورت ساختاریافته به فرمت اکسل تبدیل شد. 7608 رکورد دارای چکیده بودند. 797 رکورد فاقد چکیده بودند که پس از حذف رکوردهای خالی تعداد 7608 رکورد برای تجزیه وتحلیل مورداستفاده قرار گرفتند. بدین منظور ابتدا توکن یابی انجام شد. سپس علائم نگارشی و ایست واژه ها نیز حذف شدند و در ادامه، ریشه یابی کلمات انجام گرفت و در مرحله بعد برای اعمال فنون یادگیری ماشینی در داده های متنی، ابتدا متون به بردارهای عددی تبدیل شدند و سرانجام با الگوریتم LDA مدل سازی موضوعات انجام گرفت. پس از پاک سازی داده ها، چکیده و عناوین این مقالات با استفاده از کتابخانه های Pandas، PyLDAvis، sklearn، PyLDAvis، numpy، Setuptools، NLTK، Gensim، Wordcloud، Seaborn، مورد تجزیه وتحلیل قرار گرفت و مدل سازی موضوع انجام شد. یافته ها: تجزیه وتحلیل مقالات بازیابی شده با استفاده از الگوریتم TF_IDF نشان داد بالاترین فراوانی وزنی مربوط به «بیماران»، «روان»، «سلامت روان»، «اطلاعات» و «مراقبت» هستند. با استفاده از الگوریتم تخصیص دیریکله پنهان، 7 خوشه موضوعی نتیجه مدل سازی موضوعی شامل «اطلاع جویی سلامت برخط و سواد سلامت دیجیتال»؛ «تأثیر سواد سلامت در تصمیم گیری»؛ «خوانایی منابع آموزش به بیمار»؛ «سواد سلامت در همه گیری کووید 19»؛ «سواد سلامت روان»؛ «سواد سلامت دهان و دندان»؛ و «ارتباطات در مراقبت های بهداشتی» شناسایی شد.تحلیل روند رشد تولیدات علمی در هر یک از موضوعات استخراج شده نشان داد که دو موضوع "تأثیر سواد سلامت بر تصمیم گیری" و " اطلاع جویی سلامت برخط و سواد سلامت دیجیتال " بیشترین رشد را در طول زمان داشتند. در مقابل، موضوع "سواد سلامت در همه گیری کووید-19" روند کاهشی را نشان داد. ازنظر درصد مقالات علمی در حوزه سواد اطلاعاتی مشخص شد که موضوع «سواد سلامت روان» با 22 درصد بالاترین و «سواد سلامت در همه گیری کووید 19» با 2 درصد کمترین درصد تولیدات علمی را به خود اختصاص داده اند. نتیجه گیری: خوشه های موضوعی استخراج شده از تولیدات علمی سواد اطلاعاتی انسجام مناسب و روابط موضوعی قوی را نشان دادند؛ بنابراین این پژوهش می تواند کمک شایانی به پژوهشگران در راستای ارتقای تولیدات علمی حوزه سواد اطلاعات سلامت کند.Analysis of Topics and Trends in Scientific Productions in the Field of Health Information Literacy Using Text Mining Techniques
Purpose : Diverse research in information literacy necessitates analyzing the topics of these studies to gain a clear and comprehensive understanding of this area. The current research aims to apply topic modeling to published scientific productions related to health information literacy using the PubMed database. Method: This study employed a quantitative approach with an applied focus, utilizing text-mining techniques. Scientific publications in information literacy were extracted from the PubMed database using the MeSH term "information literacy" [Majr] without any time constraints. A search on August 5, 2024, yielded 8407 records from 1519 journals and books. Subsequently, the abstracts and titles of the articles were saved in text format and then converted into a structured Excel format for analysis. After removing null records, 7608 records with abstracts were used for analysis. The process involved tokenization, removal of punctuation and stop words, stemming, and conversion of text data into numerical vectors to apply machine learning techniques. Finally, topic modeling was performed using the Latent Dirichlet Allocation (LDA) algorithm. After data cleaning, the abstracts and titles of these articles were analyzed and topic modeled using the Pandas, PyLDAvis, sklearn, PyLDAvis, numpy, Setuptools, NLTK, Gensim, Wordcloud, and Seaborn libraries. Findings: Analysis of the retrieved articles using the TF-IDF algorithm revealed that the terms "patients", "mental", "mental health", "information", and "care" had the highest term frequency-inverse document frequency weights. Using Latent Dirichlet Allocation, seven thematic clusters were identified, including "Online Health Information Seeking and Digital Health Literacy"; "Impact of Health Literacy on Decision-Making"; "Readability of Patient Education Materials"; "Health Literacy During the COVID-19 Pandemic";"Mental Health Literacy"; "Oral Health Literacy"; and "Communication in Healthcare."The analysis of scientific articles in the field of health information literacy revealed that the topics “Impact of Health Literacy on Decision-Making” and “Online Health Information Seeking and Digital Health Literacy” experienced the highest growth over time. In contrast, the topic “Health Literacy During the COVID-19 Pandemic” showed a declining trend. Additionally, the distribution of publications showed that “Mental Health Literacy” accounted for the largest share at 22%, while “Health Literacy During the COVID-19 Pandemic” represented the smallest share, making up only 2% of the total publications in the field of health information literacy. Conclusion : The extracted thematic clusters from the scientific productions on information literacy demonstrated good coherence and strong thematic relationships; therefore, this research can significantly contribute to researchers in improving scientific production in the field of health information literacy.







