Xây dựng hệ thống phát hiện tin giả mạo dạng văn bản bằng phương pháp học sâu
Abstract
In this paper, we present an approach to textual fake information detection using deep learning method combined with a Vietnamese-specific pre-training model (PhoBert) to recognize and detect fake information. Our proposed models are tested and evaluated on Vietnamese dataset [1]. From input data through input data processing steps, we need to apply word embedding to convert text data into vector model. Then this vector will go through the PhoBert transform model to get the feature vectors as output. This output vector is the input of an MLP (Multi-layer Perceptron) feedforward network to be used for the classification process. The softmax() function for the classification process is used. This function returns the probability on two sets of fake or real labels. In this paper, three different models such as CNN, LSTM and PhoBert are used to perform model training on the same data set, then compare the obtained results and the accuracy of each model to then select the best and most accurate model to build a fake information detection system. The test results show that the PhoBert model has achieved outstanding results with 89% accuracy in detecting fake information.
Tóm tắt
Trong bài báo này, chúng tôi trình bày cách tiếp cận đối với việc phát hiện tin tức giả mạo dạng văn bản bằng cách sử dụng phương pháp học sâu kết hợp với mô hình huấn luyện trước dành riêng cho tiếng Việt(PhoBert) để nhận dạng và phát hiện tin giả mạo. Các mô hình đề xuất của chúng tôi được huấn luyện và đánh giá trên bộ dữ liệu tiếng Việt [1]. Từ dữ liệu đầu vào qua các bước xử lý dữ liệu đầu vào, chúng tôi cần nhúng từ (word embedding) để chuyển đổi dữ liệu văn bản sang mô hình vector. Sau đó vector này sẽ đi qua mô hình biến đổi PhoBert để có được các vector đặc trưng làm đầu ra. Vector đầu ra này là đầu vào của một mạng truyền thẳng MLP(Multi-layer Perceptron) để sử dụng cho quá trình phân lớp. Chúng tôi sử dụng hàm softmax() cho quá trình phân lớp. Hàm này trả ra xác suất trên hai tập nhãn fake hoặc real. Trong bài báo này chúng tôi sử dụng ba mô hình khác nhau như CNN, LSTM và PhoBert để thực hiện huấn luyện mô hình trên cùng một tập dữ liệu, sau đó so sánh kết quả đạt được và độ chính xác của từng mô hình để từ đó lựa chọn mô hình tốt nhất, có độ chính xác nhất để xây dựng hệ thống phát hiện tin giả mạo. Kết quả thử nghiệm cho thấy mô hình PhoBert đã đạt được kết quả vượt trội với 89% độ chính xác phát hiện tin giả mạo.
Tài liệu tham khảo
[1] Thanh, Ho Quang, Ho Quang Thanh and ninh-pm-se, “thanhhocse96/vfnd- vietnamese-fake-news-bộ dữ liệus: Tập hợp các bài báo tiếng Việt và các bài post Facebook phân loại 2 nhãn Thật & Giả (228 bài)”. Zenodo, 27-Feb-2019, 2019.
[2] Ahmed H, Traore I, Saad S, Ahmed H, Traore I, Saad S (2017) Detection of online fake news using N-gram analysis and machine learning techniques. In: International conference on intelligent, secure, and dependable systems in distributed and cloud environments. Springer, Cham, pp 127, 2017.
[3] Ghanem B, Rosso P, Rangel F, Ghanem B, Rosso P, Rangel F (2018) Stance detection in fake news a combined feature representation. In: Proceedings of the first workshop on fact extraction and VERification (FEVER), pp 66–71, 2018.
[4] Ruchansky N, Seo S, Liu Y, Ruchansky N, Seo S, Liu Y (2017) Csi: A hybrid deep model for fake news detection. In: Proceedings of the 2017 ACM on conference on information and knowledge management. ACM, pp 797-806, 2017.
[5] Kaliyar RK, Goswami A, Narang P, Sinha S, Kaliyar RK, Goswami A, Narang P, Sinha S (2020) FNDNetA deep convolutional neural network for fake news detection. Cognitive Systems Research 61:32-44, 2020.
[6] QĐND, QĐND, “qdnd.vn," 22 11 2021. [Online]. Available: https://www.qdnd.vn/
phong-chong-dien-bien-hoa-binh/tin-gia-hiem-hoa-that-678175..
[7] M.-W. C. K. L. a. K. T. B. P.-t. o. Jacob Devlin, Jacob Devlin, Ming-Wei Chang,
Kenton Lee, and Kristina Toutanova, Bert: Pre-training of, 2019.
[8] Nguyen, Dat Quoc and Anh Gia-Tuan Nguyen. “PhoBERT: Pre-trained, Nguyen, Dat
Quoc and Anh Gia-Tuan Nguyen. “PhoBERT: Pre-trained.
[9] Jwa H, Oh D, Park K, Kang JM, Lim H, Jwa H, Oh D, Park K, Kang JM, Lim H (2019) exBAKE: Automatic fake news detection model based on bidirectional encoder representations from transformers (BERT). Appl Sci 9(19):4062, 2019.
[10] Karimi H, Roy P, Saba-Sadiya S, Tang J (2018) Multi-source multi-class fake news detection.In: Proceedings of the 27th international conference on computational linguistics, pp 1546-1557, 2018.