Luận văn automatic presentation slides generator for scientific academic paper

Máy tạo slide trình bày tự động cho bài báo khoa học giúp tăng hiệu quả trình bày, tiết kiệm thời gian và nâng cao chất lượng bài thuyết trình.

Chuyên ngành

Khoa học máy tính

Người đăng

Ẩn danh

Thể loại

Đồ án tốt nghiệp

2023

85
0
0

Phí lưu trữ

30 Point

Tóm tắt

I. Giới Thiệu Về Tạo Slide Thuyết Trình Tự Động

Tạo slide thuyết trình tự động là một giải pháp hiện đại giúp tiết kiệm thời gian và công sức trong quá trình chuẩn bị bài thuyết trình từ bài báo khoa học. Trong thời đại kỹ thuật số, việc xử lý thông tin từ các tài liệu học thuật trở nên ngày càng quan trọng. Công nghệ này kết hợp xử lý ngôn ngữ tự nhiên (NLP)trí tuệ nhân tạo để trích xuất nội dung chính từ file PDF hoặc dữ liệu có cấu trúc. Từ đó, hệ thống tự động tóm tắt văn bản và xây dựng các slide trình bày chuyên nghiệp. Điều này đặc biệt hữu ích cho giảng viên, sinh viên và nhà nghiên cứu khi họ cần nhanh chóng chuẩn bị thuyết trình từ các bài báo khoa học phức tạp.

1.1. Định Nghĩa Và Tầm Quan Trọng

Tạo slide thuyết trình tự động là một ứng dụng công nghệ giúp tự động hóa quá trình thuyết trình từ tài liệu học thuật. Việc này giúp giảm thời gian chuẩn bị từ hàng giờ xuống còn vài phút. Trong môi trường giáo dục và nghiên cứu hiện đại, khả năng trích xuất thông tin hiệu quả từ bài báo khoa học là một kỹ năng thiết yếu.

1.2. Ứng Dụng Thực Tiễn

Công nghệ này được ứng dụng rộng rãi trong các lĩnh vực giáo dục, nghiên cứu khoa học, và kinh doanh. Sinh viên có thể nhanh chóng chuẩn bị bài thuyết trình tốt nghiệp, giảng viên có thể tạo bài giảng từ bài báo mới nhất, và các nhà nghiên cứu có thể chia sẻ kết quả một cách hiệu quả.

II. Công Nghệ Xử Lý Ngôn Ngữ Tự Nhiên Trong Tạo Slide

Xử lý ngôn ngữ tự nhiên (NLP) là nền tảng của hệ thống tạo slide tự động. Công nghệ này cho phép máy tính hiểu và phân tích văn bản từ bài báo khoa học một cách chính xác. Quy trình bao gồm phân tích cú pháp, nhận dạng thực thể, và tóm tắt văn bản để trích xuất những thông tin quan trọng nhất. Khi hệ thống xử lý dữ liệu từ file PDF, nó sẽ loại bỏ các thông tin dư thừa và giữ lại nội dung chính của bài báo. Học máyhọc sâu được sử dụng để cải thiện độ chính xác của việc phân tích và tóm tắt. Các mô hình này được huấn luyện trên tập dữ liệu lớn để nhận diện mẫu ngôn ngữ khoa học.

2.1. Các Bước Xử Lý Văn Bản

Quá trình xử lý văn bản bao gồm: trích xuất dữ liệu từ PDF, chuẩn hóa dữ liệu, phân đoạn câu, phân tích từloại bỏ từ dừng. Tiếp theo là nhận dạng các khái niệm chínhtóm tắt nội dung. Cuối cùng, hệ thống sẽ tổ chức thông tin thành các slide trình bày hợp lý.

2.2. Ứng Dụng Học Máy Và Học Sâu

Học máy giúp hệ thống học từ các bài báo trước đó để cải thiện chất lượng tóm tắt tự động. Mạng thần kinh sâu có thể phân tích các mối quan hệ phức tạp giữa các câu và đoạn văn, từ đó tạo slide có nội dung liên kết chặt chẽ hơn.

III. Các Bước Phát Triển Hệ Thống Tạo Slide Tự Động

Phát triển một hệ thống tạo slide thuyết trình tự động cần tuân theo một quy trình có cấu trúc. Giai đoạn đầu tiên là tìm hiểu lý thuyết về xử lý ngôn ngữ tự nhiên và học máy. Tiếp theo, nhóm phát triển cần thu thập và phân tích dữ liệu từ các bài báo khoa học đa dạng để chuẩn hóa dữ liệu. Giai đoạn thứ ba là cài đặt và đánh giá các mô hình hiện có. Sau đó, nhóm sẽ đề xuất các giải pháp mớitối ưu hóa hiệu suất. Cuối cùng, phải xây dựng giao diện người dùngtích hợp mô hình vào ứng dụng web. Mỗi bước cần được kiểm thử kỹ lưỡng để đảm bảo chất lượng.

3.1. Giai Đoạn Tìm Hiểu Và Thu Thập Dữ Liệu

Trong giai đoạn này, nhóm nghiên cứu tìm hiểu các công trình liên quan về bài toán tóm tắt tự độngtrích xuất thông tin từ PDF. Công việc thu thập dữ liệu từ các nguồn đáng tin cậy là rất quan trọng. Chuẩn hóa dữ liệu giúp đảm bảo tính nhất quán và chất lượng cao cho các bước tiếp theo.

3.2. Giai Đoạn Phát Triển Và Tối Ưu Hóa

Giai đoạn này bao gồm hiện thực hóa mô hình đã được đề xuất và xây dựng ứng dụng web. Nhóm sẽ cải tiến hệ thống dựa trên kết quả thử nghiệm và đánh giá hiệu suất của các mô hình. Việc tối ưu hóa độ chính xác của quá trình tóm tắttạo slide là mục tiêu chính.

IV. Lợi Ích Và Triển Vọng Phát Triển Trong Tương Lai

Hệ thống tạo slide thuyết trình tự động mang lại nhiều lợi ích đáng kể cho cộng đồng học thuật. Trước hết, nó tiết kiệm thời gian đáng kể khi chuẩn bị bài thuyết trình từ bài báo khoa học. Thứ hai, hệ thống giúp chuẩn hóa cách trình bày thông tin, đảm bảo tính nhất quán và chuyên nghiệp. Thứ ba, nó cho phép tập trung vào nội dung thay vì lo lắng về thiết kế slide. Trong tương lai, công nghệ này có tiềm năng hỗ trợ đa ngôn ngữ, tích hợp hình ảnh thông minh từ bài báo, và tối ưu hóa theo từng đối tượng người nghe. Sự kết hợp với trí tuệ nhân tạogiao diện người dùng thân thiện sẽ mở ra những khả năng mới trong lĩnh vực automation công nghiệp.

4.1. Những Lợi Ích Hiện Tại

Tiết kiệm thời gian là lợi ích chính của hệ thống tạo slide tự động. Giảng viên và sinh viên không còn phải dành nhiều giờ thực hiện tóm tắt thủ công từ bài báo khoa học. Chất lượng thuyết trình được nâng cao nhờ việc trích xuất thông tin chính xác từ các tài liệu học thuật.

4.2. Hướng Phát Triển Trong Tương Lai

Trong tương lai, tạo slide tự động sẽ được cải thiện thêm với hỗ trợ đa ngôn ngữ, nhận dạng hình ảnh, và tối ưu hóa động dựa trên nhu cầu người dùng. Sự phát triển công nghệ AI sẽ giúp hệ thống hiểu sâu hơn nội dung bài báo khoa họctạo slide phức tạp hơn.

11/12/2025
Luận văn automatic presentation slides generator for scientific academic paper

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING REPORT CAPSTONE PROJECT AUTOMATIC PRESENTATION SLIDES GENERATOR FOR SCIENTIFIC ACADEMIC PAPER Major: COMPUTER SCIENCE THESIS COMMITTEE: Computer Science 07 SUPERVISOR 1: Vo Thanh Hung, MSCS SUPERVISOR 2: Nguyen Gia Huy, BCS REVIEWER: Tran Huy, BCS —o0o— STUDENT: Dinh Gia Quang - 1911900 Ho Chi Minh City, October 2023 ĐẠI HỌC QUỐC GIA TP.HCM CỘNG HÒA XÃ HỘI CHỦ NGHĨA VIỆT NAM ---------- Độc lập - Tự do - Hạnh phúc TRƯỜNG ĐẠI HỌC BÁCH KHOA KHOA:KH & KT Máy tính ____ NHIỆM VỤ LUẬN VĂN/ ĐỒ ÁN TỐT NGHIỆP BỘ MÔN:KHMT ____________ Chú ý: Sinh viên phải dán tờ này vào trang nhất của bản thuyết trình HỌ VÀ TÊN: Đinh Gia Quang __________________________ MSSV: 1911900 _______ NGÀNH: Khoa học máy tính _____________________ LỚP: MTKH09 _______________ 1. Đầu đề luận văn/ đồ án tốt nghiệp: Tự động tạo slide thuyết trình cho bài báo nghiên cứu khoa học (Automatic presentation slides generator for scientific academic paper) 2. Nhiệm vụ (yêu cầu về nội dung và số liệu ban đầu): Thuyết trình là một hoạt động phổ biến hiện nay để tổng quát và cung cấp thông tin cho người nghe. Để chuẩn bị tốt, người thuyết trình cần dùng nhiều thời gian đọc các tài liệu từ các nguồn đáng tin cậy, trong đó có các bài báo khoa học.

Đề tài này nghiên cứu các giải pháp xử lý ngôn ngữ tự nhiên, trích xuất văn bản từ dữ liệu file pdf của bài báo, từ đó tóm tắt văn bản từ dữ liệu text hoặc dữ liệu có cấu trúc với các điều kiện đảm bảo trích được nội dung chính của bài. Từ kết quả trích xuất tự động xây dựng các slide thuyết trình cơ bản. Giai đoạn đề cương (*Giai đoạn 1*): - Tìm hiểu về xử lý ngôn ngữ tự nhiên, học máy, học sâu - Tìm hiểu các công trình liên quan về bài toán trong đề tài - Thu thập dữ liệu cho bài toán - Phân tích, đánh giá và chuẩn hóa dữ liệu - Cài đặt và đánh giá một số công trình quan trọng theo hướng liên quan - Đề xuất, thử nghiệm và đánh giá các giải pháp Giai đoạn luận văn (*Giai đoạn 2*): - Hiện thực mô hình theo giải pháp đã đề xuất - Xây dựng trang web tích hợp mô hình - Cải tiến và đánh giá hệ thống 3. Ngày giao nhiệm vụ: 05/06/2023 4.

Ngày hoàn thành nhiệm vụ: 25/09/2023 5. Họ tên giảng viên hướng dẫn: Phần hướng dẫn: 1) Võ Thanh Hùng _________________________________________________________________ 2) Nguyễn Gia Huy ________________________________________________________________ Nội dung và yêu cầu LVTN/ ĐATN đã được thông qua Bộ môn. CHỦ NHIỆM BỘ MÔN GIẢNG VIÊN HƯỚNG DẪN CHÍNH (Ký và ghi rõ họ tên) (Ký và ghi rõ họ tên) Võ Thanh Hùng PHẦN DÀNH CHO KHOA, BỘ MÔN: Người duyệt (chấm sơ bộ):________________________ Đơn vị: _______________________________________ Ngày bảo vệ: __________________________________ Điểm tổng kết: _________________________________ Nơi lưu trữ LVTN/ĐATN: ________________________ Declaration of Authenticity We declare that this research is our work, conducted under the supervision and guid- ance of Mr. Vo Thanh Hung and Mr.

Nguyen Gia Huy. The result of our research is legitimate and has not been published in any form before this. All materials used in this research are collected independently from various sources and are appropriately listed in the references section. In addition, we used the results of several other authors and organizations within this research.

They have all been aptly referenced. In any case of plagiarism, we stand by our actions and will be responsible for it. Therefore, the University of Technology - Vietnam National University HCMC is not responsible for any copyright infringements conducted within our research. Ho Chi Minh City, June 2023 Authors Dinh Gia Quang v Acknowledgements First and foremost, we would like to extend our profound gratitude to our supervisor, MSCS Vo Thanh Hung, for his tolerance, direction, and support throughout virtually all of our student years.

We have gotten a lot out of your vast expertise and careful editing. We are appreciative that you accepted us as students and kept believing in us throughout the years. Second, we acknowledge Mr. Nguyen Gia Huy’s assistance with gratitude.

We need to keep his counsel and encouragement safe since they are priceless. Families are deserving of our unending thanks for their unwavering affection, self- lessness, and assistance in keeping us inspired and self-assured. They trusted in us, and as a result, we have achieved success. We cherish our family more than words can express.

vii Abstract Presentations are commonly used in the fields of business, education, and research because they can effectively summarize and clarify large amounts of information us- ing visual aids. Crafting a high-quality presentation requires a significant investment of time and commitment and the ability to simplify complex concepts into concise and aesthetically pleasing language. By arranging layouts, providing visuals, and utilizing effects, we can enhance the impact of our presentations. The main body of a presenta- tion provides an overview of the subject matter.

Many tasks have been automated with advancements in artificial intelligence and deep learning models, saving time and labor. For example, DALL-E [1] can automatically generate images from a single sentence, while GPT-3 [2] can write paragraphs and answer questions. Naturally, we aim to create a deep-learning model that can produce presentation slides on demand. This solution in- volves document summarizing, image and text retrieval, and slide organization to ensure that important components are presented in a suitable format.

Our system is designed to help researchers efficiently create presentations on their respective topics. ix Contents Declaration of Authenticity v Acknowledgements vii Abstract ix 1 Introduction 1 1.1 Overview of scientific papers .2 Overview of slide presentation .1 Question Answering problem .2 Frequently used question forms .2 Question answering system architecture .1 Information retrieval-based system .2 Reading comprehension based question answering system .4 Sequence to Sequence .1 Recurrent Neural Network .2 Seq2seq’s architecture .2 Encoder and Decoder stacks .3 Scaled Dot-Product Attention .4 Multi-Head Attention .5 Position-wise Feed-Forward Networks .6 Embeddings and Softmax .1 Automatic slide generation .3 State-of-the-Art model utilization. 24 5 Solutions for Automatic Slide Generation 27 5.1 Query-based text summarization approach (Baseline) .2 Dense IR module .4 Figure extraction module .2 Filtering data by proposed semantic search method .2 Filtering data by using semantic search .3 Fine-tuning Dense IR module .1 Dense IR model .2 Data processing approach .1 Text extraction from papers .2 Text extraction from slides .1 State-of-the-Art model utilization .2 Fine-tune Dense IR model .3 Filtering data by using semantic search .1 Functional and non-functional requirement. 66 7 Summary 69 References 71 xiii List of Figures 3.1 Information retrieval based system with three main components [3] .2 Information retrieval based system with four main components [4] .3 Reading comprehension based question answering system [5] .4 Encoder and Decoder [6] .5 The Transformer architecture [7] .6 Scaled Dot-Product Attention [7] .7 Multi-Head Attention [7] .1 The architecture of DOC2PPT [8] .1 The system architecture of D2S model [9].2 The structure of semantic search using DistilBERT and Faiss .3 The data filtering process using semantic search .4 The difficulty in finding benchmark and data .5 The difficulty in finding pre-trained model and data .6 An example in the training dataset .7 Using CATTS model to generate summaries for data augmentation .1 Title and author of a paper in the dataset .2 Title and author are extracted using GROBID .3 An example of slide extraction by using Azure OCR .4 The flow of generating slides .7 Input slide title and keywords page .8 Generated slide page .9 Generate multi-slide page.

68 xvi List of Tables 5.1 Top 5 sentences have the highest similarity score with the given sentence.1 Descriptive statistics of documents about total count and the average number(in parenthesis) .2 Descriptive statistics of presentations about total count and the average number(in parenthesis) .3 Descriptive statistics of documents about total count and the average number(in parenthesis) .4 Descriptive statistics of presentations about total count and the average number(in parenthesis) .5 Evaluation results on the test set when using BERT and RoBERTa model 55 6.6 Evaluation results of fine-tuned models .7 Evaluation results on the test set when using fine-tuned DistilBERT model 56 6.8 The number of examples in each prefiltered and filtered dataset .9 Evaluation results of fine-tuned models using datasets filtered by using semantic search.10 Evaluation results on the test set when using DistilBERT model fine- tuned on the datasets filtered by using semantic search .11 The number of examples in each augmented data .12 Evaluation results of fine-tuned models using datasets combining filtered data and augmented data.13 Evaluation results on the test set when using DistilBERT model fine- tuned on the datasets filtered by using random forest method and aug- mented data .14 Evaluation results on the test set when using DistilBERT model fine- tuned on the datasets filtered by using semantic search (t=0.4) and aug- mented data .15 Evaluation results on the test set when using DistilBERT model fine- tuned on the datasets filtered by using semantic search (t=0.3) and aug- mented data. 62 xviii Chapter 1 Introduction 1.1 Problem Statement Presentations are commonly used in business, education, and research as they are visually effective in summarizing and explaining work to an audience. Designing a presentation is often considered an art form, requiring patience and effort to create an excellent presentation. It is essential to be able to concisely and aesthetically explain complex ideas while abstracting them.

We can enhance presentations by selecting lay- outs, adding illustrations, and incorporating effects. The content presented significantly affects the presentation’s quality as it is the summary of the subject matter. In natural language processing, text summarization is a common task. It involves breaking down lengthy text into manageable paragraphs or sentences.

There are two types of text summarization: extractive and abstractive. Extractive summarization se- lects a subset of the text’s sentences to create a summary. Abstractive summarization reorganizes the vocabulary in the source text and may add new words or phrases to cre- ate a more ”human-friendly” summary. With the effectiveness of deep learning models, researchers have integrated the selection strategy into the scoring model to predict the relative importance of previously selected sentences, resulting in better extractive sum- maries [10].

It can be challenging to generate abstract summaries for long documents. With an encoder-decoder architecture, the researchers in [11] have added deep commu- nicating agents to address this issue. A group of cooperating agents, each in charge of a 1 segment of the input text, divide the duty of encoding a long text. Each of them individ- ually encodes the given text and broadcasts their encoding to other agents, which allows agents to share global context information regarding various sections of the document.

Text summarization models are employed to summarize scientific research papers. It helps researchers save time in the literature review phase. Also, it facilitates the speedy selection of pertinent research papers and the elimination of less significant ones. As part of a search engine that allows users to select categorized values like scientific tasks, datasets, and more, IBM researchers introduced a novel system in 2019 [12] that pro- vides summaries for articles in the field of computer science.

Researchers have devel- oped the CATT model to generate TLDRs for scientific papers in 2020 [13] based on the success of deep learning models utilizing transformer architecture, such as BART and BERT. When presenting scientific research in seminars, a presentation is necessary. To make this easier, researchers have developed ways to automate the creation of slides. They use a method called extractive summarization to select important sentences or phrases in the document and use an algorithm to divide them into slides.

However, this method only focuses on the text and does not include graphic elements.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ