Luận Án Tiến Sĩ: Nghiên Cứu Về Giải Đáp Truy Vấn Cooperative XML (CoXML)

Luận án tiến sĩ về Cooperative XML (CoXML) tập trung vào phương pháp trả lời truy vấn hiệu quả, ứng dụng trong quản lý dữ liệu XML phân tán và tích hợp.

Trường đại học

University of California, Los Angeles

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

dissertation

2006

154
5
0

Phí lưu trữ

45 Point

Mục lục chi tiết

ACKNOWLEDGMENTS

ABSTRACT OF THE DISSERTATION

1. CHAPTER 1: Introduction

1.1. The Problem

1.2. CoXML Framework

1.3. Outline of the Dissertation

2. Foundation of XML Relaxation

2.1. Query Relaxation Types

2.2. Query Relaxation Properties

3. XML Query Relaxation Language

3.1. Query Relaxation Language Syntax

3.2. Query Relaxation Language Examples

4. XML Relaxation Index Structure - XTAH

4.1. XML Type Abstraction Hierarchy - XTAH

4.2. Assigning Internal Representatives in XTAH

4.3. XTAH-Guided Query Relaxation Process

4.4. Relaxing Queries with Similar Structures

5. XML Ranking

5.1. Weighted Term Frequency

5.2. Inverse Element Frequency

5.3. Extended Vector Space Model

5.4. Semantics-Oriented Structure Distance

6. The System Architecture and Relaxation Control

6.1. The System Architecture

6.2. RLXQuery: A Relaxation-Enabled XQuery

6.3. Local Relaxation Control Operators

6.4. Global Relaxation Control Operators

6.5. Functionalities of System Components

6.6. The Relaxation Control Flow

7. Evaluation Studies

7.1. The INEX Test Collection

7.2. Query Sets & Tasks

7.3. Experimental Studies on the Content Similarity

7.4. Evaluation Studies on Different Node Weight Configurations

7.5. Evaluation Studies on Different Term Modifier Weight Configurations

7.6. Performance Studies with all the 30 CAS Topics in INEX 03

7.7. Evaluation Studies on the CoXML Testbed

7.8. INEX 05 Query Sets

References

LIST OF FIGURES

LIST OF TABLES

Tóm tắt

I. Giới thiệu

Luận án tiến sĩ này tập trung vào việc giải quyết các vấn đề liên quan đến truy vấn hiệu quả trong Cooperative XML (CoXML). Với sự phát triển của công nghệ XML trong các kho dữ liệu khoa học, thư viện số và ứng dụng web, nhu cầu về các phương pháp tìm kiếm XML linh hoạt và hiệu quả ngày càng tăng. CoXML được đề xuất như một hệ thống giải đáp truy vấn XML hợp tác, cho phép tối ưu hóa truy vấn bằng cách thư giãn điều kiện truy vấn để tạo ra các câu trả lời gần đúng. Hệ thống truy vấn XML này không chỉ tập trung vào phân tích ngữ nghĩa XML mà còn đưa ra các phương pháp truy vấn XML hiệu quả, giúp người dùng tìm kiếm thông tin một cách chính xác hơn.

1.1 Vấn đề

Sự phức tạp và không đồng nhất của cấu trúc dữ liệu XML khiến người dùng khó nắm bắt hoàn toàn các thuộc tính cấu trúc trước khi đặt truy vấn. Điều này dẫn đến việc không có câu trả lời chính xác cho truy vấn. CoXML giải quyết vấn đề này bằng cách thư giãn truy vấn, tạo ra các câu trả lời gần đúng. Hệ thống quản lý dữ liệu XML này cung cấp các công cụ truy vấn XMLphương pháp truy vấn XML hiệu quả, giúp người dùng tìm kiếm thông tin một cách linh hoạt.

II. Nền tảng của XML Relaxation

Phần này trình bày các loại thư giãn truy vấntính chất của thư giãn truy vấn trong hệ thống thông tin XML. XML query relaxation là quá trình mở rộng phạm vi điều kiện truy vấn để tạo ra các câu trả lời gần đúng. CoXML sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hệ thống. Phân tích ngữ nghĩa XMLtối ưu hóa truy vấn là hai yếu tố quan trọng trong việc đảm bảo tính hiệu quả của hệ thống truy vấn XML.

2.1 Các loại thư giãn truy vấn

Có nhiều loại thư giãn truy vấn khác nhau được áp dụng trong CoXML, bao gồm thư giãn điều kiện nội dung và thư giãn điều kiện cấu trúc. XML query optimization là quá trình tối ưu hóa các điều kiện truy vấn để tạo ra các câu trả lời gần đúng. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

2.2 Tính chất của thư giãn truy vấn

Các tính chất của thư giãn truy vấn bao gồm tính linh hoạt, tính hệ thống và tính hiệu quả. CoXML đảm bảo rằng các câu trả lời gần đúng được tạo ra một cách chính xác và phù hợp với yêu cầu của người dùng. Phân tích ngữ nghĩa XMLtối ưu hóa truy vấn là hai yếu tố quan trọng trong việc đảm bảo tính hiệu quả của hệ thống truy vấn XML.

III. Ngôn ngữ thư giãn truy vấn XML

Phần này giới thiệu ngôn ngữ thư giãn truy vấn XML được sử dụng trong CoXML. Ngôn ngữ truy vấn XML này mở rộng các truy vấn tiêu chuẩn với các cấu trúc thư giãnđiều khiển thư giãn, cho phép người dùng chỉ định các điều kiện gần đúng và kiểm soát quá trình khớp gần đúng. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

3.1 Cú pháp ngôn ngữ thư giãn truy vấn

Cú pháp ngôn ngữ thư giãn truy vấn bao gồm các cấu trúc thư giãn và điều khiển thư giãn, cho phép người dùng chỉ định các điều kiện gần đúng và kiểm soát quá trình khớp gần đúng. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

3.2 Ví dụ về ngôn ngữ thư giãn truy vấn

Các ví dụ về ngôn ngữ thư giãn truy vấn minh họa cách sử dụng các cấu trúc thư giãn và điều khiển thư giãn trong CoXML. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

IV. Cấu trúc chỉ mục thư giãn XML XTAH

Phần này giới thiệu cấu trúc chỉ mục thư giãn XML (XTAH) được sử dụng trong CoXML. XTAH là một cấu trúc chỉ mục phân cấp đa cấp, cung cấp hướng dẫn và kiểm soát khớp gần đúng một cách hệ thống. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

4.1 Cấu trúc XTAH

Cấu trúc XTAH bao gồm các nhóm đa cấp, mỗi nhóm chứa một tập hợp các cấu trúc thư giãn tương ứng với một đặc tả thư giãn cụ thể. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

4.2 Quá trình thư giãn truy vấn với XTAH

Quá trình thư giãn truy vấn với XTAH bao gồm việc tham khảo các nhóm tương ứng để thư giãn truy vấn một cách hiệu quả. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

V. Đánh giá hiệu suất

Phần này trình bày các nghiên cứu đánh giá hiệu suất của CoXML sử dụng bộ sưu tập thử nghiệm INEX. Kết quả cho thấy các cấu trúc thư giãnđiều khiển thư giãn cho phép người dùng biểu đạt các đặc tả khớp gần đúng một cách hiệu quả, giúp hệ thống cung cấp các câu trả lời chính xác hơn. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

5.1 Bộ sưu tập thử nghiệm INEX

Bộ sưu tập thử nghiệm INEX được sử dụng để đánh giá hiệu suất của CoXML. Kết quả cho thấy các cấu trúc thư giãnđiều khiển thư giãn cho phép người dùng biểu đạt các đặc tả khớp gần đúng một cách hiệu quả, giúp hệ thống cung cấp các câu trả lời chính xác hơn. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

5.2 Kết quả đánh giá

Các kết quả đánh giá cho thấy CoXML không chỉ hệ thống hóa việc truy xuất các câu trả lời gần đúng mà còn đảm bảo tính liên quan cao hơn so với các hệ thống khác. Hệ thống truy vấn XML này sử dụng các phương pháp truy vấn XMLcông cụ truy vấn XML để thực hiện quá trình này một cách hiệu quả.

21/02/2025

Trích đoạn nội dung tài liệu

UNIVERSITY OF CALIFORNIA Los Angeles Cooperative XML (CoXML) Query Answering A dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Computer Science by Shaorong Liu 2006 UMI Number: 3240875 INFORMATION TO USERS The quality of this reproduction is dependent upon the quality of the copy submitted. Broken or indistinct print, colored or poor quality illustrations and photographs, print bleed-through, substandard margins, and improper alignment can adversely affect reproduction. In the unlikely event that the author did not send a complete manuscript and there are missing pages, these will be noted. Also, if unauthorized copyright material had to be removed, a note will indicate the deletion.

® UMI UMI Microform 3240875 Copyright 2007 by ProQuest Information and Learning Company. All rights reserved. This microform edition is protected against unauthorized copying under Title 17, United States Code. ProQuest Information and Learning Company 300 North Zeeb Road P.

Box 1346 Ann Arbor, MI 48106-1346 © Copyright by Shaorong Liu 2006 The dissertation of Shaorong Liu is approved. C2 TS Junghoo Cho Carlo Zaniolo LÒ, ce Yifgnian Wu A:— Wesley W. Chu, Committee Chair University of California, Los Angeles 2006 il To my family 1H TABLE OF CONTENTS 1 Introduction. 0000 2 gà kẻ va 1 11 The Problem.

cuc 1 12 CoXML Framework.3 Outline of the Dissertation. 6 2 Foundation of XML Relaxation.4 Query Relaxation Types .5 Query Relaxation Properties. eee 15 3 XML Query Relaxation Language. gà AT kia 19 3.2 Query Relaxation Language Syntax .3 Query Relaxation Language Examples.

ee 25 4 XML Relaxation Index Structure - XTAH. cv vn v v kg KV kg va 26 4.2 XML Type Abstraction Hierarchy- XTAH. eee ee eens 30 lv 4.4 Assigning Internal Representativesin XTAH .5 XTAH-Guided Query Relaxation Process.6 Relaxing Queries with Similar Structures. kg xa 45 XML Ranking .1 Weighted Term Frequency.2 Inverse Element FTequency.3 Extended Vector Space Model.3 Semantics-Oriented Structure Distance.

ee và va 57 5.1 The System ArchiteetUre.2 RLXQuery: A Relaxation-Enabled XQuery.2 Local Relaxation Control Operators.3 Global Relaxation Control Operators .3 Functionalities of System Components .4 The Relaxation Control Flow. ga 73 7 Evaluation Studies. eee ee ees 79 7.1 The INEX Test Collection.2 Query Sets & Tasks.2 Experimental Studies on the Content Similarity .1 Evaluation Studies on Different Node Weight Configurations 82 7.2 Evaluation Studies on Different Term Modifier Weight Con- figurations.3 Performance Studies with all the 30 CAS Topics in INEX 03 89 7.3 Evalation Studies on the CoXML Testbed .1 INEX 05 Query Sets .00 0c 2 cv Vy nt 101 References. AAaaaaaa 130 vi LIST OF FIGURES 11 The CoXML framework .1 A sample XML document.2 The tree representation of the sample XML document in Figure 2.3 Asample XML twig .4 Examples of structure relaxations for the twig in Figure2.1 A sample relaxation-enabled XML query.2 Topic 267 in the INEX 05 query set.3 Representing topic 267 in the INEX 05 query set with our query relaxation language.1 An example of re-using relaxed twigs for relaxing queries with the same tree structure .2 An example of XML relaxation index structure for the twig in Figure2.3 An example of using the upper and lower distance bounds in de- termining the search paths .4 The XTAH with some internal groups pruned based on the relax- ation controls !del($4) A gen(esis›) A !gen(esiss) A UseRT ype (node_delete, edge_generalization) in Figure3.5 A query with its twig structure similar to that in Figure2.1 An example of weighted term frequency.1 The CoXML testbed architecture .2 The CoXML testbed query relaxation control flow .3 The screen shot of a sample relaxation process: before relaxation .4 The screen shot of how a user edits the domain knowledge 73 6.ð The screen shot of a sample relaxation process: 1# relaxation.6 The screen shot of a sample relaxation process: 2” relaxation 76 6.7 The screen shot for a sample relaxation process: 3"? relaxation .1 The structure of a sample XML document in the INEX test collection 80 7.2 The structure summary of the INEX XML dataset .3 The precision/recall curves for topic 65 using the node weight con- figurations N,, N, & N, and the strict quantization function .4 The precision/recall curves for topic 65 using the node weight con- figurations N,, N, & N, and the generalized quantization function 85 7.5 The precision/recall curves for the 30 CAS topics in INEX 03 with the node weight configuration N, and the term modifer weight configuration M, 7.6 Comparisions of the number of returned results by our system vs.

the number of relevant results in the relevance assessment using the generalized quantization function .7 The precision/recall curves for the 30 CAS topics in INEX 03 with the node weight configuration Ny and term modifier weight con- figuration My using the AND-OR implmentation.8 The decay functiona®TM. ee vill LIST OF TABLES 2.1 Summary of notations related to the XML data model .2 Summary of notations related to the XML query model.3 Summary of notations related to XML query relaxation.1 Summary of notations related to XTAH internal nodes. 35 71 The meanings of the node labels used in Figure7.2 Three sets of node weight configurations used in experiments.3 The average precisions for topic 65 using the node weight configu- ration Nz, N, and Ng.4 Three sets of term modifier weight configurations used in experiments 87 7.5 The average precisions for topic 62 using the term modifier weights M,, Mp and Mz.6 The average precisions for all the 30 SCAS topics with the node weight configuration A and the term modifier weight configutation 7.7 The set of multi-branch queries in the INEX 05 CAS query set with their relevance assessment available .8 Comparisons of the performance evaluations (nxCG@10) for the results using the semantics-oriented vs. the uniform-cost distance functions.9 Comparisons of the performance evaluations (nxCG@25) for the results using semantics-oriented vs.

uniform-cost distance functions 98 1X 7.10 Comparisons of performance evaluations for the results with relax- ation controls vs. without relaxation controls (a =0.11 Comparisons of the performance evaluations for our results (a = 0. the official INEX 05 top-1 results in the VSCAS subtask 100 ACKNOWLEDGMENTS First, I would like to thank my advisor, Professor Wesley W. Chu, for being such a great advisor.

It is his profound knowledge and extreme patience that guided me through this dissertation. I am very grateful to Dr. Chu for the amount of effort he spent in helping me overcome various hurdles in research problems during the past four years. I also appreciate his extreme patience in advising me on how to write research papers and how to present research ideas.

I still recalled how Dr. Chu patiently helped me before I did my very first paper presentation at the SIGIR conference. Chu attended all four of my dry runs and gave me very constructive feedbacks in each dry run. I am fortunate to have not only a great research advisor but also a wonderful mentor in life.

During the past four years, Dr. Chu has kindly given me many invaluable advices, which alway inspire me and will be my lifelong assets. Second, I want to thank Professors Carlo Zaniolo and Junghoo Cho for helps during my Ph. I especially want to thank Professor Junghoo Cho for his advices on how to write research papers and how to do presentations.

I also want to thank Professors Junghoo Cho, Carlo Zaniolo and Yingnian Wu for participating in my Ph. committee and taking time to guide me through my dissertation. Third, I want to thank the two visiting professors, Arne and Ingeborg Solvberg, from Norwegian University of Science and Technology. They have given me many comments regarding my Ph.

research during their visit at UCLA. Professor Arne Solvberg has provided me with many constructive feedback on the various aspects of my defense slides, such as the organizations and logical connections. The research and development of CoXML has been a team effort. I would like Xi to thank our CoXML members, Tony Lee, Eric Sung, Anna Putnam, Christian Cardenas, Joseph Chen and Ruzan Shahinian, for their contributions in imple- mentation, testing and performance evaluation efforts.

I especially want to thank Ruzan, who has works with me during the past three years. Since Fall 2002, I have been actively participating in organizing dbUCLA sem- inars, which has been the most joyful extra-curriculum activity during my Ph. These seminars expose me to diverse state-of-the-art research topics and projects. I am grateful to all those students who were involved in organizing the seminars with me: Dr.

Zhenyu (Victor) Liu, Dr. Yi Xia, Alexandros Ntoulas, Ka Cheung Sia (Richard), Jianming He, Hetal Thakkar, Feng Qiu, Laura Yu Chen and many others. I am also indebted to all those volunteer seminar speakers, such as Raymond Pon, who has always been a great backup speaker. Further- more, special thanks goes to Professors Carlo Zaniolo and Junghoo Cho for their support in these activities.

I also want to thank our CoBase alumni, Dr. Wenlei Mao, Dr. Zhenyu (Victor) Liu and Dr. Wenlei gave me many suggestions during my first year working on XML relaxation.

Victor helped me a lot in both research and technical problems during the past few years. Qinghua was my first research collaborator and his hard work made the collaboration productive. I am also thankful to our current CoBase members, Jianming He and Laura Yu Chen, who are not only great colleagues but also good friends. Last but not least, I would like to thank my parents for their love, encour- agement and support over the years.

I also want to thank my fiance, Jiaxing, for his love, for always being very supportive and for sharing all the frustrating, sad, exciting and happy moments with me. His love and support greatly relieves me from all the stresses during my Ph. Summer 2005 Research Intern, Siemens Corporate Research(SCR). Fall 2005 Teaching Assistant, Computer Science Department, UCLA.

Fall 2004 Teaching Assistant, Computer Science Department, UCLA. 2002 - 2006 Research Assistant, Computer Science Department, UCLA. PUBLICATIONS Fusheng Wang, Shaorong Liu, Peiya Liu and Yijian Bai. Bridging Physical and Virtual Worlds: Complex Event Processing for RFID Data Streams.

In Proceedings of 10th International Conference on Extending Database Technology (EDBT 2006), Munich, Germany, March, 2006. Fusheng Wang, Shaorong Liu and Peiya Liu. Complex RFID Event Processing. Submitted for Journal Publication, 2006.

xii Shaorong Liu, Fusheng Wang and Peiya Liu. Integrated RFID Data Modeling: An Approach for Querying Physical Objects in Pervasive Computing. Submitted for Conference Publication, 2006. Yijian Bai, Fusheng Wang, Peiya Liu and Shaorong Liu.

RFID Data Processing with a Data Stream Query Language. Submitted for Conference Publication, 2006. Shaorong Liu and Wesley W. CoXML: A Cooperative XML Query An- swering System.

Submitted for Conference Publication, 2006. Chu and Shaorong Liu. Cooperative XML (CoXML) Query An- swering. Encyclopedia and Electronic Engineering, John Wiley & Son, Inc., 2006 Shaorong Liu, Wesley W.

Chu and Ruzan S. Vague Content and Structure (VCAS) Retrieval for Document-Centric XML Collections. In Proceed- ings of the 8th International Workshop on Web and Database (WebDB 2005), Baltimore, Maryland, USA, June, 2005. Shaorong Liu, Qinghua Zou, Wesley W.

Configurable Indexing and Rank- ing for XML Information Retrieval. In Proceedings of 27th Annual International ACM Special Interest Group on Information Retrieval (SIGIR 2004) Conference, Sheffield, UK, July, 2004. Qinghua Zou, Shaorong Liu, Wesley W. Using a Compact Tree to Index XIV and Query XML Data.

In Proceedings of Thirteenth Conference on Information and Knowledge Management (CIKM 2004), Washington D. Qinghua Zou, Shaorong Liu, Wesley W. CTree: A Compact Tree for Indexing XML Data. In Proceedings of 6th International Workshop on Web In- formation and Data Management (WIDM 2004), Washington D., USA, Novem- ber, 2004.

Shaorong Liu and Wesley W. Cooperative XML (CoXML) Query An- swering at INEX 2003. In Proceedings of the 2nd Initiative of the Evaluation of XML Retrieval (INEX 2003) Workshop, Schloss Dagstuhl, Germany, December, 2003. Xiaoyan Hong, Nam Nguyen, Shaorong Liu and Ying Teng.

Dynamic Group Support in LANMAR Routing Ad Hoc Networks.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Luận Án Tiến Sĩ: Giải Đáp Truy Vấn Cooperative XML (CoXML) Hiệu Quả là một nghiên cứu chuyên sâu về việc tối ưu hóa quá trình truy vấn dữ liệu XML trong môi trường hợp tác (Cooperative XML). Tài liệu này tập trung vào việc đề xuất các giải pháp và thuật toán nhằm nâng cao hiệu quả xử lý truy vấn, giảm thiểu thời gian và tài nguyên tính toán. Đây là nguồn tài liệu quý giá cho các nhà nghiên cứu, lập trình viên và chuyên gia công nghệ thông tin đang tìm hiểu về xử lý dữ liệu XML và các hệ thống hợp tác.

Để mở rộng kiến thức về các nghiên cứu liên quan, bạn có thể tham khảo 2 tóm tắt luận án tiến sĩ tiếng việt ncs nguyễn khắc tấn, cung cấp cái nhìn tổng quan về các phương pháp nghiên cứu hiện đại. Ngoài ra, Luận văn thạc sĩ xây dựng thuật toán trích xuất số phách trên phiếu trả lời trắc nghiệm của trường đại học phan thiết sẽ giúp bạn hiểu rõ hơn về ứng dụng thuật toán trong thực tiễn. Cuối cùng, Luận văn đề xuất các giải pháp nhằm nâng cao hiệu quả áp dụng là tài liệu hữu ích để khám phá các phương pháp cải thiện hiệu suất trong nghiên cứu và ứng dụng công nghệ.