VIETNAM NATIONAL UNIVERISTY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY NGO TRIEU LONG IDENTITY RESOLUTION OF DEVICES FROM WEB DATA Major: COMPUTER SCIENCE Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, June 2024 THIS THESIS IS COMPLETED AT HO CHI MINH UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisors: Assoc. Huynh Tuong Nguyen Assoc. Quan Thanh Tho Examiner 1: Dr. Vo Thi Ngoc Chau Examiner 2: Dr.
Nguyen Thi Thuy Loan This master’s thesis is defended at Ho Chi Minh City University of Technology (HCMUT) – VNU-HCM on 17th June 2024. Master’s Thesis Committee: 1. Le Hong Trang Chairman 2. Vo Thi Ngoc Chau Examiner 1 3.
Nguyen Thi Thuy Loan Examiner 2 4. Mai Hoang Bao An Commissioner 5. Le Thanh Van Secretary Approval of the Chairperson of the Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any). CHAIRMAN OF DEAN OF FACULTY OF THESIS COMMITTEE COMPUTER SCIENCE AND ENGINEERING VIỆT NAM NANTIONAL UNIVERSITY – HO CHI MINH CITY SOCIALIST REPUBLIC OF VIET NAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Ngo Trieu Long Student ID: 2270386 Date of birth: 09/06/1996 Place of birth: Quang Ngai Major: Computer Science Major ID: 8480101 I.
THESIS TITLE (in English): Identity resolution based on web script II. THESIS TITLE (in Vietnamese): Định danh khách hàng ẩn danh dựa trên hành vi web III. TASKS AND CONTENTS: a. Research and design a model capable of identifying users from web browsing data.
Implement, test and evaluate the model. THESIS START DAY: 15/01/2024 V. THESIS COMPLETION DAY: 20/05/2024 VI. HUYNH TUONG NGUYEN 2.
QUAN THANH THO Ho Chi Minh City, date 05/08/2024 SUPERVISOR SUPERVISOR CHAIRMAN OF PROGRAM (Full name and signature) (Full name and signature) COMMITTEE (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness i ACKNOWLEDGEMENTS To complete this thesis, I received substantial support from many sources. First and foremost, I would like to extend my sincere gratitude to my direct supervisors, Assoc. Quan Thanh Tho and Assoc. Huynh Tuong Nguyen.
They have been the principal guides, providing resources, monitoring the progress of this topic, and offering support whenever I encountered difficulties. Above all, they have inspired me with a passion for machine learning, deep learning, natural language processing, and many other areas in the field of Computer Science since my days as a student at the Polytechnic University. My heartfelt gratitude also goes to the dedicated teachers and assistants in the Department of Computer Science and Engineering at Ho Chi Minh City University of Technology. The knowledge I have gained from them is invaluable and has greatly assisted me in completing this thesis.
I also wish to express my thanks to the Center for Applied Data Science at FPT Corporation for providing me with the opportunity to delve into research and enhance my professional knowledge, as well as for supporting resources for model training which contributed to the completion of this thesis. Lastly, I want to thank my family, relatives, and friends, all of whom have shown concern, encouraged, and supported me both physically and mentally, enabling me to maintain the strength and health needed to successfully complete this thesis. With sincere gratitude, I wish health and all the best to all the professors and lecturers in the Department of Computer Science and Engineering at Ho Chi Minh City University of Technology, National University of Ho Chi Minh City." This revision ensures consistency in verb tenses and agreement, provides clearer attribution of support and inspiration, and maintains a formal but heartfelt tone appropriate for thesis acknowledgments. ii VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness ABSTRACT In the context of rapid technological diversification, individuals frequently use multiple electronic devices—such as personal computers, tablets, and smartphones—to access the Internet.
This multiplicity enhances the complexity of consumer behaviors, which now vary significantly across different platforms. Moreover, with the tightening of personal privacy regulations, user data on the Internet increasingly requires anonymization, complicating the tracking of user activities across devices. This thesis develops a cross- device matching framework to address these challenges using real-world data from the website fpt. The framework is designed to accommodate both device groups that receive sparse logs and those that receive frequent logs, encompassing two main stages: retrieval and re-ranking.
In the retrieval stage, a simple rule-based method utilizing the number of shared IP addresses is employed. The re-ranking stage, however, applies more sophisticated techniques. Initially, various methods were explored to represent device information, employing advanced NLP techniques such as TF-IDF and Doc2Vec for device embedding. Subsequently, extensive feature engineering on the input vectors was conducted, and different supervised classification models, as well as a Siamese Network, were utilized to determine the likelihood of device pairs belonging to the same user.
iii VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness TÓM TẮT LUẬN VĂN THẠC SĨ Trong bối cảnh công nghệ đang phát triển nhanh chóng, người dùng thường xuyên sử dụng nhiều thiết bị điện tử khác nhau—như máy tính cá nhân, máy tính bảng và điện thoại thông minh—để truy cập Internet. Sự đa dạng này làm tăng độ phức tạp của hành vi tiêu dùng, vốn thay đổi đáng kể trên các nền tảng khác nhau. Hơn nữa, với việc thắt chặt các quy định về bảo mật thông tin cá nhân, dữ liệu người dùng trên Internet ngày càng cần phải được ẩn danh, làm phức tạp hóa việc theo dõi các hoạt động của người dùng trên nhiều thiết bị. Luận văn này phát triển một khung làm việc để ghép nối các thiết bị chéo nhằm giải quyết các thách thức này bằng cách sử dụng dữ liệu thực tế từ trang web fpt.
Khung làm việc này được thiết kế để đáp ứng cả những nhóm thiết bị nhận được ít nhật ký và những nhóm nhận được nhật ký thường xuyên, bao gồm hai giai đoạn chính: truy xuất và xếp hạng lại. Trong giai đoạn truy xuất, một phương pháp đơn giản dựa trên quy tắc sử dụng số lượng địa chỉ IP chia sẻ được áp dụng. Tuy nhiên, giai đoạn xếp hạng lại áp dụng các kỹ thuật tinh vi hơn. Ban đầu, nhiều phương pháp khác nhau đã được khám phá để đại diện cho thông tin thiết bị, sử dụng các kỹ thuật NLP tiên tiến như TF-IDF và Doc2Vec để tạo vector cho thiết bị.
Sau đó, việc kỹ thuật tính năng rộng rãi trên các vector đầu vào đã được thực hiện, và các mô hình phân loại có giám sát khác nhau cũng như Mạng nơ-ron Siamese đã được sử dụng để xác định khả năng các cặp thiết bị thuộc về cùng một người dùng. iv VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness DECLARATION OF AUTHORSHIP I hereby declare that the thesis titled: “IDENTITY RESOLUTION OF DEVICES FROM WEB DATA” is my own research work. The documentation used in this thesis has been clearly stated in the References section. The data and results presented in this thesis are entirely truthful, and I am fully responsible for any inaccuracies and will accept any discipline set forth by the department and the university.
SUPERVISOR SUPERVISOR STUDENT (Full name and signature) (Full name and signature) (Full name and signature) Ngo Trieu Long v Contents 1 Topic Introduction 1 1.2 Overview about Cross-Device matching .4 Scope of Thesis .1 An Overview of Cross-Device matching techniques .2 Supervised Machine learning method .3 Candidate Filtering Techniques .3 Graph-Based method .1 Logistic Regression: A Key Technique .3 Learning to Rank. 20 vi Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 3.2 Optimization and Learning .1 Gradient Boosting Framework .5 Text Information Retrieval .2 Dataset and data pre-processing method .3 Proposed Model and Evaluation Metrics .1 Supervised Training Set Construction .4 Under-3-logs Framework .1 Motivation and idea .2 Feature Extraction and Engineering .3 Experimental results and discussion .5 Over-3-logs Framework .1 Motivation and Idea .2 Vectorization and Feature Extraction .3 Experimental results and discussion. 47 5 Conclusion 50 vii Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 5.2 Ongoing Issues and Challenges. 54 viii List of Figures 1.1 Overview of the Cross-Device User Linking problem.
Events are grouped into sessions in preprocessing.1 Overall Flow and Architecture of Cross-Device matching pipeline (Phan, 2017) [4].2 Candidate Filtering method of (Lin, 2021) [11].3 Resolving user identity: Transitioning from device-level to user- level view using device graph [12].1 Supervised Learning Model.2 An overview of Learning to Rank method.3 A simple Information Retrieval pipeline.1 Overview of the solution process [11].2 Data taken from fptshop.3 The distribution of log entries per CDP ID.4 Illustration of classification training set construction.5 Parsing of user agent string for feature extraction.6 Extract similarity feature for matching comparison.7 Transformation of CDP ID to vector form by the Doc2vec algorithm.8 Extract similarity feature from CDP ID vector.9 Example of siamese network. 45 ix List of Tables 4.1 Evaluation scores on CIKM Cup 2016 [11] .2 Performance of different algorithms at various AP thresholds .3 Performance of different algorithms at various AP thresholds .4 Performance comparison of different log usage strategies. 47 x Chapter 1 Topic Introduction 1.1 General Introduction The advent of the digital era has fundamentally changed the way we inter- act with the internet. As the digital landscape evolves, users are no longer limited to traditional PCs for web access; smartphones, tablets, and smartwatches have become equally prevalent.
This shift presents a unique challenge: accurately linking the multiple device activities of users to a single identity. Addressing this challenge is crucial for creating a comprehensive view of the customer jour- ney, which is instrumental in delivering personalized experiences that enhance customer satisfaction and drive business growth. Various applications of this technology, such as: • Impression Capping: Ensuring a user doesn’t encounter the same adver- tisement excessively across devices. • Conversion Uplift: Identifying effective media channels to increase user conversion into permanent subscribers.
• Audience Understanding: Gaining accurate audience insights by consoli- dating user identities across devices. 1 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Historically, these professionals have depended on cookies to track and tar- get users across the web. However, as users increasingly diversify their device usage beyond traditional web browsers, cookies are losing their effectiveness in tracking user behavior comprehensively. To overcome these limitations, there is a growing need for a new paradigm Cross-Device Matching [1], which proposes sophisticated methods for consistent and accurate user tracking that does not overly depend on cookie-based mapping.
This thesis is driven by the necessity of developing such innovative frameworks to meet the demands of modern digital analytics.2 Overview about Cross-Device matching Cross-Device matching is a methodological challenge in data science that aims to identify and associate multiple online user activities to a single individual across a range of personal devices. The input of the problem often contains: • Anonymized User Data: This includes browsing logs and clickstream data, with identifiers that are unique to each device but do not reveal the actual identity of the user. • Device Logs: These contain anonymized records of user activities across different devices, including timestamps, visited URLs (obfuscated), and HTML titles. There are two main approaches for solving cross-device matching task.