Luận án tiến sĩ performance comparison of adaptive filtering in time and frequency domain for unmixing acoustic sources in real reverberant environments for close microphone applications

Luận án so sánh hiệu năng lọc thích nghi trong miền thời gian và tần số để tách nguồn âm trong môi trường phản xạ thực cho ứng dụng micro gần.

Trường đại học

Toyohashi University of Technology

Chuyên ngành

Electrical and Electronic Information Engineering

Người đăng

Ẩn danh

Thể loại

master's thesis

2013

46
3
0

Phí lưu trữ

30 Point

Tóm tắt

I. Giới thiệu

Trong lĩnh vực ghi âm âm thanh, một vấn đề quan trọng là hiện tượng rò rỉ microphone, khi âm thanh từ các nguồn không mong muốn được ghi lại. Kỹ thuật microphone gần được sử dụng để giảm thiểu hiện tượng này bằng cách đặt microphone gần nguồn âm thanh mong muốn. Mục tiêu của nghiên cứu này là so sánh hiệu suất của hai kỹ thuật lọc thích nghi trong việc tách nguồn âm trong môi trường phản xạ thực. Hai kỹ thuật này sử dụng tiêu chí sai số bình phương nhỏ nhất (MMSE) để xác định phản hồi xung kênh. Một kỹ thuật sử dụng bộ lọc Wiener trong miền tần số, trong khi kỹ thuật còn lại sử dụng thuật toán NLMS trong miền thời gian. Nghiên cứu này sẽ phân tích hiệu suất của mô hình sử dụng thuật toán NLMS trong việc tách biệt hai nguồn âm trong các môi trường phản xạ khác nhau.

1.1. Vấn đề rò rỉ microphone

Rò rỉ microphone là một vấn đề phổ biến trong ghi âm âm thanh, đặc biệt khi nhiều nhạc cụ được ghi âm cùng một lúc. Kỹ thuật microphone gần giúp giảm thiểu hiện tượng này bằng cách tối ưu hóa vị trí của microphone. Tuy nhiên, ngay cả với kỹ thuật này, âm thanh từ các nguồn không mong muốn vẫn có thể ảnh hưởng đến chất lượng ghi âm. Nghiên cứu này sẽ xem xét cách mà các kỹ thuật lọc thích nghi có thể cải thiện khả năng tách biệt âm thanh trong các tình huống thực tế.

II. Kỹ thuật lọc thích nghi

Kỹ thuật lọc thích nghi là một phương pháp quan trọng trong xử lý tín hiệu âm thanh, cho phép điều chỉnh các tham số của bộ lọc dựa trên tín hiệu đầu vào. Bộ lọc Wiener và thuật toán NLMS là hai trong số những kỹ thuật phổ biến nhất. Bộ lọc Wiener được biết đến với khả năng tối ưu hóa sai số giữa tín hiệu đầu ra và tín hiệu mong muốn, trong khi thuật toán NLMS cung cấp một cách tiếp cận linh hoạt hơn trong miền thời gian. Nghiên cứu này sẽ phân tích hiệu suất của cả hai kỹ thuật trong việc tách biệt nguồn âm trong môi trường phản xạ thực.

2.1. Bộ lọc Wiener

Bộ lọc Wiener là một giải pháp tối ưu cho vấn đề tách biệt nguồn âm. Nó hoạt động bằng cách ước lượng tín hiệu mong muốn từ tín hiệu đầu vào, bao gồm cả tín hiệu mong muốn và tín hiệu nhiễu. Bộ lọc này thường được sử dụng trong các ứng dụng như tách biệt nguồn âm và khử nhiễu. Nghiên cứu này sẽ so sánh hiệu suất của bộ lọc Wiener với thuật toán NLMS trong việc tách biệt âm thanh trong các môi trường khác nhau.

2.2. Thuật toán NLMS

Thuật toán NLMS là một phương pháp lọc thích nghi trong miền thời gian, cho phép điều chỉnh các tham số của bộ lọc dựa trên tín hiệu đầu vào. Nó thường được sử dụng trong các ứng dụng như khử nhiễu và điều chỉnh âm thanh. Nghiên cứu này sẽ phân tích cách mà thuật toán NLMS có thể cải thiện hiệu suất tách biệt âm thanh trong các môi trường phản xạ thực, so với bộ lọc Wiener.

III. Phân tích hiệu suất

Hiệu suất của các kỹ thuật lọc thích nghi được đánh giá thông qua các chỉ số như tỷ lệ tín hiệu trên nhiễu (SIR) và tỷ lệ tín hiệu trên biến dạng (SDR). Nghiên cứu cho thấy rằng bộ lọc Wiener hoạt động tốt hơn trong các điều kiện nhất định, trong khi thuật toán NLMS cho kết quả tốt hơn khi khoảng cách giữa nguồn âm và microphone tăng lên. Điều này cho thấy rằng việc lựa chọn kỹ thuật lọc thích nghi phù hợp có thể ảnh hưởng lớn đến chất lượng âm thanh cuối cùng.

3.1. Tỷ lệ tín hiệu trên nhiễu SIR

Tỷ lệ tín hiệu trên nhiễu (SIR) là một chỉ số quan trọng để đánh giá hiệu suất của các kỹ thuật lọc. Nghiên cứu cho thấy rằng bộ lọc Wiener cho hiệu suất tốt hơn khi khoảng cách giữa nguồn âm và microphone thấp. Tuy nhiên, khi khoảng cách này tăng lên, thuật toán NLMS lại cho kết quả tốt hơn. Điều này cho thấy rằng việc lựa chọn kỹ thuật lọc thích nghi cần phải dựa trên điều kiện cụ thể của môi trường ghi âm.

3.2. Tỷ lệ tín hiệu trên biến dạng SDR

Tỷ lệ tín hiệu trên biến dạng (SDR) là một chỉ số khác để đánh giá chất lượng âm thanh sau khi xử lý. Nghiên cứu cho thấy rằng thuật toán NLMS luôn cho hiệu suất tốt hơn so với bộ lọc Wiener trong việc giảm thiểu biến dạng âm thanh. Điều này cho thấy rằng thuật toán NLMS có thể là lựa chọn tốt hơn trong nhiều tình huống thực tế.

06/02/2025

Trích đoạn nội dung tài liệu

PERFORMANCE COMPARISON OF ADAPTIVE FILTERING IN TIME AND FREQUENCY DOMAIN FOR UNMIXING ACOUSTIC SOURCES IN REAL REVERBERANT ENVIRONMENTS FOR CLOSE – MICROPHONE APPLICATIONS 2013 MASTER OF ENGINEERING Department of Electrical and Electronic Information Engineering DANG NGUYEN CHAU M125212 TOYOHASHI UNIVERSITY OF TECHYNOLOGY DATE: 2013/07/25 Department of Electrical and Electronic Information ID M125212 Engineering Supervisor H. Uehara Name DANG NGUYEN CHAU Abstract PERFORMANCE COMPARISON OF ADAPTIVE FILTERING IN TIME AND Title FREQUENCY DOMAIN FOR UNMIXING ACOUSTIC SOURCES IN REAL REVERBERANT ENVIRONMENTS FOR CLOSE – MICROPHONE APPLICATIONS (800 words) One significant problem in audio recording is microphone leakage. That is when the sound of an instrument or other sources is picked by the microphone other than the desirable sources. For example, when a group of musicians is playing together, the individual microphone will not only take the signal from one instrument but also capture the interference signals that are generated by other instruments.

The close-microphone technique, in which the microphone is placed in close to the source of interest, is used in order to make the microphone capture as much of the sound of interest as possible and reduce the effect of microphone leakage. The purpose of this work is comparing the performance of two adaptive filtering techniques for solving the problem of unmixing acoustic sources. These two techniques identify channel impulse response based on minimum mean squared error (MMSE) criterion. However, one uses the calculation in frequency domain with the Wiener filter while the other calculates it in time domain by using NLMS algorithm.

Besides that, the calculation in time domain uses the “solo interval”, which is usually used in music performance. With the using of “solo interval”, the two calculations have different performance. This work aims examining the performance the model using NLMS algorithm to the realistic problem of unmixing and separating of two interfering sources in the several reverberant environments with the close-microphone technique. Moreover, the performance of the model in this work will be compared with the model using Wiener filter, which is presented as good algorithm for blind source separation (BSS) problem in unmixing acoustic sources.

In the experiment, two speakers, which are used to produce the anechoic recording of music from instruments, are used as the two sources and two omnidirectional microphones are used as the sensors. In the music performance, there are usually time interval that there is only one instrument are in active. In this period, the NLMS algorithm is used to estimate the channel response from this source to another microphone. This channel response could be called as the leakage path response.

After that period, two of sources are in active. The subtraction of the signal from this microphone with the leakage signal, which is the multiplication of the leakage path response and the other microphone signal, could be seen as the estimated signal of interest. The experiment is produced in different rooms, which have different time reverberant, for examining the system performance in various real environments. The length of the weighting vector, which is used in this work as the approximation of the leakage path response, is changed too.

The result of this changing is used to examine the performance of system when changing the length of weighting vector. For the performance evaluation, by using the orthogonal projection, the output signal of the system could be decomposed into three components: a version of the original source signal, the error term that depends on the interference signal and the other error term that depends on the other noise. Two parameters are used to show the performance of the algorithm is signal-to-interference ratio (SIR) and signal-to-distortion ratio (SDR). The SIR is used to show the remaining interference signal in the unmixed signal while the SDR is used as the quality measure of the remaining interference and noise in the output signal.

The result of the experimental shows that in the way of SIR, the Wiener filter gives the better performance than the NLMS algorithm when the distance of source and microphone is low (10cm-25cm). When this distance increases, the NLMS algorithm give the better performance than the Wiener filter. In the other hand, in the way of SDR, the NLMS algorithm always gives the better performance than the Wiener filter. The room reverberation time also has effect on the algorithm performance.

The room with longer reverberation time will gives the better performance in two cases SIR and SDR. The weighting vector length has effect on the system performance too. However, the result of this work shows that it is not effective when choosing increasing the weighting vector length for increasing the performance. ADAPTIVE NOISE CANCELLATION .2 Recursive Least Squares (RLS) adaptive filter .3 The Steepest descent method .4 The Least Mean Squared (LMS) adaptation method.

NLMS ALGORITHM APPROACH FOR UNMIXING ACOUSTIC SOURCES .1 Signal before and after processing .3 Effect of room acoustic .4 Effect of the length of weighting vector. UEHARA Page 1 DANG NGUYEN CHAU LIST OF FIGURES Fig.1 Microphone leakage in Close-Microphone applications Fig.1 A frequency domain Wiener filter for reducing additive noise Fig.2 Wiener filter structure Fig.3 Illustration of an adaptive filter Fig.1 Block diagram of the blind source separation problem Fig4.1 Block diagram of Wiener filter Fig.2 Block diagram of Wiener filter for two sources-two microphones case Fig.1 Block diagram of system using NLMS algorithm Fig.1 Reverberant time Fig.2 Learning curve of the NLMS algorithm Fig.1 Clean signal (a), Signal at the microphone (b) and the Output signal of the system (c) Fig.2 System performance (SIR) with recording in room 3 and the 2048 – filter length Fig.3 System performance (SDR) with recording in room 3 and the 2048 – filter length Fig.4 System performance (SIR) with recording in room 1 (a) and room 2 (b) with the 2048 – filter length Fig.5 System performance (SDR) with recording in room 1 (a) and room 2 (b) with the 2048 – filter length Fig.6 System SDR performance and SIR performance of Wiener filter in various room Fig.7 System performance (SIR) of model using NLMS algorithm in various rooms with 2048 – filter length Fig.8 System performance (SDR) model using NLMS algorithm in various rooms with 2048 – filter length Fig.9 Algorithm performance (SIR) with various weighting vector length in room 1 and room 3 Fig.10 Algorithm performance (SDR) with various weighting vector length in room 1 and room Supervisor: H. UEHARA Page 2 DANG NGUYEN CHAU LIST OF TABLES Table 7.1 Experiment parameter Table 7.2 Properties of rooms in which recordings took Supervisor: H. UEHARA Page 3 DANG NGUYEN CHAU ACKNOWLEDGEMENT In the first, I want to say thank you to Prof.

Uehara, who is my supervisor. In the time I am in TUT, Prof. Uehara has been helping me everything from problems in the school to the problem in life. Uehara has given me the ideas and suggesting, which is really value, for my research.

Besides that, Prof. Uehara has been helped me to have to best condition for finishing my research. Kitayama is my tutor in the time in Japan. He is the person who has helped me, who has the first time far from country, to be acquainted with the life in Japan.

He has always joined the seminar with me, given the ideas for me… I really want to show my thankful to him. Ad-hoc group is one of group in Wireless Communication Laboratory, which is the group I belong to. I want to say thank you to everyone in group for the suggesting in the seminar. With your help, I could do my work more easy.

With the other in my Laboratory, I want to say thank you with your friendship. All of you are very friendly with the, that make me feel no strange with the new life. Finally, I want to say thank you to Vietnamese friends in TUT. You are really good with me.

Toyohashi, June 15, 2013 Dang Nguyen Chau Supervisor: H. UEHARA Page 4 DANG NGUYEN CHAU PERFORMANCE COMPARISON OF ADAPTIVE FILTERING IN TIME AND FREQUENCY DOMAIN FOR UNMIXING ACOUSTIC SOURCES IN REAL REVERBERANT ENVIRONMENTS FOR CLOSE – MICROPHONE APPLICATIONS 1. INTRODUCTION In the modern music, it often involves a number of musicians playing together inside the same room, with a number of microphones, which are set to capture the sound from their instrument (see Fig. A common technique for setting microphones in this situation is to place a dedicated microphone to reproduce each sound source.

Ideally, the microphone has to pick the only signal from the interested instrument. However, due to the effect of the various instruments, the microphones will pick not only the signal from interested source but also the signal from other instruments. This is known as the microphone leakage, which is undesired effect. The close-microphone technique is the technique in which the microphone is placed close to the interested source.

The microphone is close to the interested source to capture as much of the interested sound as possible and reduce the microphone leakage effect. This technique is also used to minimize the effect of the room acoustics on the received signal. UEHARA Page 5 DANG NGUYEN CHAU Fig.1 Microphone leakage in Close-Microphone applications In order to address this problem, sound engineers suggest some ways: using the directional microphone, optimal placements of sources and microphones… However, the problem is discussed is more general, noting that this problem and the need for source separation and interference suppression arise in various other applications. The purpose of this work is suggesting a model to solve the problem unmixing two interference sources recorded in various reverberant environments with the close-microphone technique.

Besides that, it is aimed examining the performance of this model and the model using Wiener filter in [2]. In realistic case, the sources are located in enclose spaces such as concert hall or studio room. The source signal arriving at the microphone will be largely dominated by the room impulse response. Therefore, under such condition, the system from sources to microphones is the same as a multi input – multi output system.

This system is the set of room impulse response from the sources to the microphones. However, the inversion of this system is not simple. The main reasons are given in [2] such as: the non-minimum phase property of room impulse response, the unstable inversion of the system… One more reason is the room impulse response for audio processing is considered as a lengthy filter. The inverting for the matrix with each element has ten thousand elements has large computational cost.

So, it is reasonable for suggesting a model to solve the general problem of unmixing sources. Such model is presented in [2] with a model using Wiener filter. UEHARA Page 6 DANG NGUYEN CHAU Wiener filter is an alternative for solving the problem source separation. It provides a way to estimate the interested source sˆ(n) from the signal x(n) which contains the interested signal x(n) and interfering signal x(n).

The Wiener filter is frequently used for the problem source separation. NLMS is also frequently used in sound processing. However, it is frequently used for noise removing or echo equalization. This work suggests a model using NLMS with the close-microphone set up for unmixing the audio sources.

The system performance will be compared with the model using Wiener filter. Moreover, the system is examined in various real environments for unmixing sources successfully. This work is organized as follows. Section 2, the adaptive filters are presented.

In section 3, the problem formulation is presented. After that, the related work is discussed in section 4. In section 5, the suggested model using NLMS algorithm is presented. Section 6 is used to discuss about the performance computing.

In the section 7 and 8, the experimental and results are presented. Finally, the conclusion is in section 9. ADAPTIVE NOISE CANCELLATION In telecommunication from noisy acoustic environment, it is often that the interested signal is observed with an additive noise.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Bài viết "So sánh hiệu suất lọc thích nghi trong miền thời gian và tần số để tách nguồn âm trong môi trường phản xạ thực cho ứng dụng micro gần" cung cấp cái nhìn sâu sắc về các phương pháp lọc âm thanh, đặc biệt là trong các môi trường có phản xạ phức tạp. Tác giả phân tích hiệu suất của các kỹ thuật lọc trong hai miền thời gian và tần số, từ đó đưa ra những ưu điểm và nhược điểm của từng phương pháp. Điều này không chỉ giúp người đọc hiểu rõ hơn về cách thức hoạt động của các thuật toán lọc mà còn cung cấp thông tin hữu ích cho các ứng dụng thực tiễn trong lĩnh vực xử lý tín hiệu âm thanh.

Nếu bạn muốn mở rộng kiến thức về tách nguồn âm thanh, hãy tham khảo bài viết Luận văn thạc sĩ khoa học máy tính tách nguồn âm thanh dựa trên tiếp cận học máy, nơi bạn sẽ tìm thấy các phương pháp học máy hiện đại trong việc tách nguồn âm. Ngoài ra, bài viết Luận văn thạc sĩ hcmute nhận dạng và phân loại các tín hiệu quá độ dựa vào mạng neuron kết hợp với phân tích wavelets sẽ giúp bạn hiểu rõ hơn về cách phân loại tín hiệu âm thanh. Cuối cùng, bài viết Luận án tiến sĩ nghiên cứu phương pháp hiệu chỉnh các sai lệch kênh trong adc ghép xen thời gian cung cấp cái nhìn sâu sắc về các kỹ thuật hiệu chỉnh tín hiệu, rất hữu ích cho việc cải thiện chất lượng âm thanh trong các ứng dụng thực tế.