VIETNAM NATIONAL UNIVERSITY – HO CHI MINH CITY UNIVERSITY OF SOCIAL SCIENCES & HUMANITIES FACULTY OF ENGLISH LINGUISTICS & LITERATURE ERRORS IN THE USE OF FORMULAIC SEQUENCES BY ENGLISH-MAJOR STUDENTS A thesis submitted to the Faculty of English Linguistics & Literature in partial fulfilment of the Master’s degree in TESOL By TRAN UYEN PHUONG Supervised by PHO PHUONG DUNG, Ph. HO CHI MINH CITY, 2022 i ACKNOWLEDGEMENTS It has been a long journey from leaving my fulltime job to going back to school and completing this graduate thesis. Life happened, and I stumbled quite a few times along the way. But here I am.
Words are not enough to express my gratitude to my supervisor, Dr. Pho Phuong Dung, for her immense patience in helping me with this thesis. Whenever I needed her guidance, I would send her an email and rest assured that her response would have already been in my inbox by the time I woke up the next morning. And she would always remember to check up on me from time to time just to make sure I was doing alright.
How can a person be so knowledgeable, dedicated, and compassionate? To this day, it remains a mystery. What I know is that, deep down in my heart, I am grateful for Dr. Pho Phuong Dung’s overflowing support. Not only did she set me on the right path and never give up on me, but she also inspired me to always strive for excellence.
I would like to thank Dr. Nguyen Thi Nhu Ngoc for her support with my pilot study, as well as Mr. Nguyen Khoa Nam and Ms. Pham Thi Kieu Tien for allowing me access to their classes and accommodating me as much as they could.
I am thankful to all my participants whose names I cannot disclose, but I still want to acknowledge their valuable contributions. Without them, I would have been left with no data and no way to continue my research. Lastly, I would like to dedicate this thesis to my mother who always has my back. She will probably never see this because I am too shy to show her, but it feels right to say that she is my greatest role model in life, the one who gave me the courage and the strength to leave my comfort zone, to persist despite hardships, and to continue pursuing my passion for life-long learning.
ii STATEMENT OF ORIGINALITY I hereby certify that this thesis entitled “ERRORS IN THE USE OF FORMULAIC SEQUENCES BY ENGLISH-MAJOR STUDENTS” is my original work. This thesis has not been submitted for the award of any degree or diploma in any other institution. Ho Chi Minh City, September 2022. Trần Uyên Phương iii RETENTION OF USE I hereby state that I, Trần Uyên Phương, being the candidate for the degree of Master in TESOL, accept the requirements of the University relating to the retention and use of the Master’s Thesis deposited in the library.
In terms of these conditions, I agree that the original of my thesis deposited in the library should be accessible for the purpose of study and research, in accordance with the normal conditions established by the library for the care, loan, or reproduction of theses. Ho Chi Minh City, September 2022. iv TABLE OF CONTENTS ACKNOWLEDGEMENTS. ii STATEMENT OF ORIGINALITY.
iii RETENTION OF USE. iv LIST OF ABBREVIATIONS. viii LIST OF TABLES. ix LIST OF FIGURES.
xiii CHAPTER 1 INTRODUCTION. Background to the study. Rationale for the study. Aims of the study.
Significance of the study. Scope of the study. Organization of thesis chapters. 9 CHAPTER 2 LITERATURE REVIEW.
Definitions of FSs. Classification of FSs. Collocations, idioms, and phrasal verbs in the present study. Development of FS acquisition.
FSs and error analysis. Types of errors in the use of FSs. Sources of errors in the use of FSs. Current research on errors in the use of FSs.
Context of the study. The translation test. Corpus of Contemporary American English (COCA). Data collection procedure.
Data analysis procedure. Step 1 – Data collection. Step 2 – Identification of FS errors and correct output. Step 3 – Classification of FS errors.
Step 4 – Quantification of FS errors and correct output. Step 5 – Analysis of error sources. 55 CHAPTER 4 FINDINGS AND DISCUSSION. The participants’ common errors in the use of FSs.
Common FS errors by linguistic categories. Relative distribution of FS errors by FS types. The relationships between the participants’ formulaic performance and their learning factors. The participants’ learning factors.
The relationship between the participants’ FS errors and their learning factors. The relationship between the participants’ correct output and their learning factors. The relationship between the participants’ overall formulaic performance and their learning factors. 95 CHAPTER 5 CONCLUSION AND RECOMMENDATION.
Pedagogical implications and suggestions. Limitations of the study. Recommendations for further research. 130 vii LIST OF ABBREVIATIONS BA Bachelor of Arts BNC British National Corpus COCA Corpus of Contemporary American English EA Error Analysis EFL English as a Foreign Language FREQ Raw Frequency FS Formulaic Sequence MI Mutual Information SPSS Statistical Package for the Social Sciences viii LIST OF TABLES Table 2.
Different terms used to describe the phenomenon of formulaic language (Wray & Perkins, 2000, p.61) six classes of FS. Summary of Howarth’s (1998) phraseological categories. Summary of Wood’s (2019) FS classification. Summary of surface strategy taxonomy (Dulay, Burt, & Krashen, 1982; James, 1998).
An error taxonomy for FSs (Qi & Ding, 2011, p. The error taxonomy used for the present study (Adapted from Qi and Ding, 2011). Description of the initial questionnaire. Description of the questionnaire items used for the main study.
Participants’ FS errors by linguistic categories. Eta correlation ratio values and interpretations. Eta values of FS errors and learning factor variables. FS errors by levels of Learning Context.
FS errors by levels of Study Techniques. FS errors by levels of Active Use in Writing. The relationship between the participants’ FS errors and their learning factors. Eta values of correct output and learning factor variables.
Correct output by levels of Learning Context. Correct output by levels of Sources of Exposure. The relationship between the participants’ correct output and their learning factors. Correlations between participants’ formulaic performance and their learning factors.
96 x LIST OF FIGURES Figure 2. Six-step error analysis procedure (Gass & Selinker, 2008, p. Conceptual framework of the present study. COCA search interface.
Example of the minimize+consequence combination searched in COCA. Actual FREQ and MI score of the combination minimize+consequence. Examples of errors involving pronouns. Number of attempts made by the participants to use FSs.
Comparison of the participants’ FS errors and total attempts to use FSs. Distribution of the participants’ attempts to use FSs. FS attempts by the individual participants. Different ways to convey the message “write it down on a piece of paper” in Vietnamese.
Relative distribution of FS errors. Distribution of FS errors in comparison to the total attempts to use FSs. The participants’ starting age. The participants’ learning contexts.
The participants’ sources of exposure to English. The participants’ awareness of the benefits of FSs. The participants’ techniques to learn FSs. The participants’ habits of incorporating FSs in writing.
79 xii ABSTRACT Formulaic sequences are word sequences that are prefabricated, stored, and retrieved holistically from memory. Formulaic sequences are a fundamental component of communicative and social interactions, playing an essential role in language learning. Unfortunately, formulaic competence is slow to develop and learners at all levels of competency are prone to formulaic sequence errors. To investigate errors in the use of formulaic sequences made by 50 English-major students in a Vietnamese-to-English translation test, this study adapted Gass and Selinker’s (2008) error analysis model with two adjustments.
First, it considered formulaic sequence errors and correct output to get a more holistic view of the participants’ overall formulaic performance. Second, it included a correlation analysis to explore the potential relationship between the participants’ formulaic performance and major learning factors. In the 20,813-word collection of translation products, a total of 1,924 attempts to use formulaic sequences were identified. Among the 1,924 attempts, 1,389 word sequences met the criteria of formulaic sequences and were considered correct output, while 535 word sequences failed to meet the criteria.
These 535 word sequences were considered errors in the use of formulaic sequences. Analyzing these 535 formulaic sequence errors, the study found that the three most common types of formulaic sequence errors were inaccurate verb choices (37.31%), inaccurate preposition choices (18.7%), and inaccurate noun choices (11. Moreover, the participants’ overall formulaic performance has been found to correlate with their learning contexts and sources of exposure to English, their study techniques, and their active use of formulaic sequences in writing. On this basis, the study offers several pedagogical implications and suggestions to improve the teaching and learning of formulaic sequences.
Key words: formulaic sequences, formulaic performance, formulaic sequence errors xiii CHAPTER 1 INTRODUCTION This chapter introduces important information about the present study, including the background to the study, the rationale for the study, the aim of the study, research questions, significance of the study, the scope of the study, and the organization of the study. Background to the study Formulaic sequences are prefabricated patterns that are readily available in the mind of native speakers and can be retrieved effortlessly. They are considered a fundamental component of both spoken and written discourse. One of the most prominent studies into the distribution of formulaic sequences in the English language was conducted by Erman and Warren (2000).
According to their calculation, various types of formulaic sequences made up 58.6% of the spoken discourse and 52.3% of the written discourse. Foster (2001) used different criteria and procedures to analyze the spontaneous speech of native speakers and concluded that formulaic sequences constituted 32.3% of the data gathered. More recently, Wei and Li (2013) proposed an improved measure using the Mutual Information (MI) scores to extract formulaic sequences from corpora and found that the sequences of two to six words extracted made up 58.75% of their corpus. Their result is quite close to Erman and Warren’s (2000) manual analysis, despite the fact that Wei and Li (2013) used a corpus of academic language and automatic extraction.
In general, it is safe to say that formulaic sequences make up around one-third to one-half of the whole English language. Being a widespread linguistic phenomenon, formulaic sequences play an essential role in communication. It is argued that a native speaker could only process eight to ten words of novel discourse at a moment (Kuiper & Haggo, 1984; Pawley & Syder, 1983). Yet, on a daily basis, the average human processes and produces language much beyond this cognitive limitation thanks to their storage of formulaic sequences in long-term memory.
Because formulaic sequences are 1 stored as units, they are processed more efficiently compared to creatively generated sequences of words. As Gibbs Jr (2007) puts it, the ability to use formulaic sequences, in other words, formulaic competence, provides us with “mental shortcuts in both language production and comprehension” (p. A sufficient repertoire of formulaic language allows for ease of mental processing, contributing to natural and effective communication. Formulaic competence serves to fulfill not only communicative demands but also social needs (Wray & Perkins, 2000).
From a sociocultural perspective, the use of formulaic sequences reinforces an individual’s identity and membership within a certain social group or community. Each speech community has its own unique way of using language and a set of formulaic sequences.