VIETHAAM IIATI0I1IAL UIIVEITSTV, hAII0I UHTIVET'STTV 0E EHGTIIIEETTIG AIID TEchII0L0GV HGUYEH MIHh ThUAH Enhancing the qualily 0f Machine Translati0n System Using cr0ss-Lingual W0rd Emedding M0dels (Hâng ca0 chất lơjợng của hệ thống dịch máy dựa lrên các mô hình vecl0r nhúng siêu diễn lừ giữa hai ngôn ngữ) Pr0gram: cOmputer Science MajOr: cOmputer Science c0de: 8480101.01 MASTETI' ThESIS: c0MPUƯTELI' SelIEHeE SUPELT'VIS0L: Ass0c. IGUVYEH PhU0HG ThAI han0i — 11/2018 Enhancing the quality 0f Machine TranslahOn System Using cr0ss-Lingual W0rd Emedding M0dels IIguyen Minh Thuan Faculty Of InfOrmatOn Techn0l0gy University Of Engineering and Techn0l0gy Vietnam IatiOnal University, han0i Supervised by AssOciate PrOfessOr. Iguyen PhuOng Thai A thesis submitted in fulfillment Of the requirements fOr the degree Of Master Of Science in cOmputer Science 10vemper 2018 0TTGIHALITV STATEMEITT ‘T hereby declare that this submission is my Own work amd to the best of my knowledge it contains m0 materials previously published or written by another person, or supstan- tial proportions of material which have been accepted for the award of amy other degree or diploma at University of Engineering and Technology (UET/coliech) or any other educational institution, except where due acknowledgement is made in the thesis. Any contribution made to the research by Others, with whom I have worked at UET/coltech or elsewhere, is explicitly acknowledged in the thesis.
I also declare that the intellectual content of this thesis is the product of my own work, except to the extent that assistance from others in the projects design and conception or in style, presentation and linguistic expression is acknowledged.’ han0i, WOvemser 15’", 2018 ii AbSTTAcT In zecent yeaes, Machine TeanslatiOn has shOwn pe0mising eesulls and eeœived much _intezest Of zeseaechees. TwO appeOaches that have been widely used f0z machine feans- latiOn aze Phzase-vased Statistical Machine TeanslatiOn (PBSMT) and TTeueal Ma- chine TzanslatiOn (IIMT). Duzing teanslatiOn, v0th app2Oaches 2ely feavily On lazge amOunts Of vilingual cOzp02za which zequize much eff02t and financial supp0et. The lack Of vilingual data leads 10 a p002 phezase-tavle, which is One Of the main COmp0- nents Of PBSMT, and the unknOwn w0ed p2Ovolem in IIMT.
In Onteast, mOnOlingual data aze available fO2 mOst Of the languages. Thanks 10 the advantage, many mOdels Of w0zd embedding and ce0ss-lingual w0ed embedding Fave veen appeazed 10 imp2zOve the quality Of vaziOus tasks in natueal language p2Ocessing. The puepOse Of this thesis is 10 p2zOpOse twO mOdels f02 using ceOss- lingual w0ed embedding mOdels 10 addzess the av0ve impediment. The fizst mOdel enhances the quality Of the phzase-lavle in SMT, and the zemaining mOdel tackles the unkn0wn w0ed pe0olem in TIMT.
PublicaHOns: x Minh-Thuan IIguyen, Van-Tan bui, huy-hien Vu, Phuong-Thai IIguyen and chi-Mai Luong. Enhancing the quality of Phrase-laple in Slalisical Machine Translation for Less-common and Low-Tesource Languages. In ¿he 2018 Inieenafional eonƒfeeenœ on Asian Language PeOœssing (IALP 2018). iii AckTIOWLEDGEMETITS I wOuld like 10 express my sincere gratitude 10 my lecturers in university, and especially 10 my supervisOrs - AssOc.
Iguyen PhuOng Thai, Dr. IIguyen Van Vinh and MSc. Vu huy hien. They are my inspiratiOn, guiding me 10 get the peter Of mary Osstacles in the cOmpletiOn this thesis.
I am grateful 10 my family. They usually emcOurage, mOtivate and create the pest cOnditiOns fOr me 10 accOmplish this thesis. I wOuld like 10 alsO thank my bprOlher, Ilguyen Minh ThOng, my friends, Tran Minh Luyen, hOang cOng Tuan Auh, fOr giving me many useful advices and suppOrting my thesis, my studying and my living. Finally, I sincerely acknOwledge the Vietnam [latiOnal University, han0i and especially, Tc.02-2016-03 prOject named “building a machine translaiOn system 10 suppOrt lranslaiO0n Of dOcuments between Vieliamese and Japanese 10 help managers amd businesses in han0i apprOach Japanese market’ fOr suppOrting finance 10 my master study.
T0 my family iv Table Of cOntents 1 Intr0ducti0n1 2 Lileralure review4 2.4 Open-SOurce Machine TranslatiOn .1 MOses - am Open Statistical Machine TranslaH0n h0.2 0penIIMT - an 0pen Ileural Machine TranslaH0n h0.1 M0n0lingual W0rd Empedding M0dels.2 cr0ss-Lingual W0rd Empedding MOdels .----- --- -- 13 3 Using cr0ss-Lingual W0rd Emxedding M0dels f0r MachineTrans- lali0n Systems17 3.1 Enhancing the quality Of Phrase-table in SMT Using cr0ss-Lingual WOrd 2012550000115 00107.1 TecOmpuhing Phrase-taple weights 3.2 Generaling new phrase palTS .2 Addressing the UnknOwn WOrd PrOslem in ITMT Using cr0ss-Lingual 'W0rd Empeddimg MOdels .ccseccsseseeseeeeseeseeeeeeesceseeeeseesceececseeaeeeeeeaeenes 21 4 Experiments and Fesults27 4. 31 v TABLE 0E c0IITEIITS vi 4.1 W0rd TranslaliOH Task .2 Impact Of Enriching the Phrase-table On SMT system. Impacl 0 ITem0ving the Unkn0wn W0rds 0nIIMT syslem 5_ c0nclusi0n38 List 0f Figures 2.3 The cb0W mOdel predicts the current wOrd based On the cOntext, ard the Skip-gram predicts surrOunding wOrds based On the current wOrd. 13 TOy illustratiOn Of the cr0ss-lingual emeedding m0del.
22 FlOw Of testing phrase. 25 vii List Of Tables 3.8 The sample Of new phrase pairs generated by using prOjectiOns Of wOrd vectOr represerntatiOns. 29 The precisiOn Of wOrd translatiOn retrieval t0p-k nearest neighbOrs in Vietnamese-English and Japanese- Vietnamese language paifs. 32 Tesults On UET and TED dataset in the PhSMT system fOr Vietnamese- English and Japanese-VIelramese respecliVeÌy.
--- --- «+ -s=+s+ 33 TranslaH0n examples 0f the PbSMT 1m Vieliamese-English. 34 Tesults Of remOving unkn0wn wOrds On UET and TED dataset in the TIMT syslem f0r Vielriamese-English and Japanese- Vielramese U51 177. TranslatiOn examples 0f the IIMT syslemin Vielrnamese-English vi Lisi 0f AbbreviaH0ms MT SMT PbSMT TIMT TILP THI UHMT Machine TranslaH0n Slalishcal Machine Translali0n Phrase-based Statistical MachineTranslatiOn Tleural Machine Translali0n Tlatural Language Pr0cessing Tecurrent Meural MetwOrk cll cOnvOlutiOnal Teural MetwOrk Unsupervised Ileural Machine Translali0n ix chapter 1 IntrOduch0On Machine TranslatiOn (MT) is a sub-field Of COmputatiOnal linguistics. It is autO- mated lranslah0n, which translates text Or speech frOm One natural language 10 anOther by using cCOmputer sOftware.
In the PbSMT system, the cOre Of this system is the phrase-table, which cOntains wOrds and phrases fOr SMT system 10 translate. In the translatiOn prOcess, sertences are split imt0 distinguished parts as shOwn in (KOehn et al. At each step, fOr a given sOurce phrase, the system will try 10 find the best candidate amOngst many target phrases as its translatiOn based mainly On phrase-taple. hence, having a g00d phrase-table pOssibly makes translatiOn systems imprOve the quality Of translatiOn.
hOwever, attaining a rich phrase-taple is a chal- lenge since the phrase-table is extracted and trained frOm large amOunts Of bilingual cOrpOra which require much effOrt and financial suppOrt, especially fOr less-cCOmmOn languages such as Vietnamese, La0s, etc. In the IMT system, †w0 main cÖOmp0nenis are enc0der and decOder. the encOder cOmpOnent uses a neural netwOrk, such as the recurrent neural netw0rk (TTIID, 10 encOde the sOurce sertterce, ard the decOder cOmpOnent alsO uses a neural netwOrk 10 predict wOrds in the target language. SOme IIMT mOdels incOrpOrate attentiOn mechanisms 10 imprOve the translatiOn quality.
TO reduce the cOmputahOnal cOmplexity, COnventiOnal IIMT systems Often limit their vOcabularies 10 be the t0p 30K-80K m0st frequent wOrds in the sOurce and target language, and all wOrds Outside the vOcabulary, called unkn0wn wOrds, are replaced imt0 a single unk syms01. This appr0ach leads 10 the inability 10 generate the pr0per lranslahi0n f0r this unkn0wn wOrds during lesing as sh0wn 1m (LuOng e† al.,20155) (Lï et al.,2016) Latterly, there are several apprOaches 10 address the abO0ve impediments. With the prOslem in the PbSMT system. (Passpan et al.,2016) pr0p0sed a methOd Of using new scOres generated by a COnvOlutiOn Heural TetwOrk which indicates the se- mantic relatedness Of phrase pairs.
They attained an imprOvement Of apprOximately 0. hOwever, their methOd is suitable fOr medium-size cOrpOra and creates mOre scOres fOr the phrase-table which cam increase cCOmputatiOn cOmplexity Of all translatiOn systems. (cui et al.,2013) utilized techniques Of pivOt languages 10 enrich their phrase-table. Their phrase-table is made Of sOurce-pivOt and pivOt-larget phrase-taples.
As a result Of this cCOmpinatiOn, they attained a significant imprOvement Of translatiOn. Similarly, (Zhu et al.,2014) used a methOd based On pivO0t languages 10 calculate the translatiOn prObapilities Of sOurce-target phrase pairs and achieved a slight enhance- ment. UnfOrtunately, the methOds based On pivOt languages are n0t able 10 apply fOr the Vietnamese language since the the less-cOmm0n nature Of this language. (VOgel and M0ns0n,2004) impr0ved the translatiOn quality by using phrase pairs frOm an augmented dichOnary.
They first augmented the dichOnary using simple m0rph010gical variatiOns and then assigned prOsabilities 10 entries Of this dictiOnary by using cO-Occurrence frequencies cOllected fr0m bilingual data. hOwever, their methOd needs a 10t Of bilingual cOrpOra 10 estimate accurately the prOsabilities fOr dichOnary entries, which are n0t available fOr 1|Ow-resOurce languages. In Order 10 address the unknOwn wOrd pr0slem in IIMT system. (LuOng et al., 20155) annOtated the training bilingual cOrpus with explicit alignmentinfOrmatiOn that allOws the IIMT system 10 emit, fOr each unknOwn wOrd in the target semtence, the p0siH0n Of its cOrrespOnding wOrd in the sOurce sentence.
This infOrmahOn is then used im a pOst-pr0cessing step 10 translate every unknOwn wOrd by using a bilingual diclOnary. The methOd shOwed a susstantial imprOvement Of up 10 2.8 BLEU p0ints Over vari0us IIMT systems On WMT’14 English-French translatiOn task. hOwever, having the g00d dictiOnary, which is utilized in the pOst-pr0cessing step, is alsO cOstHly and time-cOnsuming. (Semnrich et al.,2016) intrOduced a simple appr0ach 10 handle the translatiOn Of unknOwn wOrds in IMT py encOding unkn0wn wOrds as asequence Of subwOrd units.
This methOd based 0n the intuitiOn that a variety Of wOrd classes are translated via smaller units than wOrds. FOrexample, names are translated by character cOpying Or lranslderah0n, c0mpOunds are lranslaled via cOmp0siHOnal translatiOn, etc. The appr0ach indicated am imprOvement up 10 1.3 BLEU Over a sack-Off dichOnary baseline mOdel On WMT 15 English-lussian translatiOn task. (Li et al.,2016) prOpOsed a n0vel supstitutiOn-translahOn-restOratiOn methOd 10 tackle the pr0slem Of the IIMT unkn0wn w0rd.
In this methOd, the susstitutiOn step replaces the unkn0Own wOrds in a testing sentence with similar in-vOcabulary wOrds based On a similarity mOdel learned frOm mOnOlingual data. The translatiOn step then translates the testing sentence with a mOdel trained On bilingual data with unkn0wn wOrds replaced. Finally, the restOratiOn step susstitutes the translatiOns Of the replaced wOrds by that Of Original Ones. This methOd demOnstrated a significant imprOvement up 10 4 BLEU p0ints Over the attentiOn-based IIMT On chinese-t0- English translatiOn.
Tecemtly, techniques using wOrd embedding receive much interest frOm natural language pr0cessing cCOmmunities. WOrd embedding is a vectOr represertahOn Of wOrds which cOnserves semantic infOrmatiOn and their cOntexts wOrds. AdditiOnally, we can expl0it the advantage Of embedding 10 represent wOrds in diverse distinchOn spaces as shOwn in (MikOI0v et al. besides, crOss-lingual wOrd embedding mOdels are alsO receiving a 10t Of interest, which learn cr0ss-lingualrepresemtatiOns Of wOrds in a jOint embedding space 10 represent meaning and transfer knOwledge in cr0ss-lingual scerariOs.
Inspired by the advantages Of the cr0ss-lingual embedding m0dels, the wOrk Of (MikO10v et al.,20135) and (Li et al.,2016), we prOpOse a mOdel 10 enhance the quality Of a phrase-table by recOmputing the phrase weights and generating new phrase pairs fOr the phrase-taple, and a mOdel 10 address the unkn0wn w0rd pr0blem in the IIMT system by replacing the unknOwn wOrds with the mOst apprOpriate in- vOcabulary wOrds. The rest Of this thesis is Organized as fO0llOws: chapter 2 gives an Overview Of related backgrOunds. In chapter 3, we describe Our twO pr0pO0sed mOdels. A mOdel enhances the quality Of phrase-table in SMT, and the remaining mOdel tackles the unkn0wn w0rd pr0slem in IIMT.
Settings and results Of Our experiments are shOwn in chapter 4. We indicate Our cOnclusiOm and future wOrks im chapter 5. chapter 2 Literature review In this chapter, we indicate an Overview Of Machine TranslatiOn (MT) research and WOrd Embedding mOdels in sectiOn 2.1 shOws the histOry, appr0aches, evaluatiOn and Open-sOurce in MT.2, we imtrOduce an Overview Of WOrd Embedding including MOn0lingual and cr0ss-Lingual W0rd Embedding mOdels.1 hislúry Machine TranslatiOn is a sub-field Of cOmputatiOnal linguistics. It is autOmated translaliOn, which translates text Or speech frOm One natural language 10 anOther by using cCOmputer sOftware.
The first ideas Of machine translaliOn may have ap- peared in the seventh century.