Details

Title Исследование и разработка метода машинного перевода для пары малоресурсных языков вьетнамский – кхмерский с использованием крупной языковой модели SEA-LION: выпускная квалификационная работа магистра: направление 09.04.04 «Программная инженерия» ; образовательная программа 09.04.04_02 «Основы анализа и разработки приложений с большими объемами распределенных данных» = Research and development of a machine translation method for the low-resource Vietnamese – Khmer language pair using the SEA-LION large language model
Creators Фам Ван Ньян
Scientific adviser Малеев Олег Геннадьевич
Organization Санкт-Петербургский политехнический университет Петра Великого. Институт компьютерных наук и кибербезопасности
Imprint Санкт-Петербург, 2026
Collection Выпускные квалификационные работы ; Общая коллекция
Subjects машинный перевод ; малоресурсные языки ; вьетнамский – кхмерский ; большая языковая модель ; SEA-LION ; machine translation ; low-resource languages ; Vietnamese – Khmer ; large language model
Document type Master graduation qualification work
Language Russian
Level of education Master
Speciality code (FGOS) 09.04.04
Speciality group (FGOS) 090000 - Информатика и вычислительная техника
DOI 10.18720/SPBPU/3/2026/vr/vr26-5575
Rights Доступ по паролю из сети Интернет (чтение)
Additionally New arrival
Record key ru\spstu\vkr\44839
Record create date 9/4/2026

Allowed Actions

Action 'Read' will be available if administrator prepare required files

Group Anonymous
Network Internet

В ходе работы был проведен анализ предметной области, включающий обзор существующих подходов к машинному переводу для малоресурсных языковых пар (transfer learning и многоязычное обучение, back-translation, перевод через промежуточный язык, а также предварительно обученные многоязычные модели mBART, M2M-100 и NLLB-200) и анализ лингвистических особенностей вьетнамского и кхмерского языков. На основе изученных теоретических основ (нейронный машинный перевод, архитектура Transformer, параметрически-эффективное дообучение QLoRA, дообучение по инструкциям с completion-only loss и метрики оценки перевода) был разработан процесс (pipeline) адаптации крупной языковой модели. Был собран и предварительно обработан двуязычный корпус VOV, приведенный к формату instruction-response; модель Gemma-SEA-LION-9B была дообучена методом QLoRA на одном GPU, а также исследовано расширение данных путем конкатенации реальных переводных пар. В результате работы на независимом тестовом наборе из 1996 предложений были оценены четыре конфигурации. Базовая модель в режиме zero-shot показала слабый результат (12,68 spBLEU и 28,10 chrF++), уступая внешнему baseline NLLB-200-distilled-1.3B (19,91 и 44,98). После дообучения методом QLoRA на исходных данных качество выросло до 51,25 spBLEU и 57,60 chrF++, а добавление 10 000 аугментированных примеров обеспечило наилучший результат – 57,14 spBLEU и 69,60 chrF++, что превышает baseline NLLB примерно на 37 пунктов spBLEU и 25 пунктов chrF++. Качественный анализ ошибок показал, что дообученная модель существенно снижает повторение последовательностей, обеспечивает более связный кхмерский текст и лучше сохраняет имена собственные. Полученные результаты подтверждают, что адаптация модели Gemma-SEA-LION-9B с помощью QLoRA является осуществимым и эффективным подходом к машинному переводу для малоресурсной пары вьетнамский – кхмерский в условиях ограниченного аппаратного обеспечения.

During the work, an analysis of the subject area was carried out, including a review of existing approaches to machine translation for low-resource language pairs (transfer learning and multilingual training, back-translation, pivot-language translation, and the pre-trained multilingual models mBART, M2M-100 and NLLB-200), as well as an analysis of the linguistic features of the Vietnamese and Khmer languages. Based on the studied theoretical foundations (neural machine translation, the Transformer architecture, parameter-efficient fine-tuning with QLoRA, instruction fine-tuning with completion-only loss, and translation evaluation metrics), a processing pipeline for adapting the large language model was developed. The VOV bilingual corpus was collected and pre-processed and converted into an instruction-response format; the Gemma-SEA-LION-9B model was fine-tuned using QLoRA on a single GPU, and data augmentation by concatenating real translation pairs was investigated. As a result of the work, four configurations were evaluated on an independent test set of 1,996 sentences. The base model in the zero-shot setting performed poorly (12.68 spBLEU and 28.10 chrF++), falling behind the external baseline NLLB-200-distilled-1.3B (19.91 and 44.98). After QLoRA fine-tuning on the original data, quality rose to 51.25 spBLEU and 57.60 chrF++, while adding 10,000 augmented examples yielded the best result – 57.14 spBLEU and 69.60 chrF++, exceeding the NLLB baseline by about 37 spBLEU and 25 chrF++ points. The qualitative error analysis showed that the fine-tuned model substantially reduces sequence repetition, produces more coherent Khmer text, and better preserves named entities. The obtained results confirm that adapting the Gemma-SEA-LION-9B model with QLoRA is a feasible and effective approach to machine translation for the low-resource Vietnamese – Khmer pair under limited hardware.

Network User group Action
ILC SPbPU Local Network All
Internet Authorized users SPbPU
Internet Anonymous
...