Details
| Title | Исследование и разработка метода машинного перевода для пары малоресурсных языков вьетнамский – кхмерский с использованием крупной языковой модели SEA-LION: выпускная квалификационная работа магистра: направление 09.04.04 «Программная инженерия» ; образовательная программа 09.04.04_02 «Основы анализа и разработки приложений с большими объемами распределенных данных» = Research and development of a machine translation method for the low-resource Vietnamese – Khmer language pair using the SEA-LION large language model |
|---|---|
| Creators | Фам Ван Ньян |
| Scientific adviser | Малеев Олег Геннадьевич |
| Organization | Санкт-Петербургский политехнический университет Петра Великого. Институт компьютерных наук и кибербезопасности |
| Imprint | Санкт-Петербург, 2026 |
| Collection | Выпускные квалификационные работы ; Общая коллекция |
| Subjects | машинный перевод ; малоресурсные языки ; вьетнамский – кхмерский ; большая языковая модель ; SEA-LION ; machine translation ; low-resource languages ; Vietnamese – Khmer ; large language model |
| Document type | Master graduation qualification work |
| Language | Russian |
| Level of education | Master |
| Speciality code (FGOS) | 09.04.04 |
| Speciality group (FGOS) | 090000 - Информатика и вычислительная техника |
| DOI | 10.18720/SPBPU/3/2026/vr/vr26-5575 |
| Rights | Доступ по паролю из сети Интернет (чтение) |
| Additionally | New arrival |
| Record key | ru\spstu\vkr\44839 |
| Record create date | 9/4/2026 |
Allowed Actions
–
Action 'Read' will be available if administrator prepare required files
| Group | Anonymous |
|---|---|
| Network | Internet |
В ходе работы был проведен анализ предметной области, включающий обзор существующих подходов к машинному переводу для малоресурсных языковых пар (transfer learning и многоязычное обучение, back-translation, перевод через промежуточный язык, а также предварительно обученные многоязычные модели mBART, M2M-100 и NLLB-200) и анализ лингвистических особенностей вьетнамского и кхмерского языков. На основе изученных теоретических основ (нейронный машинный перевод, архитектура Transformer, параметрически-эффективное дообучение QLoRA, дообучение по инструкциям с completion-only loss и метрики оценки перевода) был разработан процесс (pipeline) адаптации крупной языковой модели. Был собран и предварительно обработан двуязычный корпус VOV, приведенный к формату instruction-response; модель Gemma-SEA-LION-9B была дообучена методом QLoRA на одном GPU, а также исследовано расширение данных путем конкатенации реальных переводных пар. В результате работы на независимом тестовом наборе из 1996 предложений были оценены четыре конфигурации. Базовая модель в режиме zero-shot показала слабый результат (12,68 spBLEU и 28,10 chrF++), уступая внешнему baseline NLLB-200-distilled-1.3B (19,91 и 44,98). После дообучения методом QLoRA на исходных данных качество выросло до 51,25 spBLEU и 57,60 chrF++, а добавление 10 000 аугментированных примеров обеспечило наилучший результат – 57,14 spBLEU и 69,60 chrF++, что превышает baseline NLLB примерно на 37 пунктов spBLEU и 25 пунктов chrF++. Качественный анализ ошибок показал, что дообученная модель существенно снижает повторение последовательностей, обеспечивает более связный кхмерский текст и лучше сохраняет имена собственные. Полученные результаты подтверждают, что адаптация модели Gemma-SEA-LION-9B с помощью QLoRA является осуществимым и эффективным подходом к машинному переводу для малоресурсной пары вьетнамский – кхмерский в условиях ограниченного аппаратного обеспечения.
During the work, an analysis of the subject area was carried out, including a review of existing approaches to machine translation for low-resource language pairs (transfer learning and multilingual training, back-translation, pivot-language translation, and the pre-trained multilingual models mBART, M2M-100 and NLLB-200), as well as an analysis of the linguistic features of the Vietnamese and Khmer languages. Based on the studied theoretical foundations (neural machine translation, the Transformer architecture, parameter-efficient fine-tuning with QLoRA, instruction fine-tuning with completion-only loss, and translation evaluation metrics), a processing pipeline for adapting the large language model was developed. The VOV bilingual corpus was collected and pre-processed and converted into an instruction-response format; the Gemma-SEA-LION-9B model was fine-tuned using QLoRA on a single GPU, and data augmentation by concatenating real translation pairs was investigated. As a result of the work, four configurations were evaluated on an independent test set of 1,996 sentences. The base model in the zero-shot setting performed poorly (12.68 spBLEU and 28.10 chrF++), falling behind the external baseline NLLB-200-distilled-1.3B (19.91 and 44.98). After QLoRA fine-tuning on the original data, quality rose to 51.25 spBLEU and 57.60 chrF++, while adding 10,000 augmented examples yielded the best result – 57.14 spBLEU and 69.60 chrF++, exceeding the NLLB baseline by about 37 spBLEU and 25 chrF++ points. The qualitative error analysis showed that the fine-tuned model substantially reduces sequence repetition, produces more coherent Khmer text, and better preserves named entities. The obtained results confirm that adapting the Gemma-SEA-LION-9B model with QLoRA is a feasible and effective approach to machine translation for the low-resource Vietnamese – Khmer pair under limited hardware.
| Network | User group | Action |
|---|---|---|
| ILC SPbPU Local Network | All |
|
| Internet | Authorized users SPbPU |
|
| Internet | Anonymous |
|