Quick answer
Vietnamese NLP data annotation is the labelling of Vietnamese text — for NER, sentiment, intent, or text classification — to train AI models. It requires native Vietnamese speakers because the six-tone system encoded in diacritic combinations, the absence of word-boundary spaces (compound words are multi-syllabic), Northern vs Southern dialect divergence in vocabulary and pragmatic markers, and the Teenicode convention that strips tonal diacritics in social media create structural annotation challenges that generic multilingual pipelines cannot handle. Production pipelines must use underthesea or pyvi for word segmentation, vncorenlp or PhoBERT for pre-annotation, and explicit Teenicode normalisation and dialect routing policies — or risk datasets with 30–45% silent quality errors that degrade model performance without a clear diagnostic signal.
Why Vietnamese Breaks Generic NLP Annotation Pipelines
Vietnamese is the official language of Vietnam, spoken natively by approximately 96 million people and used by a further 4–5 million in diaspora communities in the United States, Australia, France, and Germany. Vietnam's digital economy, valued at $30 billion in 2023 and growing at 19% annually (Google-Temasek-Bain e-Conomy SEA report, 2023), is producing strong domestic demand for Vietnamese-language AI across e-commerce, fintech, healthcare, and government services.
Despite being written in a Latin-based script — Chữ Quốc Ngữ, introduced by French missionaries in the 17th century and standardised in the 20th — Vietnamese has linguistic properties that make it one of the most technically demanding Southeast Asian languages to annotate correctly. The features that appear familiar (Latin characters, no morphological inflection) mask properties that break every assumption generic multilingual annotation pipelines make about how text encodes meaning.
The consequences of annotation pipelines that do not account for these properties are significant. VinAI Research and other Vietnamese AI labs have documented consistent failure modes in models trained on annotation not calibrated for Vietnamese-specific properties: intent classification accuracy dropping 20–35% on informal register, NER F1 declining 15–25 percentage points on compound organisational names, and sentiment polarity errors of 30–45% on social-media corpora where Teenicode has stripped tonal diacritics. Understanding the structural causes is the prerequisite for building annotation pipelines that avoid them.
The Four Structural Challenges of Vietnamese NLP Annotation
1. Six tones, diacritic ambiguity, and the Teenicode problem
Vietnamese is a tonal language with six lexical tones, each represented by a specific diacritic mark (or absence of mark) on the vowel of each syllable. The six tones are: ngang (level, no mark — ma, ghost), huyền (falling, grave accent — mà, but), sắc (rising, acute accent — má, mother), hỏi (dipping, hook above — mả, tomb), ngã (creaky rising, tilde — mã, horse), nặng (heavy falling, dot below — mạ, rice seedling). The same Latin character string with different tone diacritics produces words with completely different meanings and grammatical roles.
This diacritic dependency creates a fundamental vulnerability when tone marks are absent or corrupted. Teenicode — the Vietnamese internet language in which diacritics are stripped for speed — systematically removes all tone information from text. 'không' (not/no) becomes 'ko' or 'k'; 'được' (can/be able to) becomes 'dc' or 'dk'; 'người' (person) becomes 'ng'. Research on Vietnamese social media corpora (Nguyen et al., VinAI Research, 2022) estimates that 40–60% of Vietnamese social media text uses Teenicode conventions at least partially. Annotation pipelines calibrated on formal text with full diacritics encounter Teenicode as near-random noise — stripped-tone forms are treated as unknown tokens — producing error rates 35–50% higher than native-speaker annotation on digital commercial corpora.
A compounding factor is mobile input method rendering inconsistency. Vietnamese input methods on Android and iOS can produce different diacritic orderings on compound vowels (the circumflex-plus-tone combinations on ô, â, ê). The grapheme 'ấ' (a with circumflex and acute tone) can be encoded as a single Unicode precomposed character (U+1EA5) or as a decomposed sequence (a + circumflex combining + acute combining). Standard Unicode normalisation (NFC/NFD) handles precomposed/decomposed differences, but annotation tools that do not apply normalisation before display will show annotators visually identical text that is tokenised differently at the character level — producing span offset mismatches that invalidate annotation provenance.
2. Word segmentation: the invisible boundary problem
Vietnamese is written with spaces between every syllabic unit — but in the linguistic sense, a 'word' in Vietnamese is often composed of two or more syllabic units that function as a single semantic unit. 'Học sinh' (student) is two syllables with a space between them, but it is one word — neither 'học' (to study) nor 'sinh' (birth/life) independently carries the meaning 'student.' 'Máy tính' (computer) is two syllables for one word. 'Ngân hàng' (bank) is two syllables for one word. The full legal name of Vietcombank — 'Ngân hàng Thương mại Cổ phần Ngoại thương Việt Nam' — spans nine space-separated syllabic tokens for a single organisation entity.
For NER annotation, this creates span boundary ambiguity that is absent from Indo-European languages and from other Southeast Asian isolating languages like Thai (which at least marks some word boundaries consistently). An annotation guideline that specifies 'label the organisation name' without specifying how to handle compound syllabic names will produce inconsistent spans across annotators: some will label the full compound name, others will label individual syllabic tokens that happen to match known organisation names. Vietnamese NLP benchmarks (Pham et al., 2017) show that word segmentation accuracy is the single largest determinant of downstream NER F1 — a 1 percentage point drop in segmentation accuracy typically causes a 2–4 percentage point drop in NER F1.
Pre-processing with a word segmentation tool before human annotation is not optional for Vietnamese NLP — it is foundational. The standard tools are underthesea (the most widely used Vietnamese NLP toolkit, providing word segmentation, POS tagging, and NER bootstrapping), pyvi (faster word segmentation based on the VnCoreNLP model), and VnCoreNLP itself (Java-based, highest F1 on Vietnamese NER benchmarks). For annotation pipelines, the recommended approach is to run word segmentation as a pre-processing step and display segmented compound words as single units to annotators — either with underscore joining ('học_sinh,' 'ngân_hàng') or with the compound word highlighted as a pre-confirmed unit — so annotators label at the word level rather than the syllabic-token level.
3. Northern vs Southern dialect divergence
Vietnamese has two major dialect groups — Northern (centred on Hanoi) and Southern (centred on Ho Chi Minh City) — with a Central dialect (Huế) occupying an intermediate position. For AI annotation purposes, Northern and Southern Vietnamese are the two dialects that drive annotation quality problems in production datasets.
The tonal distinction is the most critical for text-based annotation. Southern Vietnamese neutralises the hỏi and ngã tones — both are pronounced identically in the South, though they are still written distinctly in standard orthography. This means Southern writers sometimes use the wrong tone diacritic for these two tones, creating written forms that a Northern-calibrated annotator would classify as orthographic errors rather than regional variation. For sentiment annotation that depends on recognising intensified or negated forms, this tonal neutralisation creates a systematic misclassification risk.
Lexically, Northern and Southern Vietnamese use different words for a significant fraction of everyday vocabulary. The annotation-relevant divergences are concentrated in: food terminology (important for e-commerce and delivery annotation), kinship terms and address pronouns (critical for chatbot intent annotation — the pronoun system is much more complex in Vietnamese than in English, encoding social hierarchy and relationship distance), and discourse-final particles. Southern Vietnamese uses 'nghen' and 'nha' as positive social softeners ('okay?'/'alright?' with warmth); Northern Vietnamese uses 'nhé' and 'nào' for similar pragmatic functions. An annotator calibrated on Northern formal Vietnamese will assign neutral or question intent to Southern sentences that end in 'nghen' — where the intended reading is confirmation with positive social affect.
For annotation projects targeting national Vietnamese products — where the user base spans both Hanoi and Ho Chi Minh City — dialect-aware annotation design is essential: stratified sampling of Northern and Southern text, dialect-identified annotation routing, and dialect-balanced guideline examples. Annotation projects calibrated entirely on Northern Vietnamese formal text produce sentiment and intent errors of 15–25% when deployed on Southern Vietnamese digital corpora.
4. English code-switching in tech and startup contexts
Vietnamese tech and startup communication — customer service messages to Momo, Zalo Pay, or VNG products; reviews of app features on app stores; professional communication in the startup ecosystem — contains systematic English code-switching that creates annotation complexity comparable to Hinglish or Indonesian-English mixing.
Vietnamese-English code-switching typically inserts English technical nouns, brand names, and action verbs into Vietnamese grammatical structures. A typical Momo customer service message reads: 'Em đã nhấn nút "confirm" rồi nhưng transaction vẫn pending, em đã refresh app nhiều lần rồi' (I already pressed the "confirm" button but the transaction is still pending, I have refreshed the app many times). Vietnamese grammatical structure carries the intent (completed action, ongoing problem state, repeated attempts) while the English vocabulary identifies the domain and problem.
For intent annotation in tech and e-commerce contexts, annotators who classify by English semantic content alone miss the Vietnamese grammatical intent signals (the aspect marker 'đã' signals completed action; 'vẫn' signals persistent state contrary to expectation; 'nhiều lần' signals repeated attempts — together encoding 'I tried and it still doesn't work,' which is a high-priority escalation intent). Annotators who classify by Vietnamese grammar alone without domain knowledge of the English vocabulary miss the problem domain. Both approaches produce intent errors of 20–30% on tech-sector Vietnamese corpora.
Need native-speaker Vietnamese annotation for an NLP or AI project?
AI Taggers provides native Vietnamese-speaking annotators for NER, sentiment, intent, and text classification tasks — with Teenicode normalisation, underthesea word segmentation pre-processing, and Northern/Southern dialect routing included as standard.
See our multilingual annotation servicesThe Vietnamese NLP Tooling Ecosystem
Vietnamese NLP has a growing open-source tooling ecosystem anchored by VinAI Research, the Vietnam National University NLP groups, and the VnCoreNLP team at the University of Notre Dame and Vietnamese research institutions.
underthesea
underthesea is the most widely used Vietnamese NLP toolkit, providing word segmentation, POS tagging, NER, and dependency parsing. For annotation pre-processing, underthesea's word segmenter converts raw syllabic-token text into word-level units with underscore-joined compound words ('học_sinh,' 'ngân_hàng'), enabling consistent NER span labelling at the word rather than syllabic level. The toolkit also provides a sentence tokeniser and text normaliser that handles some Teenicode abbreviations. underthesea's NER model achieves approximately 85–89% F1 on formal Vietnamese news — useful for annotation bootstrapping on news-domain corpora, less reliable on social media text.
PhoBERT and VinAI models
PhoBERT (VinAI Research) is the standard Vietnamese pre-trained language model, available in base and large variants trained on 20GB of Vietnamese Wikipedia and news text. PhoBERT achieves state-of-the-art results on the VLSP Vietnamese NLP benchmarks for NER, POS tagging, and dependency parsing. For annotation bootstrapping, PhoBERT pre-annotations on formal Vietnamese reduce annotation time by 25–35% on news and Wikipedia-domain text, but the model degrades on Teenicode and social-media text — where pre-annotation suggestions require heavy correction. VinAI Research's PhoGPT and PhoNLP extensions expand PhoBERT coverage to informal registers but require more annotation data to fine-tune effectively.
VnCoreNLP
VnCoreNLP is a Java-based Vietnamese NLP pipeline that achieves the highest published F1 on the VLSP 2016 NER benchmark (approximately 88–92% on the CoNLL NER evaluation). For annotation pipelines that require maximum word segmentation and NER bootstrapping accuracy on formal Vietnamese, VnCoreNLP is the preferred tool despite its Java dependency. It integrates with Label Studio via the Label Studio ML backend — word-segmented and pre-annotated text can be served directly to annotators as a pre-annotation starting point, significantly reducing per-record annotation time for news-domain Vietnamese NER.
Vietnamese text normalisation for Teenicode
No standard Vietnamese Teenicode normalisation library exists with the breadth of coverage of PySastrawi for Indonesian, but several research implementations are available. The vn-nltk project and the Vietnamese Preprocessing Toolkit (by Pham, University of Engineering and Technology Vietnam) provide rule-based Teenicode expansion for the most common abbreviations. For production annotation pipelines, the recommended approach is a two-stage normalisation: first, apply the rule-based expander to convert known Teenicode abbreviations to their formal equivalents; second, display both original and normalised text to annotators, allowing human resolution of ambiguous stripped-tone forms that the rule expander cannot resolve automatically.
Case Study: Vietnamese Fintech — Customer Intent Classification Recovery
In mid-2025, a major Vietnamese digital payment and financial services platform (comparable in scale to Momo or ZaloPay) needed 48,000 annotated customer service messages in Vietnamese for an intent classification model to route customer queries automatically and reduce contact centre volume. Messages were drawn from in-app chat, including a mix of formal Northern Vietnamese, Southern Vietnamese informal register, and Teenicode, with approximately 25% English code-switching from the platform's urban millennial and Gen Z user base.
The initial annotation run used a multilingual crowdsourcing platform. After 12,000 records, the internal NLP team reviewed a validation sample against a 400-record gold standard annotated by senior native Vietnamese linguists and found:
- Intent accuracy of 61.8% against the gold standard — against a 90% project target
- 47% of Teenicode messages were partially or fully mislabelled — stripped-tone forms were treated as unknown tokens rather than standard Vietnamese words
- Southern Vietnamese discourse-final markers ('nghen,' 'nha') were assigned neutral or question intent in 33% of cases, misclassifying confirmation messages as enquiries
- NER spans for compound organisation and product names had boundary errors in 41% of cases — annotators were labelling at the syllabic token level rather than the compound word level
- Vietnamese aspect markers ('đã' for completed, 'đang' for ongoing, 'sẽ' for future) were not being used in intent attribution, causing incorrect tense-based intent classification on 28% of records
The team rebuilt the annotation pipeline with native Vietnamese-speaking annotators and proper pre-processing:
- underthesea word segmentation on all 48,000 records, displaying word-joined compound forms to annotators as pre-confirmed NER-eligible units
- Two-stage Teenicode normalisation: rule-based abbreviation expansion plus native-speaker review of ambiguous stripped-tone tokens
- Dialect identification to classify messages as Northern-dominant, Southern-dominant, or mixed — routing the approximately 38% Southern-dominant messages to annotators calibrated on Southern Vietnamese register
- Annotation guideline additions: Vietnamese aspect marker intent attribution table, Southern Vietnamese discourse-final particle reference, compound noun boundary rules, and English code-switching domain-identification policy
- Double annotation on 15% of records with kappa measurement per intent class
Results on the re-annotated corpus:
The native-speaker annotation cost was 3.4× higher per record. The 12,000 crowd-annotated records — particularly the Teenicode and Southern-dialect subsets — were structurally unusable and discarded. The downstream intent model, trained on native-speaker annotated data, achieved 89.1% accuracy on production holdout, enabling the platform to automatically route 76% of customer service queries with a false-routing rate of 3.8% — meeting the project's original target and releasing contact centre capacity for complex financial escalations.
Vietnamese Annotation Guidelines: What Generic Templates Miss
Vietnamese annotation guidelines adapted from English or generic multilingual templates routinely omit the language-specific instructions that prevent the most common errors. Critical inclusions are:
- Teenicode normalisation policy: Define whether annotators see original text, normalised text, or both. For intent and sentiment tasks, the recommended approach is to provide both — the original for context, the normalised form for consistent labelling. Include a reference table of the 50 most frequent Teenicode abbreviations with their formal Vietnamese equivalents and any tonal ambiguity notes.
- Word boundary policy for NER spans: Explicitly define that NER spans are to be placed at the word-segmented compound level, not the syllabic token level. Provide 20–30 illustrated examples of compound organisation, location, and product names with correct span boundaries, including full legal organisation names that span six or more syllabic tokens.
- Aspect marker intent attribution: Include a reference table of Vietnamese aspect markers and their intent implications: 'đã' (completed action — maps to reporting mode, not imperative), 'đang' (ongoing — maps to escalation intent if paired with negative evaluation), 'sẽ' (future — maps to request or expectation intent), 'vẫn' (still/persistent — maps to escalation intent). This table is the highest-value addition for customer service annotation projects.
- Southern Vietnamese dialect particle reference: Provide a reference table of Southern Vietnamese discourse-final particles ('nghen,' 'nha,' 'hen,' 'hén') with their pragmatic meanings — confirmation softeners with positive social affect — and their annotation treatment: do not affect intent classification, do not contribute negative sentiment. Northern dialect equivalents for comparison: 'nhé,' 'nào,' 'thôi.'
- English code-switching labelling: Define how to label English tokens in Vietnamese grammatical contexts. The standard for intent annotation is to label by the intent the Vietnamese grammatical structure expresses — the English vocabulary identifies the domain, the Vietnamese aspect markers and mood encode the intent type. For NER, English brand names and technical terms in Vietnamese text are treated as organisations or products by their semantic role.
The Vietnamese AI Market and Annotation Demand
Vietnam's National Digital Transformation Programme and the National AI Strategy (approved by the Prime Minister in 2021, updated in 2023) direct government investment toward Vietnamese-language AI, with priority areas including Vietnamese-language NLP for government digital services, healthcare AI for Vietnam's distributed hospital network, and agricultural AI for the Mekong Delta and Central Highlands farming regions.
Commercial annotation demand from Vietnam's domestic tech sector is growing rapidly. VNG Corporation's Zalo (with 74 million monthly active users) is building Vietnamese-language AI at scale, including conversational AI, moderation, and recommendation systems. Momo (35 million users) and ZaloPay are building Vietnamese financial NLP for fraud detection and customer service. FPT Corporation's AI division provides enterprise Vietnamese NLP for banking, insurance, and retail clients. VinAI Research publishes state-of-the-art Vietnamese NLP models and requires annotation data at research scale. International AI labs adding Vietnamese to Southeast Asian language rosters add a further layer of annotation demand.
Our multilingual annotation services include native Vietnamese-speaking annotators for NER, sentiment, intent, and text classification tasks — with Teenicode normalisation, underthesea word segmentation pre-processing, and Northern/Southern dialect routing included as standard. For teams building Vietnamese annotation alongside other Southeast Asian language requirements, our native-speaker annotation network covers Vietnamese alongside Thai, Indonesian, Tagalog, and 120+ other languages with consistent quality standards. For broader NLP annotation capabilities, our text annotation services provide end-to-end NLP annotation including NER, sentiment analysis, intent classification, and coreference resolution across Vietnamese and all major Southeast Asian languages.
Related Reading
If you are building Vietnamese NLP annotation alongside other language or multilingual requirements, these posts cover annotation challenges and strategies for related contexts:
- How Does Multilingual Annotation and Localization Work for Global AI? — Multi-language annotation workflows when Vietnamese is one of several target languages in a Southeast Asian product rollout
- How Much Does Using Native-Speaker Annotators Improve Multilingual AI? — Quantifying the quality lift from native-speaker annotation across language families, with case studies relevant to Southeast Asian languages
- Indonesian NLP Data Annotation: What Makes It Hard and How to Get It Right — Parallel structural challenges for teams working across Vietnamese and Indonesian (word boundary complexity, social-media register, code-switching)
Frequently Asked Questions
What is Vietnamese NLP data annotation?+
What is Teenicode and how common is it in Vietnamese corpora?+
How does Vietnamese word segmentation affect annotation accuracy?+
What is the difference between Northern and Southern Vietnamese for annotation?+
How much does Vietnamese NLP annotation cost per record?+
What is the Vietnamese AI market size?+
Start Your Vietnamese NLP Annotation Project
Tell us about your Vietnamese annotation requirements — Teenicode, dialect routing, or domain-specific — and we'll scope a native-speaker workflow for your dataset.
Neel Bennett
AI Annotation Specialist at AI Taggers
Neel has over 8 years of experience in AI training data and machine learning operations. He specializes in helping enterprises build high-quality datasets for computer vision and NLP applications across healthcare, automotive, and retail industries.
Connect on LinkedIn