Outcome. You can evaluate English and Khmer separately for token efficiency, instruction following, factual preservation, names, numbers, and code-switching.
Multilingual capability emerges from training across languages, scripts, translations, and shared patterns. Cross-language transfer can be powerful, but quality is not uniform. Training coverage, tokenizer efficiency, post-training data, evaluation quality, domain, and script all matter. “One brain, many languages” is a metaphor, not a verified description of one shared internal representation.
A localized interface proves only that interface strings were translated. The course content, retrieval index, safety behavior, tool arguments, and generated answers may perform differently in Khmer. Translation can preserve general meaning while damaging names, amounts, dates, technical terms, politeness, or legal nuance. Code-switching adds another boundary because the tokenizer and model must handle multiple conventions in one request.
Noesis should make Khmer a first-class eval dimension. Use parallel cases written or reviewed by fluent speakers; score meaning, instruction compliance, naturalness, terminology, evidence use, and safety separately. Record token counts because unequal segmentation changes cost and usable context.
Mental model. Multilingual quality is a matrix of language × task × domain—not a single “supports Khmer” checkbox.
Evidence trail — reviewed 23 July 2026. BLOOM documents explicit multilingual training design: https://arxiv.org/abs/2211.05100. Tokenizer inequity across languages: https://arxiv.org/abs/2305.15425. Use official token counters for each current model.
Create 20 parallel EN/KM cases across summarization, extraction, safety, instruction following, local names/addresses, KHR/USD amounts, dates, and mixed Khmer–English technical text. Use fluent review. Count tokens and score semantic accuracy, naturalness, terminology, format, and refusal parity for each language.