多语言模型在乌尔都语生成中存在文化语言缺陷
Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu
朋友,这篇论文挺有意思的,它专门测试了GPT-5.1、Qwen-3-Max和DeepSeek-3.1这三个大模型在生成乌尔都语故事时的表现,发现它们在文化方面确实很弱。
这篇论文分析了GPT-5.1、Qwen-3-Max和DeepSeek-3.1三个模型在生成乌尔都语故事时的表现。研究发现这些模型在语法和语义上存在基本错误,生成的故事缺乏连贯性,存在不自然的重复,并且普遍表现出文化浅薄的问题。通过少量示例提示测试,这些文化和语境错误大多未能解决。
Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu
Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.