论文

LLM 能否生成阿拉伯古典 Maqama 文体?首个对照评估研究发布

Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

精选理由

有人认真测了五个大模型能不能写阿拉伯古典韵散体 Maqama,连 GPT-5.4-mini 都在榜上,提示词策略的影响挺有意思。

一篇 arXiv 论文首次对 LLM 生成阿拉伯古典散文体 Maqama 做了对照评估,比较五个模型在 zero-shot、few-shot 和 rule-based 三种提示策略下的表现,并通过人工标注与 LLM-as-a-judge 框架从修辞丰富度、saj 韵文密度、结构连贯性等维度打分。结果显示 few-shot 提示最能稳定提升 saj 密度,而 GPT-4o 和 GPT-5.4-mini 在修辞与连贯性上更受益于 rule-based 提示,zero-shot 则在全部五个模型上拿到最高综合分。研究还用第二个独立 LLM 评委、配对显著性检验和非 LLM 的 saj 代理指标做了交叉验证。

原文 · arXiv cs.AI

Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical ornamentation, and episodic narrative structure, making it a challenging testbed for evaluating whether LLMs can move beyond surface fluency toward deeper literary competence. In this paper, we present the first controlled evaluation study of maqama generation with LLMs, comparing five models under zero-shot, few-shot, and rule-based prompting, and evaluating outputs through both human annotation and an LLM-as-a-judge framework across dimensions such as rhetorical richness, saj density, structural coherence, and stylistic authenticity. Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models. We further observe systematic differences between models in stylistic alignment with Arabic maqama conventions, and corroborate our findings with a second independent LLM judge, paired statistical significance testing, and non-LLM proxy measures of saj.