论文

研究用 Bengali 数据集测高校立场极化,Llama 4 Maverick 零样本 F1 达 0.931

Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media

精选理由

孟加拉语这种低资源语言上,Llama 4 Maverick 不微调就拿下 0.931 的 F1,比 BanglaBERT+XGBoost 还高,做低资源语言分类的可以看看。

一篇 arXiv 论文发布了 4,060 条人工标注的孟加拉语评论数据集,标注为支持公立、支持私立或中立,Fleiss's Kappa 一致性达 0.89。论文比较了 SVM、XGBoost、BiLSTM、BanglaBERT+XGBoost 等监督模型和 Claude Sonnet 4、DeepSeek-V3.1、Llama 4 Maverick、Kimi K2 Thinking、Qwen3-235B Thinking 五个零样本 LLM。BanglaBERT+XGBoost 准确率 91.81%、macro F1 91.70%,而零样本 Llama 4 Maverick 在不做微调的情况下 macro F1 达 0.931、准确率 93.31%,超过所有监督基线。分析还显示 Pro-Private 评论强调设施与按时毕业,Pro-Public 评论强调学费低和政府工作。

原文 · arXiv: DeepSeek

Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media

Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduces a manually annotated dataset of 4,060 Bengali comments labeled as Pro-Public, Pro-Private, or Neutral. We evaluated the quality of our annotations by Fleiss's Kappa agreement that was 0.89, corresponding to a high agreement among annotators. The classical ML (SVM, Random Forest, XGBoost), BiLSTM network, hybrid BanglaBERT+XGBoost models and the state-of-the-art zero-shot LLMs (Claude Sonnet 4, DeepSeek-V3.1, Llama 4 Maverick, Kimi K2 Thinking, Qwen3-235B Thinking) models are evaluated. The accuracy of BanglaBERT+XGBoost is 91.81% and macro F1 score is 91.70%, which is higher than all the supervised baselines. The zero-shot Llama 4 Maverick Thinking achieves a macro F1 of 0.931 (overall accuracy of 93.31%) without any fine-tuning. All machine learning (ML), deep machine learning (DL) and transformer models were outperformed by the zero-shot Llama 4 Maverick model. Polarization also is evident, in some ways more clearly in the Pro-Private comments, which emphasize modern facilities and timely graduation, versus the Pro-Public comments, which emphasize affordability and government jobs. Our findings open new directions for analyzing social media polarization in low-resource languages.