DeepEdu-v1:面向越南教育的智能体教学系统
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
越南团队把开源模型改造成能跑在消费级 GPU 上的教学助手,检索调用少了 7.7 倍,准确率还从 70% 涨到 79.5%,做本地部署的可以看看。
论文提出 DeepEdu-v1,一个基于 SCALE 框架的越南教育 AI 辅导系统,应对数据主权法规如越南 Decree 53 的本地化部署需求。其长上下文推理引擎将 token 选择粒度从子块级改为簇级,检索调用次数比现有选择性注意力基线少 7.7 倍,TTFT 降低约 35%。部署配置下,DeepEdu 相比标准 vLLM 服务实现近 2 倍 TTFT 加速,复杂任务上的智能体准确率从 70.0% 提升到 79.5%,在金融推理和交互式智能体基准上增益最大。
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.