技巧

Fireworks 联合 HUD 推出强化学习微调 cookbook,无需自建训练基础设施

精选理由

用 Fireworks 和 HUD 这套流程,你不用自己搭 RL 训练基础设施,定义一次任务就能边训边评,自己微调模型的可以看看。

Fireworks AI 与 HUD 合作发布了一份强化学习训练 cookbook。开发者在 HUD 中定义任务和评分器,然后在 Fireworks 平台上对模型进行采样和 RL 训练。任务只需定义一次,训练和评估可使用同一套环境。相关文档已发布在 docs.fireworks.ai 的 fine-tuning/training 页面。

图片来源 · Fireworks AI
原文 · Fireworks AI

Want to RL-train a model on your own task without building the infra? Use our new cookbook with @hud_evals : → Define your task + grader in HUD → Sample + train the model on Fireworks Define the task once. Train and evaluate against the same environment. docs.fireworks.ai/fine-tuning/tr… 💬 2 🔄 3 ❤️ 28 👀 5498 📊 9 ⚡