Paweł Huryn 实测各家模型查代码 bug 能力,成本与性能呈指数关系
Paweł Huryn tested models' ability to find bugs planted in code. Cost increases exponentially with p...
有人真金白银测了各模型找 bug 的实力和花费,画了条性价比分界线,买 API 前可以先看看谁在线上。
Paweł Huryn 让多个模型排查预埋在代码中的 bug,并对比各自的表现与调用成本。数据显示模型性能每提升一档,成本呈指数级上升,图中横轴采用对数刻度。他还用 ChatGPT 拟合出一条趋势线,位于线上方的模型被视为性价比更高的“便宜货”。
Paweł Huryn tested models' ability to find bugs planted in code. Cost increases exponentially with p...
Paweł Huryn tested models' ability to find bugs planted in code. Cost increases exponentially with performance (note the log scale on the x axis). I had ChatGPT superimpose a trend line. Bargain = above the line. 💬 18 🔄 2 ❤️ 23 👀 9772 📊 16 ⚡