Furniture Assembly Benchmark 推理成绩10个月从28%涨到80%
有人专门做了个家具组装基准,测模型空间推理,10个月分数从28%飙到80%,连以前模型搞不定的活也开始拿下了。
一个名为 Furniture Assembly Benchmark 的空间推理基准被发布。该基准测试模型组装家具的空间推理能力,最高成绩在10个月内从28%提升到80%。空间推理此前一直被认为是模型的短板领域。基准数据需要人工故意错误组装家具来构建。
So many different kind of AI benchmarks are being released. This one is cool, "Furniture Assembly Benchmark"
Spatial reasoning used to be the safe example of what these models couldn't do.
top score has gone from 28% to 80% in just 10 months.
Somebody presumably had to build a lot of furniture wrong on purpose for this.