OpenAI过去的产品线一直因为太乱饱受诟病,直到上一次的5.4-mini、5.4、5.5才很好的解决。
这次的 5.6-luna、5.6-terra、5.6-sol 其实就是对标 5.4-mini、5.4、5.5 三个模型。但是这次每个模型都带来了low、medium、high、xhigh、max、ultra 6个挡位,于是又乱了起来。
由 DeepSWE 的数据可以看到,最具性价比的挡位应该是gpt-5.6-sol medium和gpt-5.6-terra xhigh。如果是复杂任务,可以进阶到gpt-5.6-sol high或xhigh。
Note
不建议ultra模式,会启用大量子代理并行工作,速度很慢且非常浪费token。
| MODEL | PASS@1 条形图与误差范围 | PASS@1 | AVG COST | OUT TOK | STEPS |
|---|---|---|---|---|---|
| gpt-5.6-sol [max] | 73% ±3% | $8.39 | 60k | 61 | |
| gpt-5.6-sol [xhigh] | 71% ±1% | $4.70 | 41k | 44 | |
| gpt-5.6-terra [max] | 70% ±3% | $4.95 | 72k | 76 | |
| gpt-5.6-sol [high] | 69% ±1% | $3.47 | 28k | 37 | |
| gpt-5.6-luna [max] | 67% ±4% | $3.03 | 73k | 102 | |
| gpt-5.5 [xhigh] | 67% ±6% | $7.23 | 46k | 82 | |
| gpt-5.5 [high] | 64% ±3% | $5.10 | 31k | 62 | |
| gpt-5.6-sol [medium] | 61% ±2% | $1.86 | 18k | 31 | |
| gpt-5.6-terra [xhigh] | 60% ±2% | $2.13 | 40k | 43 | |
| gpt-5.6-luna [xhigh] | 57% ±2% | $1.54 | 45k | 71 | |
| gpt-5.5 [medium] | 54% ±3% | $2.75 | 20k | 46 | |
| gpt-5.6-terra [high] | 54% ±4% | $1.13 | 22k | 34 | |
| gpt-5.6-sol [low] | 45% ±2% | $1.07 | 11k | 23 | |
| gpt-5.6-luna [high] | 44% ±3% | $0.78 | 26k | 49 | |
| gpt-5.6-terra [medium] | 35% ±3% | $0.58 | 12k | 25 | |
| gpt-5.5 [low] | 27% ±2% | $1.20 | 9.4k | 28 | |
| gpt-5.6-terra [low] | 24% ±1% | $0.43 | 8.6k | 21 | |
| gpt-5.6-luna [medium] | 11% ±1% | $0.22 | 8.2k | 24 | |
| gpt-5.6-luna [low] | 2% ±1% | $0.07 | 3.1k | 12 | |
| 0%
20%
40%
60%
80%
|