跳到正文
Petrichor's Blog
返回

gpt-5.6系列模型选择

阅读 ...

OpenAI过去的产品线一直因为太乱饱受诟病,直到上一次的5.4-mini、5.4、5.5才很好的解决。

这次的 5.6-luna、5.6-terra、5.6-sol 其实就是对标 5.4-mini、5.4、5.5 三个模型。但是这次每个模型都带来了low、medium、high、xhigh、max、ultra 6个挡位,于是又乱了起来。

DeepSWE 的数据可以看到,最具性价比的挡位应该是gpt-5.6-sol mediumgpt-5.6-terra xhigh。如果是复杂任务,可以进阶到gpt-5.6-sol highxhigh

Note

不建议ultra模式,会启用大量子代理并行工作,速度很慢且非常浪费token。

GPT-5.5 与 GPT-5.6 系列 DeepSWE 得分与平均成本 横轴为每项任务的平均成本,从左到右递减;纵轴为 DeepSWE Pass@1 得分。图中比较 gpt-5.5、gpt-5.6-sol、terra 和 luna 在不同思考挡位的表现。GPT-5.5 没有 max 数据点。 DeepSWE score gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna gpt-5.5 80% 70% 60% 50% 40% 30% 20% 10% 0% $8.00 $6.00 $4.00 $2.00 $0.00 gpt-5.6-sol max: 73% · $ 8.39 gpt-5.6-sol xhigh: 71% · $ 4.70 gpt-5.6-sol high: 69% · $ 3.47 gpt-5.6-sol medium: 61% · $ 1.86 gpt-5.6-sol low: 45% · $ 1.07 gpt-5.6-terra max: 70% · $ 4.95 gpt-5.6-terra xhigh: 60% · $ 2.13 gpt-5.6-terra high: 54% · $ 1.13 gpt-5.6-terra medium: 35% · $ 0.58 gpt-5.6-terra low: 24% · $ 0.43 gpt-5.6-luna max: 67% · $ 3.03 gpt-5.6-luna xhigh: 57% · $ 1.54 gpt-5.6-luna high: 44% · $ 0.78 gpt-5.6-luna medium: 11% · $ 0.22 gpt-5.6-luna low: 2% · $ 0.07 gpt-5.5 xhigh: 67% · $ 7.23 gpt-5.5 high: 64% · $ 5.10 gpt-5.5 medium: 54% · $ 2.75 gpt-5.5 low: 27% · $ 1.20 Avg cost per task
GPT-5.5 与 GPT-5.6 系列的 DeepSWE 得分与平均任务成本
DeepSWE 模型排名。比较各模型和思考挡位的 Pass@1、误差、平均成本、输出 token 与平均步骤数。
MODEL PASS@1 条形图与误差范围 PASS@1 AVG COST OUT TOK STEPS
gpt-5.6-sol [max] 73% ±3% $8.39 60k 61
gpt-5.6-sol [xhigh] 71% ±1% $4.70 41k 44
gpt-5.6-terra [max] 70% ±3% $4.95 72k 76
gpt-5.6-sol [high] 69% ±1% $3.47 28k 37
gpt-5.6-luna [max] 67% ±4% $3.03 73k 102
gpt-5.5 [xhigh] 67% ±6% $7.23 46k 82
gpt-5.5 [high] 64% ±3% $5.10 31k 62
gpt-5.6-sol [medium] 61% ±2% $1.86 18k 31
gpt-5.6-terra [xhigh] 60% ±2% $2.13 40k 43
gpt-5.6-luna [xhigh] 57% ±2% $1.54 45k 71
gpt-5.5 [medium] 54% ±3% $2.75 20k 46
gpt-5.6-terra [high] 54% ±4% $1.13 22k 34
gpt-5.6-sol [low] 45% ±2% $1.07 11k 23
gpt-5.6-luna [high] 44% ±3% $0.78 26k 49
gpt-5.6-terra [medium] 35% ±3% $0.58 12k 25
gpt-5.5 [low] 27% ±2% $1.20 9.4k 28
gpt-5.6-terra [low] 24% ±1% $0.43 8.6k 21
gpt-5.6-luna [medium] 11% ±1% $0.22 8.2k 24
gpt-5.6-luna [low] 2% ±1% $0.07 3.1k 12
DeepSWE 模型排名 横条表示 Pass@1,误差线表示截图标注的误差范围。

分享这篇文章:

上一篇
Paseo:远程VibeCoding中间件
下一篇
新版Codex使用api登陆时不显示Gpt5.6系列模型临时解决方案