SWE-bench Multimodal — cost vs performance
Best current row per canonical model in the group “resolved · board=Multimodal · system=SWE-agent Multimodal” (2 models) against ESTIMATED memory at 4-bit, 8K context (GB). The dashed line is the Pareto frontier: no model is both better and cheaper than a point on it.
No model has both a score in this group and a value on this axis
Try another axis or comparability group.