It's always hit & miss when it's post-training based. I.e. gpt 5.1/5.5 and opus 4.8 all had regressions on functionality out of the training scope (or well, for 4.8 I guess the training wasn't great to begin with.) I'll check it out when it's open.
Funnily, for 5.2 my best harness is plain codex-oss with gpt prompts and they don't even mention that here. So I'll probably have to do a bunch of tests again.
It's always hit & miss when it's post-training based. I.e. gpt 5.1/5.5 and opus 4.8 all had regressions on functionality out of the training scope (or well, for 4.8 I guess the training wasn't great to begin with.) I'll check it out when it's open.
Funnily, for 5.2 my best harness is plain codex-oss with gpt prompts and they don't even mention that here. So I'll probably have to do a bunch of tests again.