pull down to refresh

It's always hit & miss when it's post-training based. I.e. gpt 5.1/5.5 and opus 4.8 all had regressions on functionality out of the training scope (or well, for 4.8 I guess the training wasn't great to begin with.) I'll check it out when it's open.

Funnily, for 5.2 my best harness is plain codex-oss with gpt prompts and they don't even mention that here. So I'll probably have to do a bunch of tests again.