First, we gave models a scratchpad: instead of forcing a one-shot answer, let them spend more tokens reasoning through a problem. In a sense, you get more capability out of the same model by giving it more serial computation.
That reasoning is also useful because humans can read it. You can inspect how the model reached an answer, spot mistakes, and potentially debug or monitor it.
But there’s an obvious incentive to make that reasoning cheaper. Compress it, shorten it, remove redundant words. The problem is that if you keep optimizing for efficiency, the reasoning can drift into shorthand or “Neuralese” that still works for the model but becomes gibberish to us.
So there’s a tradeoff: more efficient reasoning vs. preserving one of the few windows we have into how the model reached its conclusion.
🔗 Privacy-friendly: https://yt.chocolatemoo53.com/watch?v=iuHddnIzKRA