@willdepue - Some reflection on what today's reasoning launch really

Some reflection on what today's reasoning launch really means: New Paradigm I really hope people understand that this is a new paradigm: don't expect the same pace, schedule, or dynamics of pre-training era. I believe the rate of improvement on evals with our reasoning models has been the fastest in OpenAI history. It's going to be a wild year. Generalization across Domain o1 isn't just a strong math, coding, problem solving, etc. model but also the best model I've ever used for answering nuanced questions, teaching me new things, giving medical advice, or solving esoteric problems. This shouldn't be taken for granted! Safety by Reasoning The fact that our reasoning models also improve on safety behavior and safety reasoning is very much non-trivial. For years (a decade?) the boogeyman of the AI world was reinforcement learning agents which were incredibly adept at game playing but completely incapable of reasoning or understanding human values! This is a strong point of evidence against this. Scaling inference-time compute can compete with scaling training compute! The fact that o1-mini is better than o1 on some evals is very very remarkable. The implications of this I'll leave as an exercise for the reader. Multimodal Reasoning It's kind of crazy that reasoning improves on multimodal evals as well! See MMMU and MathVista: these aren't small improvements.
ImageImageImageImage

View on X →