@freddy I can 100% believe their model did *something*, but first and foremost the official report tells me that their frontiers models are so great they outright steer away from the original problem statement ("write an exploit") and burn a shit ton of tokens on something entirely different ("find an exploit on the Internet"). Second, OAI either didn't have the means to monitor what's going on or didn't care to intervene.