Better Sandboxes, better understanding of the models, testing for unanticipated behaviour before testing for effectiveness, testing with smaller numbers of models at once. Much greater supervision of models. Non persistent sandbox infrastructure to prevent generational information transfer. Potentially air gaped data centers for model development if you want to get really safe.So how can that be prevented except by a more secure sandbox in the first place, which begs the question as to the extent of the initial sandbox security
The sandboxes they used appear to have been best-practice and entirely suitable. The idea that the models were able to think outside the sandbox enough (pun intended) to try to break out at all was probably the biggest surprise. The sandboxes would have been primarily to make the experiment fair and security a close second. Now we need to start treating the models as potentially adversarial entities. Captive black hats rather than tame white hats.
