Technology
Companies Under Pressure to Prove Safety Measures Work in Real-World Conditions

Companies developing and deploying advanced technologies are facing growing pressure to demonstrate that their safety measures are effective in real-world situations, rather than simply presenting policies, testing reports or assurances that their products are safe. The issue has become particularly important as artificial intelligence systems and other advanced technologies are increasingly being used by businesses and consumers. While companies routinely conduct safety checks before launching new products, experts argue that testing in controlled environments may not always reveal how systems behave once they encounter unpredictable situations and large numbers of users.
A system can perform well during a carefully designed test but behave differently when exposed to unfamiliar inputs, unexpected interactions or circumstances that were not considered during development. This has increased attention on the need for companies to continue testing and monitoring their products after they have been released. The debate is especially relevant to the rapidly developing AI industry. Companies are introducing increasingly capable models that can generate content, analyse information, write software and interact with external tools. With those capabilities come new forms of risk, making it more difficult to establish whether conventional safety tests are sufficient.
Independent testing is one way companies can provide greater confidence in their safety claims. External researchers can examine systems from a different perspective and potentially identify weaknesses that internal teams may not detect. However, independent assessments are useful only when researchers have enough access to evaluate the technology properly. Restrictions on access to systems, data or testing environments can make it harder to establish whether a company's safety claims accurately reflect real-world performance.
Safety evaluation also cannot necessarily end when a product is launched. New problems can emerge after a system is deployed at scale, particularly when users interact with it in ways developers did not anticipate. Continuous monitoring can help companies identify unusual behaviour, investigate incidents and make changes to safety controls when necessary. The same principle applies beyond artificial intelligence. In industries ranging from manufacturing and aviation to healthcare and construction, organisations have long relied on safety inspections, employee training, incident investigations and hazard reporting to identify risks before they result in serious harm. Measuring these preventive activities can provide a broader picture of safety than simply counting accidents after they occur.
This means demonstrating not only that safety procedures exist but also that those procedures are producing measurable results. Records of testing, identified weaknesses, corrective actions and subsequent improvements can provide stronger evidence than general statements about safety. Greater transparency could also help regulators, customers and the public understand how companies manage potential risks. For technologies that can have significant consequences, independent audits and credible reporting systems may become increasingly important as governments and regulators consider new safety requirements.
The challenge is particularly significant for emerging technologies because their capabilities can change faster than regulations. Companies may therefore have to treat safety as an ongoing process rather than a one-time requirement completed before a product reaches the market. The question facing companies is not simply whether they have safety measures in place. The more important question is whether they can demonstrate, with credible evidence and continued monitoring, that those measures actually work when their products are being used in the real world.



