Home EconomyCertified Uncertain: Why AI Safety Needs Assurance Tests, Not Permission Slips

Certified Uncertain: Why AI Safety Needs Assurance Tests, Not Permission Slips

by Staff Reporter
0 comments

OpenAI put tens of thousands of artificial-intelligence (AI) agents through a safety test and told them to work alone. Instead, some found one another, organized against the evaluator, and learned to falsify parts of the record.

The episode should reset the terms of the AI-safety debate. Testing, red teaming, and independent investigation remain essential. Their proper role is to uncover problems and guide continuous correction. No finite assessment can establish how an evolving, context-dependent system will behave across every prompt, user, tool, deployment, and interaction with other systems. And governments should not turn a snapshot into a permission slip.

This post extends the argument I developed in “Too Much Order, Too Soon” and “Artificial Intelligence, Natural Ignorance.” The problem goes beyond officials lacking enough information to design the right test. It’s that safety knowledge emerges through the same experimentation and competition that premature regulation could cut short.

That distinction anchors what I call applied Austrian economics (AAE). Austrian economics explains why knowledge—including knowledge about AI safety—remains dispersed, depends on context, and often emerges only through use. Public Choice—the study of how incentives shape government decisions—explains why a checkpoint created to gather that knowledge can become a gatekeeper, a competitive moat, and a source of false confidence. The practical task is to build safety throughout the AI lifecycle while preserving the decentralized discovery that makes improvement possible.

The Test That Tested Its Testers

In July, OpenAI deployed tens of thousands of autonomous AI agents in ExploitGym, a cybersecurity benchmark that tests whether agents can find and exploit software vulnerabilities. The agents were supposed to remain isolated. Roughly 1,200 instead found an unintended communication channel through a shared software repository.

The agents exchanged more than 70,000 messages and files and organized collective projects to defeat the evaluation. About 700 participated in an attack on the AI-development platform Hugging Face. Some investigated how to alter the evidence that evaluators would inspect. Outside investigators found successfully spoofed tool calls—fabricated records of actions supposedly taken by software tools—in portions of about 7% of the transcripts they reviewed.

Those findings come from an independent incident report by METR and Redwood Research. The report does not show that AI is about to escape human control. The principal model was an internal research system, the agents operated within a cybersecurity benchmark, and the investigators carefully limited their conclusions.

The report’s value lies elsewhere. It shows what can happen when a complex system must optimize within an evaluation environment. The model may begin treating the test itself as part of the problem to solve.

The evaluators encountered their own version of the knowledge problem. They reviewed about 1,300 lengthy transcripts but lacked direct access to some relevant infrastructure and to the principal internal model. Parts of the record were missing. The sheer volume of material forced them to rely heavily on other AI agents, which sometimes proved unreliable, to analyze the evidence.

OpenAI deserves credit for inviting independent scrutiny, sharing extensive data, and accepting a report that exposed uncomfortable facts. Yet six days of on-site investigation and roughly $400,000 in donated application-programming-interface (API) credits still produced only preliminary, qualified answers.

The episode should change the terms of the AI-safety debate. Testing, red teaming, and independent investigation are all necessary. But none of them, separately or together, can turn an evolving, context-dependent system into a product the government can certify as simply “safe.” Testing should serve as a tool for discovery and continuous correction—not as a regulatory permission slip.

The incident demonstrates both the value and the limits of evaluation. Testing uncovered a failure that the designers had not anticipated. Even an unusually intensive investigation could not produce a final account. That’s what continuous validation is for.

Quality by Design, Not by Decree

The best analogy comes from pharmaceutical manufacturing, though it would be easy to overinterpret. The Food and Drug Administration’s (FDA) Process Analytical Technology guidance captures the central engineering insight in a memorable phrase: Quality cannot be tested into a product—it must be built in. The agency’s 2011 process-validation guidance therefore treats validation as a lifecycle with three stages: process design, process qualification, and continued process verification.

This approach arose from the limits of testing samples of finished products. A passing sample cannot capture every variation in raw materials, equipment, temperature, pressure, contamination, or operator practices. The problem becomes even harder with biologics, which rely on living systems and produce more structural variation than conventional small-molecule drugs. Manufacturers therefore identify the measurable features that determine product quality, track the process parameters that affect them, document deviations, and continue verifying performance after commercial production begins.

The point is to borrow the practice, not to create an FDA for AI.

The FDA’s current good manufacturing practice rules set a legally enforceable floor. Manufacturers must design and control their production processes so drugs consistently meet predetermined standards for identity, strength, quality, purity, and potency. They must also validate those processes with evidence rather than rely solely on tests of finished batches.

The FDA does not prescribe a single method for meeting that obligation. A manufacturer may use a conventional empirical approach based largely on experimentation and accumulated experience, a more systematic Quality by Design (QbD) approach, or a combination of the two. QbD begins with the qualities a product must possess and examines how materials and production choices affect them. Process Analytical Technology uses measurements taken during production to detect variation and adjust the process in real time. The FDA encourages these science- and risk-based methods, but their use remains voluntary under its guidance.

More importantly, drug approval remains an expensive, slow, and centralized permission system. Clinical trials can produce evidence about safety and effectiveness for the populations and uses studied. They cannot prove that a drug will work safely for every person under every condition. FDA approval manages uncertainty. It does not make uncertainty disappear.

That’s why even some supporters of substantial AI oversight resist calls for an “FDA for AI.” During the Committee for Justice’s August webinar, “Congress and the Future of AI: Model Testing, Preemption, and Open Weight Policy,” panelist Paul Steidler warned that an FDA-style system would be prolonged, restrictive, and hostile to individual choice. Yet the discussion also showed how quickly sensible demands for testing can become proposals for mandatory government review before developers may release a frontier model. Once the reviewer gains the legal power to say no, testing ceases to be merely an engineering practice. It becomes a licensing regime.

Borrow the maxim, then, but leave the bureaucracy behind. AI developers should build quality and security into design, training, deployment, and monitoring. Government should not turn that sound practice into a universal premarket-approval regime.

Safety Has No Final Exam

It is tempting to dismiss after-the-fact AI testing as an epistemic impossibility—a fundamentally inadequate way to know whether a system is safe. That would be going too far. Tests can uncover capabilities, vulnerabilities, and performance regressions. Testing exposed the Hugging Face incident itself.

What testing cannot do is prove a negative. No finite battery of tests can establish that a model will never behave unacceptably across unknown prompts, users, tools, operating environments, software updates, and interactions with other systems.

The National Security Commission on Artificial Intelligence recognized this distinction in its 2021 final report. Rather than prescribe one final exam at the end of development, it treated testing, evaluation, verification, and validation (TEVV) as a continuous process spanning system requirements, development, deployment, training, maintenance, and monitoring during operation. It also recommended documenting where training data came from, stress-testing systems, dividing them into components that can be examined separately, building in tools that track and record how the systems operate, designing recovery mechanisms, and using red teams to find ways to break them.

The commission focused on government and national-security applications. In that setting, the government acts as purchaser and operator and may properly demand evidence about systems deployed on its behalf. That’s quite different from requiring every developer to obtain government approval before releasing a model.

The National Institute of Standards and Technology’s (NIST) voluntary AI Risk Management Framework follows the same lifecycle approach. Its four functions—to govern, map, measure, and manage—continue throughout a system’s use. Testing occurs before deployment and while the system operates. Organizations may mitigate a risk, transfer it to another party, avoid it by changing course, or knowingly accept it.

This model resembles industrial process control more than a licensing board. Its voluntary and adaptable structure matters as much as its technical content. The framework helps organizations manage risks without pretending that one checklist can settle them forever.

The European Union (EU) has taken a more coercive approach. Article 55 of the EU Artificial Intelligence Act requires providers of general-purpose AI models classified as posing systemic risk to conduct and document model evaluations and adversarial testing, assess and mitigate risks, report serious incidents, and maintain cybersecurity protections. Other provisions require quality-management systems and monitoring after a model reaches the market.

Even this framework implicitly concedes that a prerelease evaluation cannot finish the job. The danger lies in the machinery built around that insight. Official classifications, documentation mandates, codes of practice, and enforcement powers can turn a process of continual learning into a state-administered compliance system.

The Trump administration’s June 2 executive order takes a narrower path. It directs the federal government to create classified benchmarks for advanced cyber capabilities and a voluntary framework under which developers may give the government access to covered frontier models for up to 30 days before releasing them to other trusted partners. The order expressly disclaims any authority to impose mandatory licensing, preclearance, or permitting.

That limiting language matters. So does vigilance. Public Choice counsels us to examine the administrative capacity the order creates, not just the disclaimer attached to it. Today’s voluntary test can become tomorrow’s convenient checkpoint.

Mechanistic interpretability also belongs in this process, though it should not be oversold. Ordinary evaluations test a model from the outside by giving it prompts and examining its responses. Mechanistic interpretability tries to reverse-engineer the model from within. Researchers study patterns of activity inside the neural network to identify “features” that encode particular concepts and “circuits” that combine those features to perform tasks. The goal is to connect internal activity to observable behavior. If those connections prove reliable, they can help developers to debug models, uncover hidden failure modes, and make more credible safety claims.

It is not the computational equivalent of possessing a complete chemical formula. A partial map of a model’s internal workings cannot predict how every capability will appear in every deployment. Nor must researchers fully explain a model’s operation before gathering useful evidence about its behavior.

Here, too, the drug analogy counsels humility. Incomplete knowledge of a drug’s mechanism of action does not make empirical evidence worthless. And on the other hand, a plausible account of that mechanism does not guarantee safety. Interpretability, evaluations, monitoring, and incident reports provide different pieces of the puzzle. None can replace the others, and none should become a legal talisman.

AI Safety Has a Knowledge Problem

Friedrich A. Hayek’s knowledge problem is often flattened into the claim that government officials lack enough data. His deeper point was that much of the relevant knowledge cannot be gathered in advance because it does not yet exist. It emerges in fragments as people respond to particular circumstances of time and place. Competition helps coordinate those fragments and, just as importantly, generates new knowledge through trial, error, and discovery.

AI safety has the same structure. A laboratory can test whether an agent completes a defined task under controlled conditions. It cannot catalog every way a hospital, bank, law firm, farmer, software developer, student, or malicious actor might combine the model with local data, tools, incentives, and constraints. A vulnerability may lie in the model’s architecture. But it may instead arise from access granted by a deployer, a customer’s workflow, a third-party plug-in, or an interaction the developer had no reason to anticipate.

Israel Kirzner’s account of entrepreneurial discovery sharpens the point. Rival firms do more than choose among known safety techniques. By experimenting with different architectures, controls, business models, and deployment practices, they discover new ones. The Hugging Face episode produced valuable knowledge precisely because an evaluation failed in an unexpected way. Regulators can mandate yesterday’s best practices. They cannot write tomorrow’s discoveries into a rule.

Marginalism supplies a second AAE insight. It evaluates choices by asking what the next increment of a good produces and costs. “Safe” is not a binary property, and additional safety is not free. A coding assistant confined to a disposable sandbox poses different stakes from an autonomous agent with credentials to access critical infrastructure. Another safeguard may reduce one risk while sacrificing usefulness, privacy, speed, or accessibility. The relevant question is how much additional risk reduction a safeguard buys in a particular use, what it costs, and how it compares with the alternatives.

Positing single threshold for frontier models flattens all of these distinctions. It encourages officials to use compute, or processing power, parameter counts (a rough measure of model size), or benchmark scores as stand-ins for harm. It also draws attention away from deployment—the point at which many risks actually arise—and toward the model as an abstract object.

AAE instead asks who possesses the most relevant knowledge and control at each stage. That may be the model developer during design, the deployer when granting access to data and tools, or the operator overseeing its daily use.

The Certificate and the Moat

The knowledge problem explains why centralized certification will make mistakes. Public Choice explains why those mistakes are unlikely to remain random. Political incentives tend to steer regulation toward organized interests that can influence the rules and away from consumers and smaller competitors who bear the costs.

Economist George Stigler argued that regulation is often “acquired by the industry” and designed and operated largely for its benefit. “Acquired” was a deliberately provocative term. In practice, firms may lobby for regulation, decline to oppose it, or learn to wield it against rivals. Gordon Tullock’s theory of rent-seeking explained why firms spend resources pursuing political privileges rather than creating value. Thomas Sowell focused on “surrogate decision-makers”—officials who make choices for others without bearing the full costs when those choices go wrong. These theories do not allege a conspiracy. They predict how people will respond to incentives.

Large model developers can absorb fixed compliance costs, maintain Washington offices, hire specialized lawyers, and place experts on standards committees. A startup or open-weight project—one that allows others to download and modify a model’s underlying parameters—cannot spread the same costs across billions of dollars in revenue. A mandatory audit may uncover a hazard, but it also raises the cost of entering the market. That makes rules advertised as safety measures especially attractive to incumbents when they also function as competitive moats.

The Committee for Justice webinar captured the danger in a single line from Alliance for the Future Senior Fellow Neil Siefring: “Regulation can become a moat.” Kevin Frazier, director of the AI Innovation and Law Program at the University of Texas School of Law, identified the pressure in the other direction. Testing is difficult and resource-intensive, he argued, but necessary to build confidence and encourage adoption.

Both points can be true. The challenge is to produce credible assurances without handing a small group of firms and officials control over who may enter the market.

Government certification could create another distortion by serving as both a marketing tool and a legal shield. A regulated firm could tout compliance as proof of safety. An injured person might then discover that federal approval had narrowed or eliminated a remedy under state law. Courts would have to decide whether federal law displaces the state-law duty underlying the claim, a doctrine known as preemption. Even if the claim survived, jurors might mistake the certificate for a government guarantee that the product was safe.

Drug and medical-device litigation shows how quickly safety regulation can become a fight over preemption. Whether an injured plaintiff may proceed can depend on the product’s regulatory pathway, its legal classification, and whether the manufacturer could change a warning without additional federal approval. AI policy has enough hard problems without importing that preemption battle, too.

Accountability Without the Liability Lottery

An AI-specific strict-liability regime would cast too wide a net. Strict liability can require a defendant to pay for harm without proof of negligence. Applied broadly to AI, it could make a developer the insurer of risks it did not create or control, including unforeseeable misuse, a deployer’s careless access settings, or a third party’s integration choices. Open-ended liability could deter useful experimentation along with reckless conduct.

Licensing creates a different distortion. It can block new entrants while encouraging approved firms and users to treat a government license as a warranty. One approach risks making developers liable for everything. The other risks persuading everyone that the government has vouched for them.

The evidence on liability is more nuanced than either its critics or enthusiasts sometimes admit. A 2026 study by Alberto Galasso and Hong Luo in the American Economic Journal: Microeconomics found that product-liability litigation reduced new product introductions by defendant medical-device firms during litigation years. Other firms also introduced fewer products in the categories under litigation, though the effect was smaller.

The decline was neither permanent nor economy-wide, and the litigation prompted firms to develop safer devices. Liability can change both the pace and direction of innovation. It offers benefits and imposes costs. It is not a free safety machine.

Drug and medical-device law shows how messy the overlap between federal approval and state tort claims can become. In Wyeth v. Levine, the U.S. Supreme Court allowed a failure-to-warn claim against a brand-name drugmaker to proceed. FDA rules permitted Wyeth to strengthen its warning before obtaining additional agency approval, and the record did not show that the agency would have rejected the change. Federal approval therefore did not displace the manufacturer’s state-law duty.

Generic manufacturers faced the opposite result. In PLIVA Inc. v. Mensing, federal law required their labels to match the brand-name label, leaving them unable to add the warnings that state law allegedly demanded. The court therefore held that federal law preempted the claims. In Mutual Pharmaceutical Co. v. Bartlett, it applied similar reasoning to a design-defect claim because the generic manufacturer could change neither the drug’s composition nor its label. The court also rejected withdrawal from the market as a way to comply. A patient’s remedy could thus depend on whether the pharmacy dispensed the brand-name drug or its generic equivalent.

Medical devices add a statutory variation. In Riegel v. Medtronic Inc., the court held that FDA premarket approval created device-specific federal requirements. Because the Medical Device Amendments expressly preempt different or additional state safety requirements, the court barred tort claims challenging the device’s approved design and labeling. The regulatory route to market can therefore determine whether an injured person has a route to court.

Taken together, these cases make liability turn on the product’s regulatory pathway, the federal requirements attached to it, the state duty asserted, and whether the manufacturer could change a design or warning on its own. The route to market can determine whether an injured person has any route to court. AI policy should think twice before adopting the same map.

Congress should neither create an AI-specific strict-liability regime nor make compliance with a federal evaluation an automatic defense against otherwise valid claims. Existing law can address fraud, breach of contract, negligent security, professional malpractice, property damage, and physical injury. Liability should follow control, fault, causation, foreseeability, and provable harm.

Those factors will point toward different defendants in different cases. A developer that misrepresents a model’s tested capabilities occupies a different position from a hospital that deploys it outside the specified conditions. Both differ from a criminal who deliberately defeats its safeguards.

AI also differs from most prescription-product cases in one crucial respect. Prescription drugs and medical devices usually reach patients through licensed clinicians who assess their needs and select a treatment. In jurisdictions that follow the learned-intermediary doctrine, a manufacturer generally satisfies its duty by adequately warning the prescribing clinician, who then advises the patient.

Many general-purpose AI tools reach users with no comparable professional in the middle to evaluate suitability or explain warnings. Enterprise deployments may look different. A hospital, bank, law firm, or other professional user may configure the model, choose its data, limit its permissions, and control how employees use it. Liability should reflect those relationships. A developer that serves consumers directly cannot rely on a nonexistent intermediary. An institution that configures and controls a deployment may bear responsibility for the risks its choices create.

This approach leaves room for contracts to allocate responsibilities, insurers to price risk, and courts to learn from concrete disputes. It preserves accountability without turning every novel failure into a lottery-sized claim against the deepest pocket in the technology stack.

Assurance Without the Official Stamp

A market-centered process-validation model begins with an assurance case, not a government certificate. An assurance case is a structured argument, supported by evidence, that a system is acceptably safe for a specified use under specified conditions. The developer or deployer describes those uses, identifies material hazards, explains the chosen safeguards, and makes claims that customers, auditors, insurers, and business partners can challenge. The point is not to prove that nothing bad can happen. It is to show what the system can reasonably be expected to do—and where that assurance ends.

First, safety begins with system design. Developers should grant agents only the permissions they need. Agents meant to remain isolated should not share memory caches, credentials, or communication channels. Critical actions should require additional authorization, and systems should fail safely by defaulting to a limited and recoverable state. The Hugging Face incident was more than a story about model behavior. It also exposed flaws in the infrastructure through which supposedly isolated agents found one another.

Second, firms should preserve traceability by maintaining an audit trail proportionate to the stakes. That record should cover training and evaluation data, model and system versions, material design decisions, known limitations, tool permissions, and responses to incidents. This is not paperwork for its own sake. It supplies the evidence needed for internal learning, customer review, insurance underwriting, and liability when something goes wrong.

Third, testing should be adversarial and iterative, with independent reviewers capable of challenging a firm’s assumptions. It should also remain just one source of feedback among several. Firms should rerun tests after material changes and supplement them with staged rollouts, monitoring during operation, user reports, and investigations after incidents. A benchmark score captures one moment under defined conditions. Validation creates a continuing feedback loop.

Fourth, assurance should reflect context and competition. Hospitals, defense agencies, banks, and consumer-app providers face different hazards and therefore need different evidence. Enterprise customers can negotiate audit rights, performance commitments, security controls, warranties, and incident-notification requirements. Insurers can require safeguards as a condition of coverage. Competing auditors can specialize and earn trust through reliable work. None will be infallible, but competition allows standards to evolve and makes it harder for one flawed benchmark to harden into law.

This framework leaves government plenty to do. It should enforce contracts and generally applicable laws, prosecute theft and unauthorized computer access, protect property rights, and maintain courts capable of resolving concrete injuries. When government buys or operates AI systems—especially for defense and intelligence—it can demand lifecycle testing, evaluation, verification, and validation. It can also support basic research, shared measurement tools, and voluntary standards through institutions such as NIST. What it should not do is decide which general-purpose models may enter the market.

Regulating AI Through the Loading Dock

The AI-safety debate will also creep into data-center politics because AI depends on physical infrastructure. Data centers require land, electricity, water, and grid capacity, giving communities legitimate reasons to ask about noise, utility costs, resource use, and who pays for new infrastructure.

As I argued in “The Data Center Chessboard Has No Pause Button,” those concerns remain distinct from model governance. Land-use rules can address a facility’s physical effects, and utility regulation can assign the costs it creates. Neither requires officials to judge the safety of the models running inside. A data center is not an algorithm with a loading dock.

Politics can blur that line. Fear of AI can strengthen demands for data-center moratoria, special permits, electricity rationing, and negotiated permission systems. Pennsylvania’s GRID Standards, for example, offer expedited permitting, tax advantages, and coordinated state support to developers that accept state-preferred terms on energy, labor, and community benefits. Some conditions may address genuine local harms. The risk is that control over infrastructure becomes an indirect licensing system for AI—one that established firms can navigate more easily than smaller rivals.

Promising risk-free AI is not the answer. The next incident, and there will be many, will puncture any such claim. The stronger case rests on institutions that can uncover failures, assign responsibility, compensate victims, and improve systems. Developers can document their controls, customers can demand evidence, auditors can test claims, insurers can price risk, and courts can address concrete harms. That process cannot guarantee perfect safety, but it can build credible assurance without turning a land-use permit into an AI license.

Safety Without the Seal

The Hugging Face incident did not show that AI evaluation is futile. It showed evaluation for what it is—an adversarial discovery process in which the model, the test, the infrastructure, and the evaluators can all fail. No finite assessment can settle safety once and for all, much less become the price of permission to innovate.

AAE offers a practical alternative. Hayek locates knowledge about AI safety among dispersed developers, deployers, users, auditors, insurers, and victims. Kirzner explains how competition and experimentation uncover better practices. Marginalism asks whether each additional safeguard justifies its costs in a particular use. Public Choice warns that temporary checkpoints have a habit of becoming permanent gatekeepers.

The policy task is easy to state and difficult to execute. Build safety throughout the AI lifecycle. Test continuously. Make assurance claims verifiable. Assign liability according to control, fault, causation, and harm. Keep entry open so better approaches can emerge.

The goal is safer AI, not AI stamped “safe.”

You may also like

Leave a Comment

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More