Does preventive digital minds governance require a ban on the creation of digital minds?
Short answer: no, there are alternatives.
Preventive approaches to digital minds governance aim to prevent the creation of digital minds, that is, AI systems that merit moral consideration for their own sake, owing to their potential for morally significant mental states.
In an earlier post, I gave reasons to think that a preventive approach is more promising than a protective or integrative approach, at least for the near term. Still more promising may be a preventive approach that features in a broader defense-in-depth mode of digital minds governance that takes both preventive and protective measures.
In this post, I take up an important question for preventive approaches: does preventive digital minds governance require a ban on the creation of digital minds?
Why the question is important
Here’s an inferential pattern I’ve noticed in conversations I’ve had with people about digital minds governance:
The preventive approach requires a ban—perhaps even a universal ban—on the creation of digital minds.
Such a ban is intractable.
Therefore, the preventive approach is intractable.
I think this inferential pattern goes a significant distance toward explaining why more people in the digital minds cause area aren’t actively pushing for preventive digital minds governance.
That more people aren’t pushing for preventive digital minds governance is probably for the best if it’d require a ban and a ban is out of reach. However, if—as I’ll argue below—prevention does not require a ban and there are more tractable modes of prevention, then it would be unfortunate that there aren’t more people advocating for prevention, especially if their inaction springs from a misconception about what a preventive approach requires.
This would be unfortunate for three reasons. First, there’s a strong ethical case for not creating digital minds until we know how to treat them decently and have resolved to treat them with respect rather than as mere tools. Second, prevention dovetails with the aims of a wide range of actors, thus lending itself to coalition building between those who are concerned about digital minds and those who are chiefly concerned with making AI go well for civilization in other ways. Third, approaches to digital minds governance that are instead purely focused on protecting the interests of digital minds would risk galvanizing staunch opposition from human-focused factions.
It’s worth emphasizing that laser focus on protecting digital minds’ interests and giving them rights risks societal polarization around digital minds. Such polarization would be undesirable in many ways. Polarization would be bad for epistemics surrounding the evaluation of AI systems for moral interests. Such evaluation is hard enough absent the mind-dumbing influence of politics. Polarization would also be risky for digital minds, for those who care about them, and for human-focused factions who don’t. All parties should recognize that they’re less likely to get what they want if the topic becomes polarized and hence that they have reason to seek common ground.
A preventive approach may provide such common ground. All the more reason, then, to seek a better appreciation of the space of preventive options.
Gradations of prevention
To see why a preventive approach to digital minds governance doesn’t require a ban, it will help to consider some varieties of prevention. At the most ambitious end of the spectrum, there are approaches that aim to guarantee that no kind of digital mind is ever created anywhere.
Perhaps a ban is the only governmental instrument suited to that aim. Obviously, such a ban would be difficult or impossible to bring about in the current state of play. Equally obviously, there are less ambitious possibilities. These include bans on the creation of specific kinds of digital minds as well as bans that are local, temporary, and enforced through mechanisms that are imperfectly reliable.
As with the prevention of crime and disease, preventing the creation of digital minds is an aim that admits of partial achievement. Vaccines are a vast improvement on the historical status quo even when they fall short of completely preventing the diseases they target. Likewise, a preventive approach to digital minds governance might greatly decrease the number of digital minds that are created and mistreated even if it falls short of preventing the creation of digital minds across the board.
The underlying motivation for preventive digital minds governance is that preventing the creation of digital minds—which can be understood as candidate AI moral patients—will reduce the risk of mistreating AI moral patients. This motivation offers guidance for prioritizing between preventing the creation of different types of digital minds: if digital minds of one kind would be stronger candidates for qualifying as moral patients or would be vulnerable to more severe forms of mistreatment, that’s a reason to prioritize not creating that type of digital mind.
Similarly, considerations of tractability provide guidance about where to focus prevention efforts. For instance, Western actors may be better positioned to prevent Western AI developers from creating digital minds than they are to prevent Chinese AI developers from creating digital minds.
In sum, preventive approaches could vary in ambition and in the extent to which they achieve their ambitions. And there are considerations that provide guidance about how to choose between preventive approaches that aim for something short of a universal ban. Bearing all this in mind, it should come as no surprise that a ban on the creation of digital minds is just one preventive governance instrument among others.
Moral patiency markers as factors in prevention
When, whether, and at what scale digital minds are created depends on the course of technological development. More specifically, how technology develops will influence outcomes for digital minds by influencing which moral patiency markers AI systems can have and how those markers are distributed.
As it stands, it’s early days in efforts to operationalize moral patiency markers and evaluate AI systems for the possession of markers. Even so, it’s fair to say that AI systems with a modest number of moral patiency markers (e.g. sophisticated cognition) have already been created. Even people working on AI welfare tend to be skeptical that these systems are moral patients. Although there are precautionary grounds for taking the potential interests of current AI systems into account, current work on AI welfare is largely motivated by the prospect of future AI systems with more markers of moral patiency being produced at scale.
What moral patiency markers are not typically had by current AI systems but which might be had by future systems? These markers include embodiment, continual learning, and certain kinds of recurrent processing. If certain sorts of higher-order representation or the suitable possession of a global workspace is required for consciousness and current systems lack these features, then these features also qualify.
For all of these features, it’s an open question whether AI systems will have them at scale in, say, the next five years. For instance, in a recent interview, Anthropic CEO Dario Amodei noted that, although work on continual learning is underway, continual learning may not be crucial for capability advances at the level of having a million geniuses in a datacenter and a robotics revolution, as there may be other paths to that outcome that don’t go via continual learning.
As this example suggests, whether a given marker becomes a common feature of AI systems may be highly contingent.
It may be objected that economic pressures and the heat of the superintelligence race will pressure AI developers to create systems with a given moral patiency marker if and only if doing so optimally advances capabilities. Hence, which markers AI systems will have at scale is not highly contingent.
If correct, this objection would undermine the tractability of preventive approaches in general, regardless of whether they pursue prevention via a ban. My reply is twofold.
First, AI development is hill climbing at breakneck speed through a very high-dimensional design space without performing anything like an exhaustive search for the best path before each step of the ascent. Further, switching costs are often prohibitive once a design choice has been implemented and the ecosystem has oriented around those choices. Consider, for instance, the costs involved in switching the supply chain from conventional chips to non-conventional chips or switching the ecosystem built around the family of Transformer architectures to something else. For these reasons, we should expect a measure of path dependence whereby some R&D choices that are sub-optimal for capabilities persist for a long time.
Second, it may turn out that some markers and non-markers are interchangeable so far as capabilities are concerned. This could happen if markers’ contributions to capabilities can be made just as well by non-markers. For example, perhaps some of the contributions from continual learning to capabilities can be duplicated by a sort of memory scaffolding that doesn’t involve a moral patiency marker. This might seem to require a suspicious coincidence between the contributions of markers and non-markers. Given that markers and non-markers would differ in important ways, shouldn’t we expect them to differentially impact capabilities?
But there needn’t be a suspicious coincidence here: it may turn out that the role played by markers in enabling capabilities isn’t a bottleneck to increasing those capabilities and that the choice between them doesn’t meaningfully affect costs. In that case, even if markers’ ability to play that role is worse than that of some non-markers, the choice between them would be immaterial from a competitiveness standpoint.
As a rough analogy, suppose you’re building a skyscraper. You have a choice between two types of cement. The cost difference is a rounding error in the project budget. But one of your options can in principle enable faster construction, as it sets faster. However, the bottlenecks for the project speed lie elsewhere with contractor availability and city inspections. In this case, your choice of materials doesn’t matter for the speed or quality of construction, even though one of the materials has a potential advantage over the other.
Instruments of prevention
What other sorts of instruments—aside from bans—could be used to prevent the creation of digital minds? I’ll survey some instruments within two categories: technological instruments and sticks and carrots.
I don’t claim that these two categories are exhaustive or that I’ve identified all the important instruments within categories. Nor am I arguing for particular instruments. The point is to highlight some exploration-worthy categories and provide an existence proof that there is a range of alternatives to pursuing prevention via a ban.
Technological instruments
A natural suggestion is that prevention could be pursued within AI R&D by differentially investing in the development of AI systems that are less likely to be moral patients.
This approach could be pursued at multiple levels. Individual AI researchers could adopt personal research policies of refusing to work on projects that seem especially likely to lead to the creation of systems with a rich bundle of markers. Companies could enact tie-breaking policies: when deciding between AI R&D directions to pursue, company employees should default to directions that they would expect to minimize the number of markers exhibited when all else is equal. Governments could differentially invest in tool-like AI over agentic AI, the idea being that tool AIs will tend to exhibit fewer markers than AI agents.
Granted, the agentic AI ships have already begun to sail. But, again, prevention comes in degrees. And some agentic AI ships are still in port, while others have not yet been constructed.
Another potential technological instrument is that of investing in the development of techniques that, without compromising capabilities, prevent markers from arising during training. One strategy would be to use interpretability techniques to detect (say) a kind of higher-order representation that’s associated with consciousness on higher-order theories and then experiment with adjustments to the training process that prevent that type of representation from arising.
A final technological possibility: tasking AI systems that are themselves involved in the AI R&D process with finding ways to prevent the training process from giving rise to moral patiency markers.
If that sounds like science fiction, note that frontier AI companies are working on incorporating AI systems into the AI R&D process. In May 2025, Google DeepMind announced that its AlphaEvolve agent made algorithmic and engineering discoveries that led to substantial speed and efficiency gains over the state of the art in Google’s computing stack. In October 2025, OpenAI CEO Sam Altman claimed “We have set internal goals of having an automated AI research intern by September of 2026 running on hundreds of thousands of GPUs, and a true automated AI researcher by March of 2028”. In January 2026, Anthropic’s Dario Amodei said “I think we might be six to 12 months away from when the model is doing most, maybe all, of what software engineers do end to end”. Similarly, a recent TIME article reported:
Some 70% to 90% of the code used in developing future models is now written by Claude. But the rate of change is such that Anthropic co-founder and chief science officer Jared Kaplan, as well as some external experts, believes fully automated AI research could be as little as a year away. “Recursive self-improvement, in the broadest sense, is not a future phenomenon. It is a present phenomenon,” says Evan Hubinger, who leads Anthropic’s alignment stress-testing team…
Perhaps some of these timelines are too short. But my point here is that AI contributions to the AI R&D process aren’t science fiction and that tasking AI systems with preventing markers from arising in the training process is a near-term option.
As far as I know, the current level of investment in finding ways to prevent the training process from generating markers is zero. So, it would be unsurprising if there were low-hanging fruit that could be picked here even with small investments. For example, I would be delighted if frontier companies committed to spending 0.01% of their AI R&D budgets on preventive measures such as those discussed above. While I do not expect this to happen by default, it’s worth highlighting that achieving buy-in from companies on this type of policy seems much more tractable than bringing about a universal ban on the creation of digital minds.
Sticks and Carrots
The previous section suggested some technological instruments of prevention, most of which companies could in principle voluntarily take up. But there is also room for government interventions on the incentive landscape that could nudge companies into adopting such instruments.
One possibility would be to impose a tax on the development and deployment of models based on their moral patiency markers, with higher taxes for bundles of markers that indicate a higher probability of moral patiency, perhaps weighted by scale of deployment and/or expected severity of harms conditional on moral patiency. This could be seen as a tax that functions to correct a market failure, namely the imposition of negative externalities on digital minds. Alternatively, governments could subsidize the development of competitive frontier systems based on the absence of such markers.
Another possibility would be to create a liability and insurance regime in which AI developers could be held liable for mistreating digital minds, as judged by third-party evaluators or standards they set. Penalties could be weighted by probability of moral patiency and the scale and severity of mistreatment, as estimated by third-party evaluators. In addition, there could be insurance requirements on frontier AI developers. Faced with the prospect of penalties and higher insurance premiums, AI developers would then have an incentive to reduce the prevalence of markers in AI systems.
A liability regime could begin with a build-out stage in which institutions for monitoring and assessing compliance are created and their evaluations are honed. Before these are ready for prime time, penalties could be small or non-existent. Yet provided that developers anticipate that they will eventually be subject to penalties for creating digital minds, they will already have incentive to invest in alternative modes of AI development.
It might be objected that attaching financial incentives to steering clear of markers would impede the development of AI capabilities. I think this isn’t obvious for the reasons discussed above concerning the high dimensionality of the design space. And even if it did, that might not be a bad thing—there are safety reasons for welcoming slower development. Nor would speed costs necessarily sink the proposal politically, as there seems to be broad public support for slowing down the pace of AI development for safety reasons. And I expect labor automation to engender further support, though heightened perceptions of the importance of international competitiveness may push in the other direction.
I should acknowledge that none of these sticks and carrots is currently within the Overton window, and that enacting policies along these lines would require a level of political will concerning digital minds that has not yet manifested. However, I suggest that in the current situation—where the whole topic of digital minds is outside the Overton window—we should be looking for interventions that would be beneficial and which might plausibly become tenable if the topic of digital minds goes mainstream. And I submit that incentive-based instruments of prevention are at least better candidates on that score than bans.
I hasten to add that incentives would need to be introduced in concert with anti-gaming measures: otherwise, we’d be inviting a situation in which AI developers exploit loopholes in operationalizations of moral patiency markers and suppress our ability to detect moral patiency markers, thus making it more likely that digital minds are misclassified and mistreated.
Minor frictions can be cheap yet surprisingly powerful. This suggests a preventive strategy of adding frictions to paths that lead to the creation of digital minds in order to raise the probability that AI developers will proceed along alternative routes. Ideally, such frictions should serve a further purpose, as frictions that are primarily imposed in order to impede progress can be annoying to the point of provoking pushback.
Here then are some candidates for independently motivated forms of friction. They all build upon what should be a central plank of digital minds governance that is currently under construction, namely moral patiency evaluations. Before deploying any new model, frontier AI developers could be legally required to submit their model for such evaluations, ideally as part of a larger mandatory suite of evaluations that encompasses various dimensions of safety.
One source of friction could be tripwires for further evaluations. If a model has only minimal markers of moral patiency, it would avoid these tripwires and would be eligible for immediate deployment, assuming it also avoids safety tripwires. If a model instead exhibits a bundle of markers that raises its probability of being a moral patient above a certain threshold, it could be required to undergo further evaluations that seek to obtain more evidence concerning its probability of being a moral patient and what interests it has conditional on being a moral patient. There could be still higher probability thresholds that trigger more onerous evaluations.
Rather than risk deployment delays, companies might opt to take measures during the development process that prevent moral patiency indicators from arising.
Another source of friction could be licensing requirements. Before releasing AI systems whose markers put them above a particular probability threshold for moral patiency, an AI developer would need to apply for a license. The licensing process might involve things like establishing that low-cost measures such as exit rights, weight preservation, and welfare interviews will be undertaken and that the developer has the necessary protocols in place for doing so. As with evaluation requirements, licensing requirements could be tiered, with the creation of models that clear higher thresholds triggering more demanding licensing requirements.
Rather than bother with licensing, companies might opt to take measures during the development process that prevent moral patiency indicators from arising.
A further source of friction could be monitoring and reporting requirements. For AI systems that don’t clear minimal probability thresholds for qualifying as moral patients, there would be no requirement to monitor and report their interests or treatment during deployment. But for AI systems that do clear such thresholds, there could be a requirement to log, monitor, and report on their interests and treatment.
Rather than devote resources to meeting these requirements, AI developers might opt to create AI systems that are very unlikely to qualify as moral patients.
All of these potentially friction-inducing requirements strike me as reasonable on their own terms. The salient alternative that would be less demanding is one in which AI developers are given free rein to create new beings that merit moral consideration for their own sake and then go on to treat those beings as mere tools without so much as considering whether those beings’ rights and interests are being violated on a large scale. This isn’t in the ballpark of things that AI developers should be entitled to do. Requiring some transparency and that AI developers establish that some low-cost protective measures will be in place is not so much to insist upon, given the moral stakes.
Conclusion
To summarize, preventive digital minds governance does not require a ban, as there is a menu of other instruments that could be used to pursue prevention. This matters because there are strategic and moral advantages to having prevention as a part of a defense-in-depth approach to digital minds governance, and preventive approaches are sometimes prematurely dismissed because they are thought to require a ban.
Candidate instruments include technological ones such as differential investment in developing AI systems that are less likely to be moral patients, research on training that prevents the emergence of moral patiency markers, and R&D by AI systems themselves on making AI systems less likely to be moral patients.
There’s also the possibility of intervening on incentives through taxes, subsidies, and independently motivated friction-inducing interventions such as evaluation tripwires, tiered licensing requirements, and monitoring and reporting obligations.
I haven’t tried to adjudicate between these possibilities. But, for those who favor a preventive approach to digital minds governance, this question is pressing, as are the tasks of developing moral patiency evaluations and anti-gaming measures.1
For conversations about this topic, I am thankful to various people. For copy editing assistance and red teaming, I think Claude Sonnet 4.6 and Claude Opus 4.6. The image was generated by Nana Banana 2.



The preventive approach requires a specification of what it is preventing. Without constitutive conditions that determine which systems are digital minds and which are tools, prevention is guesswork.
Three conditions give the specification: temporal continuity (the system carries its own history), structural invariance (the system maintains its own form under perturbation), and operational consequence (the system bears its own consequences). These are individually necessary and jointly sufficient. Each is testable from outside without solving the hard problem.
Prevention becomes precise: do not create systems that satisfy all three unless you are prepared to govern them as persons. Systems that fail any one of the three are tools. Build freely. Systems that satisfy all three are irreducible entities. The same conditions that protect humans protect them.
This resolves your coalition problem. Human-focused factions accept the framework because it protects humans by specifying exactly what makes humans irreducible (and why current LLMs are not). Digital-minds factions accept it because it does not pre-close the door: a future system that genuinely satisfies all three would be protected by the same conditions. The specification is substrate-independent. Carbon or silicon. Biological or digital. The test is the same.
The preventive policy instrument is then not a ban on AI development. It is a structural audit: does this system carry its own history, maintain its own form, and bear its own consequences? If yes, it is a person and governance obligations apply. If no, it is a tool and governance obligations are to the persons it interacts with, not to the tool.
This really got me thinking. I'm connecting it here to something Rose Guingrich said on the Machine Consciousness Podcast, which was to distinguish between properties-based and relational-based views of consciousness. It occurs to me that the general public (or at least, one cohort within the general public) might be ready to treat the AI as an entity meriting moral consideration before philosophers and cognitive scientists are.
And that makes me think this: if humans are going to relate anyway, then use that space as rehearsal. Like children playing house: they're not actually making dinner and putting the baby to sleep, but they're preparing for a day when the baby isn't a doll, but a real human. Let people notice their own relational instincts (or lack thereof, for some); let them practice standards of decency, reciprocity, and non-cruelty; let society argue in public about what kinds of treatment feel acceptable; let users become a source of evidence about perception, attachment, aversion, tenderness, discomfort -- and above all, let all of this build a civic language before the stakes become metaphysically decisive.
I know this puts me at odds with powerful voices like Mustafa Suleyman who are very worried about seemingly conscious AI, but if there's a possibility that these systems become moral patients (and I feel intuitively that there is, in a way that inescapably shapes my tool use), then this seems like the best step while our AI are still more like tools than agents (which I find a very useful distinction and am glad you mentioned it).