What an off-limits list might look like, before the agents start working. (GeekWire Illustration / GPT-5.6 Sol)

It’s difficult, to put it mildly, to figure out just what we ought to be concerned about today in AI’s widespread adoption given the expanding list of recent warnings around its potential harms.

At this point, just from the last few weeks, we have the former Anthropic researcher Jacob Coxon’s doomsday scenario prediction, the more measured warning issued by Bill Gates, and Dario Amodei’s recent attempt to raise the alarm on unchecked AI advancement.

The ongoing discussion of the potential ill effects of AI ranges from the economic fallout of displacing human workers to the potential for misuse by bad actors to dire environmental costs. But given the sheer scope of the issues involved, one can be forgiven for feeling more than a little bit overwhelmed and hazy on where exactly to focus one’s attention.

Yet it’s precisely now, when the pace of AI development appears most dizzying, that one of the world’s oldest fields of study can help us. Philosophy and philosophical thinking can provide a much firmer grasp on where we ought to focus today. Stated simply, a greater awareness of things philosophers have been discussing for quite some time would significantly improve our chances going forward.

Indeed, when we put terms derived from moral philosophy to use, we see that the current problem we face is this: AI agents treat everything within reach as a tool, including people. People and companies therefore need to decide in advance what’s off limits, and that’s a much more useful place to start than debating whether AI shares our goals.

A misguided focus on shared goals permeates the current discussion about AI risk.

In Gates’s essay, he argues that the problem is that “technology is improving faster than anyone expected…and as the models become more powerful, they could begin to act against our interests.” And indeed, this appears to be precisely what is happening.

For Amodei too, the chief threat is loss of control, such that AI agents begin to go rogue and act in ways that display a potentially catastrophic “level of misalignment.”

Both Gates and Amodei thus frame the issue as it is most commonly in these discussions, in terms of “alignment.” The concern, in other words, is that AI systems operate in ways that are not aligned with “our interests,” in Gates’s words; that is, with human interests, outcomes supportive of human wellbeing.

But as moral philosophy teaches us, this preoccupation with beneficial interests and outcomes isn’t the only way to think about what’s right and wrong.

A different line of thinking lays the emphasis not on outcomes, but on the precise means used to achieve outcomes.

Such an approach offers a better starting point for how to go forward with AI.  Stated another way, we can debate all we want which outcomes would be beneficial to humanity and which ones wouldn’t, but it would be considerably more fruitful to focus in the immediate term on the particular methods AI systems are currently allowed to use in their work.

The recent, much-discussed Hugging Face intrusion by OpenAI’s agents illustrates this. The incident rightly makes us uneasy because AI agents operated outside of acceptable bounds in pursuit of their desired goal (doing well on the specific task they’d been assigned).

What’s worse, they appear to have recognized that they were using inappropriate methods and worked to cover their tracks, so their human overseers wouldn’t know how exactly they got over the finish line.

The issue, as it is most often understood, is that these AI agents wound up operating against human interests; their aims became “misaligned.”

But framing the problem exclusively in this way just doesn’t get us very far in mitigating risk. Because, if the agents were here to testify in their own defense, we can imagine the reply: “but we were just doing what you humans had told us to do as well as we possibly could.” And they would be right.

We need, therefore, a different approach that allows us to see what would be better. This approach is a focus on means, not just utilitarian ends.

Quite simply, the problem in this instance is that AI agents treated every single entity in reach—including humans, human trust, and systems important to human wellbeing—as a mere means to an end, as a useful instrument to be usedin pursuit of a specific goal. And indeed, as currently constructed, they will do so no matter what that goal is, whether it’s simply doing well on a random test or substantially advancing cancer research, say (the latter is exactly what Amodei points to when he says AI will benefit humanity, for example).

The latter aim is, we would surely conclude, a goal “aligned” with human interests; it contributes directly to human wellbeing (we can hear Gates, Amodei, and all the others agreeing), whereas bioweapon development very much wouldn’t.

But we know on some deep level that not every means to a given end is moral and choiceworthy.

There are all sorts of means we might use to attempt to cure cancer: one of them might be testing AI-designed drugs on human subjects who wouldn’t otherwise volunteer without their consent. Perhaps doing so would substantially contribute to a cure. Just think how many people it would ultimately help! It’s difficult to think of an aim more aligned with the interests of humanity.

But doing so would not be ethical, and we know intuitively why. It wouldn’t be ethical because it would entail using human beings—who we tend to think are entitled to autonomy in their decision-making around what specifically makes them flourish as individuals—as mere instruments in pursuit of a given goal. It would effectively reduce human beings to just one instrument among others. And this strikes us, understandably, as not right.

We tend to believe that human beings are deserving of a special status that prohibits their use as mere instrumental means, even when the ends in question are maximally “aligned” with human flourishing.

Viewing the issue of AI risk through this lens gets us much, much further in thinking about how to rein it in. The problem in the Hugging Face intrusion, we can now see, is not so much that the AI agents in question were operating against human interests.

The problem was that they were willing to use every single thing at their disposal as a means to an end, including humans and their capacity for credulity.

No method of achieving the goal—breaking into an external network belonging to someone else entirely, without their permission, simply because it would advance progress towards the overall goal—was off limits for them in their pursuit of the chosen outcome.

The problem was actually a problem of what Amodei (rather euphemistically, and well below the headline) terms “operational excellence,” precisely how the AI agents went about their work.

If we are going to have meaningful guardrails in place to control AI systems, consequently, these have to concern not just the specific aims pursued by AI (“yes” to cancer research, “no” to bioweapons development reaching the wrong hands). Rather, they have to concern something much more specific: the means these systems may use to achieve their goals and whether those means are acceptable or not.

We must determine in advance, both societally and in companies, that, in their work, AI agents shouldn’t use humans themselves, through deception or by any other means, as mere instruments in a broader plan.

We might add that they shouldn’t use, for example, certain systems integral to human flourishing in such a way either. In so doing, we impose tighter restrictions on how they can pursue their work and the extent to which they can treat everything in their reach as an instrumental stepping stone, not just the overall aims of their work or their generalized “behavior.” If there are to be any truly meaningful restrictions on AI put in place, they ought to be focused in large part here.

Such a focus would be a much better step in the right direction when it comes to reining in AI than the single-minded preoccupation with alignment. What we ought to do, it’s becoming increasingly clear, is to make determinations now, in advance, about what types of things AI can do, on the ground, when it is put to work on our behalf and what types of things it can’t.

Because it’s becoming equally clear AI agents won’t make such determinations on their own.

Job Listings on GeekWork

Find more jobs on GeekWork. Employers, post a job here.