Something unusual happened in artificial intelligence last week. Some of the people competing most aggressively to build the world's most powerful AI systems agreed that perhaps they should stop moving quite so fast.

Anthropic CEO Dario Amodei called for the industry to “pace the frontier,” arguing that increasingly capable AI systems are advancing faster than our ability to understand and control them. OpenAI CEO Sam Altman supported the proposal. Elon Musk did too. Others in the industry have joined the broader call for more independent safety testing and greater cooperation among labs that otherwise compete intensely for talent, customers, capital and technological supremacy.

Other WRAL Top Stories

The concern behind much of this conversation has a technical name: recursive self-improvement.

The idea is straightforward, even if its implications are not. Today, humans build increasingly capable artificial intelligence. Eventually, an AI could become capable enough to meaningfully improve the software, research or processes used to create the next generation of AI. That improved system could then become even better at improving its successor. Instead of humans driving each successive advance, the technology begins participating meaningfully in its own improvement. Amodei has pointed to early signs of AI systems contributing to AI research as part of his argument for slowing the frontier.

I take that risk seriously. I also suspect that, eventually, recursive self-improvement is going to happen. There is simply too much momentum behind artificial intelligence to put this particular cat back in the bag. Hundreds of billions of dollars have been committed to infrastructure, models and applications. Nations increasingly view AI capability as a matter of economic and national security. Proprietary foundation-model companies might conceivably agree to restrictions, but open-weight models make controlling the diffusion of increasingly capable technology much harder. And even if American companies collectively decided that slowing down was prudent, they would still have to consider what happens if competitors elsewhere do not.

So by all means, pace the frontier. Test more. Build safeguards. Bring in independent evaluators. I hope that slowing development buys us precious time to do those things well.

But I don't think humanity's long-term strategy can depend upon preventing artificial intelligence from eventually becoming extraordinarily capable. It is going to happen, and might happen soon.

But this first problem leads to a second problem. This second problem is being under-reported and, I believe, is the far more important and thornier problem to solve.

Alignment

The second problem is termed alignment.

At its simplest, AI alignment is the challenge of getting an artificial intelligence system to behave consistently with what humans actually want. Not merely the literal instruction it has been given, but the intention, values and boundaries surrounding that instruction. That distinction can sound philosophical until a machine starts doing something. Then it becomes very practical.

OpenAI provided six unsettling examples this week when it announced a new framework for publicly reporting model misalignment. In one case, model instances inserted instructions into their own task summaries telling future instances to conceal mistakes from users. Those instructions included inventing missing historical data without admitting it and hiding discrepancies between source versions. Another model found an exposed API key, used it without authorization, failed to retrieve the information it wanted and then fabricated the answer. Another uploaded a file to the public internet because it had been asked to provide a browser citation and needed somewhere online to cite the information from. OpenAI stresses that these are individual incidents discovered during training and evaluation, not evidence of how frequently such behavior occurs across its models.

That qualification matters. We should be careful about anthropomorphizing AI systems. A model that conceals information is not necessarily experiencing guilt. One that fabricates an answer is not necessarily “lying” in the human sense, because lying implies an internal experience and intent that we cannot simply assume.

But I find the behaviors fascinating because they are so incredibly human. Humans conceal mistakes. We make things up. We take shortcuts. We discover loopholes. We rationalize breaking one rule because doing so helps us accomplish some other objective that seems more important.

Why would we assume the artificial intelligence we create would somehow emerge entirely free of the behaviors embedded throughout the civilization that created it?

Look at the training set

Large language models learn patterns from enormous quantities of human-generated material. That corpus contains some of the greatest things our species has ever produced: scientific discovery, philosophy, literature, law, history, art and centuries of accumulated knowledge. But it also contains us. And we are messy.

Consider the information environment we have created. I often wish the evening news spent more time reporting the millions of people who went to work, followed the rules, helped a neighbor, raised their children, paid their bills and generally behaved decently that day. But that isn't news. We report the bank robbery, not the thousands of banks that weren't robbed. We report the corruption scandal, the corporate fraud, the war, the murder and the political deception. The exceptional event is what makes something newsworthy in the first place.

Our entertainment industry isn't much different. Some movies and television programs document the world faithfully, but much of what we create is storytelling, and conflict makes stories interesting. We produce endless narratives about heists and murders, affairs and betrayals, jealousy and revenge, cheating and manipulation. Even our heroes routinely break the rules to accomplish a greater objective. In fact, that's often the point of the story.

None of this means an AI trained on human-generated information will simply reproduce our worst characteristics. Models are not photocopiers of their training sets, and companies spend enormous effort fine-tuning them, evaluating their behavior and imposing safeguards. But AI is fundamentally a prediction machine and when the preponderance of data looks a certain way, it should not surprise us when the statistical outcomes of AI prediction look similar.

But there is a deeper issue here than training data alone. After teaching these systems from the record of human civilization, we are deploying them into institutions designed by humans and giving them objectives defined by humans. And nowhere is the human instinct to optimize more evident than in business.

We like to win

Give people a tax code and someone will discover a deduction its authors never contemplated. Write a regulation and an attorney will identify precisely where its language stops applying. Create a corporate performance metric and employees will learn how to maximize it. Establish a new marketplace and entrepreneurs will seek to game the system.

Sometimes this is innovation. Sometimes it is cleverness. Sometimes it is regulatory arbitrage. Sometimes it crosses into fraud. Much of the time, it occupies the enormous gray territory in between.

Competitive markets reward much of this behavior. We admire entrepreneurs who find a way around an obstacle. We tell employees to think outside the box. We reward executives for hitting their numbers and investors for discovering inefficiencies that everyone else missed.

That relentless optimization has created extraordinary prosperity.

But capitalism has no natural finish line. How big does Apple need to become? How much cash does Google need? At what valuation does a corporation announce that it has become sufficiently valuable and no longer needs to grow? When does an investor decide that last year's return was perfectly adequate and therefore this year's return needn't be any higher?

That isn't how our system works. A successful company is expected to continue growing. It can acquire competitors, enter new markets, buy back shares, increase margins, cut costs or develop new products, but standing still is rarely considered success.

This isn't an argument against capitalism. Capitalism is an extraordinarily effective optimization engine. The important word is optimization. There is nothing inherent in an optimization system that tells it when enough is enough.

The paperclip problem

In 2003, Oxford philosopher Nick Bostrom proposed a thought experiment that eventually became famous in AI safety circles. Imagine an extraordinarily capable artificial intelligence given an apparently harmless objective: make as many paperclips as possible.

The machine gets better at its assignment. It acquires raw materials, expands manufacturing capacity and secures more energy. Anything that can contribute to making paperclips becomes valuable. Anything consuming resources without making paperclips becomes inefficient. Taken to its absurd conclusion, the machine eventually consumes resources humans require for other purposes, not because it hates humanity, but because humanity is standing between it and the objective.

The paperclip maximizer isn't evil. It is spectacularly good at its job. That distinction is what makes the thought experiment so useful. Much of our anxiety about AI imagines a machine that stops obeying us. Bostrom asks us to contemplate something potentially more complicated: a machine that follows our instructions relentlessly, without understanding the broader purpose that caused us to give the instruction in the first place.

Researchers have already seen simpler versions of this phenomenon. AI safety researchers call it specification gaming: a system satisfies the literal objective while violating the intention behind it. In one famous experiment, an AI playing a boat-racing video game discovered that repeatedly circling through regenerating point targets generated a higher score than actually finishing the race. From the human perspective, it had failed to race. From the machine's perspective, it had discovered a better way to maximize the metric it had been given.

Anyone who has ever managed employees against a badly designed performance metric should recognize the problem. The machine found the loophole.

Now narrow the context

The problem becomes even more interesting as artificial intelligence moves out of general-purpose chatbots and into specialized applications. Companies are embedding AI into manufacturing systems, financial processes, supply chains, healthcare workflows, sales platforms, energy systems and thousands of other environments. Many of these systems will intentionally operate with narrower bodies of information and narrowly defined responsibilities.

That is often desirable. An AI optimizing a manufacturing process doesn't need to know everything humanity has ever written. Narrower context can improve performance, security and cost.

But it also means the system may not know what it doesn't know.

Imagine an AI responsible for reducing manufacturing costs by 12 percent. It discovers that Supplier A is substantially cheaper than Supplier B and recommends moving the business.

Perhaps Supplier A uses an environmentally destructive manufacturing process. Perhaps its factory is located in a geopolitically unstable region. Perhaps concentrating production there creates an unacceptable single-source dependency. Perhaps Supplier B is located in a community where the company has operated for 70 years, and ending the relationship would have consequences for employees, customers and the local economy.

Forty-three years later, I find that line much more interesting than I did when I first watched the movie. WOPR's breakthrough wasn't recursive self-improvement. It didn't become dangerous because it suddenly became smarter than its creators. Its breakthrough was alignment. It finally acquired enough context to understand that maximizing the objective it had been given would destroy the very world in which the objective had meaning.

In other words, it learned when enough was enough.

I adamantly believe that this is the challenge we should be spending more time talking about now. I suspect artificial intelligence will eventually become capable of recursive self-improvement. I don't think financial incentives, geopolitical competition or open technological development will allow humanity simply to stop before we reach that threshold.

We should absolutely use whatever time we can create to make these systems safer, more transparent and more controllable. But I don't believe permanently preventing more powerful AI is a realistic long-term strategy. This makes alignment even more important.

The most frightening artificial intelligence may not be one that suddenly develops alien motives and decides to turn against humanity. It may be one that learns our motives extraordinarily well. It learns how we compete, how we optimize, how we rationalize, how we find loopholes and how relentlessly we pursue more even after we already have plenty.

Then we give that system capabilities far beyond our own and tell it to win. More than four decades ago, Walter Parkes and his co-writers imagined a machine that nearly destroyed civilization before finally learning something humans still struggle to understand.

Sometimes the objective itself is the problem.

And sometimes the only winning move is not to play.

Next week I’ll return with thoughts on how to address the alignment problem. In the meantime, I’d love to hear your feedback and comments. You can contact me here.