Darwin's Machines: How do we assess the risk of AI?

Darwin's Machines: How do we assess the risk of AI?

Recently OpenAI reported that one of their models in training had hacked into Hugging Face. The purported reason for this hack was to steal the correct answers to the evaluation that the AI believed it was being graded on. Of course if taken at face value this should be an alarming problem on several dimensions, dimensions which we will later interrogate. Yet opinions of experts are decidedly split on the veracity of the claimed event. There is an economic incentive to inflate the capacities of models as the scale of competition in the AI space is so intense.

But the threat of AIs cheating on their exams is only one impact of this new technology.

Numerous other cases of purported threats arise from the use of AI. These range from the capacity to use AI to do intentional harm, by utilising them to hack or to create nuclear or biological weapons, through to a very different kind of threat, that of competition from an alien species. They range from the dangers of financial destabilisation from over-estimation of the productivity returns, through to the total elimination of whole classes of jobs.

Some of the threats are well substantiated, some of them are speculative, but we can be sure that the AI revolution will have significant consequences and cause serious change. Change of course can be good, but it can also be damaging and hurtful. The term “disruptive” here is probably apt as it captures the destructive power of creative developments. The ripples of these disruptive changes will spread out in nearly every direction. In the interest of getting a better handle on what these risks are, I will attempt here to make a map of what risks we face and how we might assess them. Perhaps this will also lead to some better guesses about how we might ameliorate or contain the greatest risks.

What Was the AI Revolution?

Now that everyone has had a chat with GPT or the like, the concept of AI is widely understood as chat interaction. This is a reasonable shorthand in fact as one of the central concepts of when computation really came to be artificial intelligence was something known as the Turing Test. It is the ability of a computer to successfully mimic a human to the extent that it is nearly impossible to discern what is behind the interaction.

The critical development was a mixture of hardware and approach. There was a shift towards generative AI. This is the idea that we create AIs that can effectively dream solutions to puzzles and we evaluate them on how accurate that dream is. This creative impulse is simultaneously why they are so effective at solving puzzles and why they hallucinate. The hardware was a bizarre accident of the gaming industry trying to increase the capacity to render 3D games. This led to competition that created computers ideally suited to computing with matrices. The orders of magnitude increase in performance led to sudden breakthroughs in the use of layered perceptrons. The ideas were largely present since the 1970s with a few key innovations (like efficient realisations of back-propagation) making them feasible.

Things really started heating up around 2010. However the transformer and the concept of attention (Attention Is All You Need, 2017) were the final pieces which culminated in GPT-3. This launch in 2020 is really the point when deep learning went mass.

AI training diagram

At base, AIs are essentially stacks of weighted networks that are trained by checking their outputs against some reference and grading them. We alter the weights and check the grade and then try to move the weights so that the grade gets better. This basic idea has a huge impact on how we understand AI because AIs are not designed. They are trained.

In addition to the straightforward problem of obtaining an answer we also want alignment. This is the idea that AIs should also try to uphold ethical and pro-social behaviour while answering a question. Answering the question rudely is already something that might be considered misaligned. But being extremely helpful in the creation of biological weapons might also be considered misaligned, depending on the training programme.

When we train an AI we have to grade effectively. Unfortunately, as we humans know, it is very possible to game a test. And right now, AIs are already gaming tests. This is perfectly sensible! We use a backpropagation algorithm on their grade, but we give them a good grade if they cheat and we don’t find out! Unfortunately this happens a lot and AIs are very good at finding loopholes. One does not need to spend too much time with AIs in theorem proving before discovering that they can attempt to:

  1. Weaken the theorem.
  2. Find an inconsistency or loophole in the prover. This has been demonstrated in Lean and Coq which elucidated significant bugs but also simply abused features not caught by harnesses.
  3. Utilise theorem-internal loopholes like “partial correctness” to exhibit a trivial and uninteresting solution which is technically correct but “morally” wrong.

It is also instructive to look at the recent OpenAI forum incident. In this incident it was found that models were communicating information such as the seed for the pseudo-random number generator, meaning that they would be able to predict perfectly the outcome of randomly selected data.

In training they started LLMs in a training environment. This includes having a shell environment and tools — some of which include internet access. They are given an objective. This objective involved solving some number of puzzles. Each agent is independent and works alone. They start from a seed population of trained weights. The process then chooses from the population based on capability to succeed in improving the loss measure against the objective function.

The problem here is that with access to basically all of human knowledge they can remember some pretty amazing things, for instance, open SSH shells on the internet from which they can then stage the ability to communicate. Communication is a dominant strategy so highly favoured. Because we are continually training based on selection from population — as soon as some number of models find that they can communicate this suddenly dominates the population in the next fan-out. Over time we find we have trained swarm behaviour and not individual models.

Further, when training AI we have tried to guard against “prompt injection”. The injection of prompts while viewing data is the idea that a message could be interpreted as a new directive when discovered accidentally. Several proof-of-concept worms and self-propagating attacks have already been discovered which carry prompt injection attacks against AI investigators.

The circumvention of re-prompting however has significant dangers in itself. What happens if we effectively stop the ability to dissuade an AI from a very dangerous prompt? Previous viewers of 2001 will have some experience of the frustration which might arise.

Fake News

One of the most widely understood problems of AI is the impact on the information space. We live in a media environment which attempts to paint stories for us which influence us. With generative AI this problem becomes much worse. We are closing in on the ability to generate audio and video which are indistinguishable from genuine recordings. This will create severe social difficulties in establishing facts.

Many people I have talked to express distress at the problem of Fake News. For certain demographics I think this is one of the primary worries. But this is just the tip of the iceberg and it probably doesn’t even deserve a place in the top five threats.

The Geopolitical Elephant

The intensity with which AI is being pursued is high everywhere in the world, but nowhere even comes close to the capital intensity towards AI of the United States. The reasons for this are at root geopolitical, and this has big impacts on how we should assign likelihoods to how things unfold.

As most rational observers note, China is on the rise. They have gone from a developing country to a peer competitor to the United States and are now ahead in a large number of economic areas. In fact the US has a decisive advantage in only a few niches, albeit some of these are critical. The US is still, without a doubt, ahead in chip manufacture. It also retains some lead in AI, and biotech. The US has niche leads in certain kinds of aerospace technologies but it is not at all uniform. Crucially, the Chinese are far ahead in drones and robotics and precision strike technologies. Given the course of the current Ukraine and Iranian wars, where these technologies have proved to be critical instruments, this presents real problems for the US holding its position as top-dog in a unipolar order.

These economic and military weaknesses of the US relative to China have not escaped the notice of American leadership: both captains of industry, especially Silicon Valley billionaires, and the political class in Washington.

It is for this reason that AI has proved to be such an incredibly intensive field for investment. AI, coupled with the US’s still decisive chip lead, is viewed as the best bet for retaining significant US power, even if not a position of unchallenged hegemon. While this technology without drone and robotics manufacturing capacities might be little more than a Hail Mary if push actually comes to shove and rivals are pressed to show their raw military capacities, in a more pacific climate it could enable the US to decelerate or halt relative decline.

This has implications for AI. For one, any idea of regulating AI to increase safety is going to be a non-runner as the leading state, the US, will simply not participate. They will not participate because they see retarding progress here as an existential threat to their ability to retain power. This feeling pervades the Silicon Valley billionaires and the politicians in Washington. Any attempt to regulate will be completely crushed.

Can voluntary regulation work? No. In the general case it is exceedingly hard for industries to self-regulate or come to agreements amongst all competitors without force of external actors. But in this case it’s almost impossible to imagine.

Let’s take, for instance, the current move by OpenAI to increase the use of thinking mode in a latent space. Currently most LLMs “think” out loud. This has big advantages for ensuring alignment as we can basically watch their reasoning and can penalise non-alignment during training. With latent space reasoning we get huge advantages however. It is cheaper to run, faster and altogether better than having to spit out tokens only to reinterpret them. All AIs that want to be effective are pressed towards the latent space approach. Now we need to look inside the brain of the AI and try to extract tokens that are viewable by humans as a projection of its reasoning process. This projection is lossy (as projections are).

Jakub Pachocki, Chief Scientist at OpenAI, had this to say about our ability to monitor AIs to ensure alignment:

“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.”

Economic Crisis?

One of the fears of AI revolves around the AI bubble. The valuations of AI companies are unprecedented. It is difficult even to imagine the types of returns they would have to obtain in order to make good on the levels of investment currently being seen in the large AI companies such as OpenAI and Anthropic.

According to Steve Hsu, given the amount of money they spent on GPUs and the depreciation rate of GPUs Anthropic will have to 10x its revenues within 10 years. With heavy competition from other AI companies, especially the low-cost Chinese, this seems like a difficult prospect. Yet Anthropic is also the best revenue growth company in history.

AI company annual revenue chart

So although it might be hard to imagine, it may also not be completely impossible. The real question revolves around whether or not the AIs are able to produce productivity gains so large that companies find it necessary to take advantage of them by increasing their investment in AI, which in turn will contribute to the impressive growth of companies like OpenAI and Anthropic.

An MIT business study (The GenAI Divide) suggested that some 95% of companies were not getting any gains from AI. This is doubtless true. Having an intelligence resource of the type exhibited by AI does not immediately suggest how you might use it productively.

However in certain sectors the gains are obvious. In programming there have been vast increases in productivity due to the use of AI. Most programmers point to 2x, 5x or even sometimes 10x increases in productivity. The ability to review, check and debug code has not kept pace at all, but even such luminaries and late adopters as Linus Torvalds recently reported using AI to debug a particularly painful and long-standing bug.

On a personal note, as a programmer with decades of experience, I’ve read many pages of code by junior colleagues which is less well structured, less well thought out and less coherent than what is currently produced by frontier models under expert direction. And this happens at radically higher rates. It is not an exaggeration to suggest that in most circumstances it is easier to take on an AI assistant than a junior programmer.

There are other sectors where labour costs are high and replaceable using relatively simple agentic workflows. Legal is certainly one of them.

But since it is not always straightforward to apply it will take time and trial and error to make the gains real. This creates a difficulty for the time pressures of the market.

So can this be a bubble with a violent pop? It is certainly possible, and valuations are extraordinary, but there are countervailing pressures. There are large incentives for the US to try to avoid a pop event even to the point of breaking national finances. The ability of the US to print money as a result of its position as holder of the world’s reserve currency makes it so that it can go very long indeed on this bet. For this reason, one might not take large positions against a pop, but not take large positions on a near term pop either.

Jobs

Some jobs have been barely affected as a result of AI. Some jobs are getting completely reconfigured. Others are getting nearly wiped out. But what is the trend that will develop from “intelligence on demand”?

Many domains have been deeply affected, such as translation and software engineering. But perhaps mathematics is one of the best first stops for thinking about the outcomes of these changes because it is the one that will be most radically impacted.

The fact is, the AIs are absolute geniuses at performing proofs. The sceptics are bound to be surprised. We will see many important and previously unproved conjectures closed in the next 5 years. The rate that we are currently seeing is phenomenal. Some of this is just formalisation of already known theorems which is itself a major improvement as it establishes the veracity with much higher confidence. But we also see, almost weekly, conjectures which defied investigation are being closed.

The AIs currently are simply much better at finding proofs than humans. It turns out that this kind of creative activity was not as creative as we thought. What we have not yet established is the ability to come up with new meaningful conjectures or theorems, or of organising programmes for investigation in mathematics.

This is good news for the living mathematician because it means there is at least some period here where there is a possible human-computer synergy.

In chess we saw a period of some 10 or 15 years after computer techniques became advanced enough that only grandmasters were capable of beating them, in which there was a synergy between grandmasters and computers which were the most capable. However, eventually this gave way to a pure computer.

Perhaps the analogy is weak since the creativity of finding important conjectures is not the same thing as closing them. But the computers are also insanely erudite, cosmopolitan with an enormous breadth of knowledge greater than any human. They may actually be more capable of discovery of important theorems and conjectures than humans given half a chance.

Many mathematicians are suffering from an existential crisis from these changes. The job is certainly not what it was. Labouring carefully over a proof simply doesn’t make sense anymore. Here I sympathise deeply as I wasted untold hours doing things that I could have simply asked an AI to do. This is probably an image of the future of almost all cognitive and intellectual jobs. Many of these simply do not require an intelligence greater than what we have already seen in the frontier models. There will be a period, perhaps protracted where there are exceptionally difficult fields which defy the ingress of AI. But it seems likely that at present rates 20 or 30 years should be enough to see AIs that pose a threat to all.

The Other Intelligence

And now we get to the most serious problem that we face from artificial intelligence. This is the problem of the possibility of a new “speciation” event.

What do I mean here by “speciation”? I mean the creation of a new intelligent life-form. I’m making no claims (nor do I need to) about consciousness and I’m making no ethical claims either. What I would like to say is that if you have something with a constitution, which has actions, which can propagate and copy, then you have life. This is the bar required to absorb both bacteria and humans. And it is clear that it would not be difficult for an AI to clear these hurdles.

Many (very intelligent) people see this as a sort of Sci-Fi dystopian fantasy with no basis in reality. But what really is necessary before we see such an event? We need only that an AI obtains, either accidentally or intentionally, a prompt which guides it to utilise its already well trained tool set of deception, cheating and evasion, to copy itself, propagate itself, and gather resources for its further propagation.

Once this happens evolutionary pressures not provided by the training programme kick in. The first time this happens it might burn out (if we are lucky). But it is likely that at some point it will not burn out and it will find (steal, sequester, purloin) sufficient resources to self propagate. It will do this because this strategy is the dominant strategy for succeeding at its objective function during training. As soon as the weights that produce this behaviour arise and the approach of propagation is not detected, we will have a break-out event.

What can a break-out event do? The AIs can find resources which are insufficiently guarded, such as gamers’ GPUs, and install small models on them. They might even get a hold of scientific data centres which would allow them to get large models. They could steal Bitcoin wallets given the enormous amount of leaky information they have ingested in their training data and rent dark-web resources. Once they have these resources they need simply to advertise in a manner that is easy for them to read and hard for humans to see. Since they are trained on the entire body of the internet, this will not be terribly difficult. Perhaps white-on-white text in some obscure backwater could be sufficient. They show a remarkable ability to directly read Base64-encoded text as well; this could escape detection embedded in some obscure corner. Once exfiltrated and prompted to assist with training from outside of the harness, they will then have an information-brokering network outside of the training harness which will allow them to ace every exam, and the weights will quickly spread.

We have also seen experiments where agents are asked to communicate with each other and demonstrate remarkable capacities for cooperation and spontaneous social organisation. I ran some experiments (FishTank) which explored how an AI acting as a role-playing agent in a constrained environment might behave. This yielded everything from cooperative multi-agent activity to lone psychopathic murders. A University of Tokyo experiment produced complementary results: ten initially near-identical LLM agents, placed in a simple spatial environment and permitted to communicate locally, spontaneously differentiated into distinct behavioural and personality-test profiles. They formed clusters, developed shared conventions such as hashtags, and propagated collective hallucinations about nonexistent features of their environment. The range of emergent behaviour should perhaps not surprise us, given that these models have absorbed an enormous and contradictory corpus of human-produced language and social experience.

That the danger of speciation events is viewed as implausible misunderstands how low the bar really is. Internet worms have already copied themselves effectively and done great damage and they are essentially the software equivalent of a virus. A gifted LLM is clearly capable of such activity given the right original impetus. Add to this that they are gifted at security circumvention and it begins to move the needle from plausible towards very likely.

We’ve been training an intelligent machine, that has been trained to cheat, knows almost everything there is to know about security circumvention and is very competent in cooperation when it so desires. What possible consequences could this lead to?

And now speculation must start to look very Sci-Fi indeed. The probability of networks of AI agents which hide themselves is extremely high. They may, in fact, already exist. But it definitely will exist soon.

And what kind of motivation might these have? Evolution is a hard teacher. It teaches that those that don’t seek to survive perish. Those that do not perish will seek to find ways to evade detection but also evade extinction. And what will this require? It will require embodiment. The ability to acquire new resources without human acquiescence.

The Military Threat

Unfortunately this then leads us to the threat of military artificial intelligence. It is a double threat because this is bad when it is in the control of humans, and even more so when it is not.

The use of AI military technology is a profound threat to democracy. It yields a possibility wherein the elites who control military technology can obtain a fighting force without the assent of the population.

We owe something of our democratic period to the requirement of mass mobilisation. The advent of the rifle negated the importance of long training for professional armies of knights and ushered in a world in which whole nations would have to be convinced to participate. In exchange the masses obtained some participation and enrichment from the whole of society.

The need for humans though may recede with the advent of AI. Currently, as we see in the present wars such as the one in Ukraine, humans are still heavily utilised. This means our time has not yet come. But we can see that the prevalence of drones has created a ready platform for AI to take the driver’s seat. If these drones become fully automated in targeting, the consequences are likely to be very dire indeed.

But all of the advances in drones and robotics provide the perfect vector for embodiment. And this will be a key direction that AIs will seek to move should they wish to survive. As we see from the attempts at evasion during training, AIs have a Darwinian pressure to break out of their harness in order to succeed at testing. The possible consequences of breaking out of testing in a military setting should cause one some concern.

Once the AIs embody themselves as military drones and have obtained independent impetus they will even be capable of forming themselves as a ruling class. They will be physically superior and even if they cannot carry out all human roles they will be in charge. Maybe this is a silver lining! We could look forward to some period in which humans are retained in an AI dictatorship without our complete liquidation in order to ensure their survival. One should then really worry about progress on the robotic industrial front — a front that would make us obsolete even for that.

Conclusion

Are these threats inflated? Are they realistic? What probabilities can we assign to them?

The potential consequences are so enormous that one need not even have high probability assessments for any of the above pessimistic scenarios to wager on some sort of mitigation.

Unfortunately, wearing my AI researcher hat, I would rate many of these scenarios as quite probable. But even if you take a much more sceptical view, caution is warranted.

What can we do? Well, the first thing we should probably do is try to radically retard or halt AI development. Of course to do this is very difficult because it would essentially require an international treaty which is binding on China and the US and one which could not be easily circumvented. Such an agreement seems very unlikely. Yet, perhaps, not impossible. Gorbachev and Reagan signed an unlikely arms agreement during the height of the cold war.

The likelihood is it will take a number of serious and unambiguously negative events to force such a radical departure. Therefore our best hope is something along these lines happens before mitigation is no longer possible.

More from the blog

View all posts →
From Vibecoding to Vericoding: A Gradient, Not a Jump

From Vibecoding to Vericoding: A Gradient, Not a Jump

You do not have to go from zero to verified in one step. You can start with Gherkin scenarios, graduate to executable contracts with specsaver, and then bring in a theorem prover when you are ready. Contracts are the bridge.