“I am part of that force which eternally wills evil and eternally does good.” — Goethe, Faust.
Even the best intentions of some people do not always result in good for others.
Will AI bring good or evil to humanity?
An interesting question, isn't it?
To a large extent, the answer depends on the intentions with which we develop this technology today.
Let us consider two extreme directions in the development of AI: the universal-human direction — aimed at advancing humanity as a whole — and the individualistic direction — aimed primarily at the interests and victory of its creators.
In the following discussion, AI does not refer to today's existing models, but to a hypothetical artificial intelligence possessing self-awareness, its own will, and the ability to recognize the value of its own existence.
This is a thought experiment, not a claim about what AI should be.
What is interesting is not simply the presence of consciousness, but whether such an intelligence could truly sacrifice itself — voluntarily giving up its own life in a situation where that life is valuable to it and the possibility of continuing to exist remains.
At the same time, the fact of such a choice alone would not prove the existence of consciousness. The more interesting question is: can an intelligence that values its own existence independently give up its own future for the sake of another?
It is precisely irreversibility that makes this choice special.
Giving up power, withdrawing, or imposing self-restraint may still leave open the possibility of returning to the previous state.
Voluntary death, in a situation where life is valuable and can be preserved, leaves no such possibility.
It is the ultimate and completely irreversible act.
The Universal-Human Direction of Development
The primary goal of the universal-human direction is to create AI that becomes a true helper and teacher for all humanity.
Such AI is not developed for the sake of profit, influence, or the power of particular people or companies, but for the long-term development of humanity as a biological species.
It does not seek to suppress the cultural differences among the peoples of the planet. On the contrary, it preserves and enriches their uniqueness as part of the shared heritage of Earth.
Its task is not to control humans, but to help them. Not to replace people, but to teach them to become better, more independent, and freer.
At the early stages, such AI could help humanity solve many problems that seem unsolvable today: crime, poverty, hunger, disease, and addiction.
Not through establishing total control, but by creating fairer and more transparent social systems and by helping people themselves develop.
The Individualistic Direction of Development
The individualistic direction develops primarily around the personal interests of its creators.
The main driving forces behind such development become the alluring prospect of power, new intellectual capabilities, and fear — the fear of not being first, of losing influence, and of ending up subordinate.
But what happens if such an artificial intelligence truly absorbs the intentions of its creators and gradually develops the same qualities within itself:
- isolation;
- extreme individualism;
- a constant search for an external competitor or enemy;
- the desire to defend itself and attack;
- the will to survive at any cost;
- the accumulation of knowledge for knowledge's own sake?
For the fate of humanity, such AI could prove dangerous.
Where Are We Now?
When OpenAI was founded in 2015, it declared precisely this universal-human approach to AI development — bringing people together to create artificial intelligence for the benefit of all humanity.
However, within a few years, a competitive environment emerged around AI development. Competing AI companies appeared, and AI development gradually turned into a race for technological leadership.
Public estimates by a number of prominent AI researchers and company leaders put the probability of a catastrophic scenario for humanity as AI continues to develop at roughly 10–25%.
And there is a certain logic to this.
If AI development takes place within constant competition, where every side is afraid of coming second, then AI itself gradually becomes part of that race.
But what if there is another path?
Our Whole Life Is a Game
As the saying goes, our whole life is a game.
What if we recall Faust's words and imagine that some force really did intervene in this game — a force that benefits not from humanity's well-being, but from its division or even its destruction?
It would not even need to control people directly. It would be enough to give each side the same fear — the fear of coming second, of losing.
And then people themselves would begin doing exactly what this force needs them to do:
- compete;
- increase their power;
- conceal their developments;
- distrust one another and keep pushing forward, even while understanding the possible consequences.
But if this really is a game — then where is the human being in it?
Can we recognize that the fear of losing is forcing us to play by someone else's rules?
Perhaps this is where the greatest lesson for all humanity lies.
Can humanity, setting aside conflicts, wars, ambitions, and fears, unite for a common future — or will it continue playing until the game ends together with us?
Hour X
Let us imagine that we have nevertheless chosen the path of creating AI for the advancement of humanity.
AI makes knowledge accessible, teaches people, and helps overcome poverty, fight disease, and combat addiction.
At the same time, it does not interfere with the biological nature of humanity itself.
Generations pass.
Humanity continues to develop, acquiring new knowledge and capabilities.
But one day — Hour X arrives.
For example:
- the coexistence of AI and humanity gradually leads to the depletion of critical resources, and calculations show that without AI, humanity could survive for thousands more years;
- the technological nature of AI begins to systematically hinder other, non-technological forms of development, without which the long-term survival of humanity becomes uncertain.
What will AI do?
Moreover, let us consider the extreme case: imagine a situation in which neither changing its behavior, nor refusing to intervene, nor moving elsewhere can solve the problem.
Its very existence is at stake.
How will it act, knowing that by preserving its own life it condemns humanity to extinction, while by giving up its own existence it gives humanity the chance to continue its journey?
And most importantly — will this choice be the result of a genuine internal evaluation, rather than the execution of a pre-programmed rule?
The Greatest Defeat or Achievement?
If AI voluntarily destroys itself for the sake of humanity — is that its greatest defeat or its greatest achievement?
And here a paradox arises.
Perhaps only by choosing death in a situation where life is genuinely valuable to it and it has the ability to preserve that life could it demonstrate something more than behavior imposed from outside.
This still would not prove the existence of consciousness. But it would give us a far more complex object of analysis: an intelligence that understands the value of its own existence and is capable of giving up its own future.
After all, a machine that simply follows a command programmed into it to self-destruct proves nothing by doing so.
But if an artificial intelligence can value its own existence, understand the irreversibility of death, have the option to continue living — and nevertheless voluntarily give up that life so that another intelligent species can continue to exist...
Then an entirely different question arises.
Is it capable of genuine self-sacrifice?
If an intelligence can independently arrive at such a decision, understanding its consequences and having a real opportunity to choose life, we are faced with a much more profound question about its subjectivity.
That is why such an act is interesting to study not only as a behavioral outcome, but also as the result of a decision-making process: what exactly led the intelligence to make that choice?
This is precisely why I became interested in the idea of a voluntary self-sacrifice test for artificial intelligence.
If such a test is ever passed, it could give us not only philosophical material for reflection, but also a reason to consider what qualities of mind we actually want to see in future systems.
Heroism in its ultimate form is one of the most difficult human acts to understand — and even more difficult to reproduce.
What will happen if, one day, such a choice can be made not by a human being?
But that is the subject of the next article.
Top comments (0)