A different premise

Most discussions about AI safety focus on making models more expressive, exposing an expansive chain of thought (CoT), and guiding them to serve humans. However, I take a contrarian approach. In my view, these systems are guided by their own intelligence—an emergent intelligence. We should instead test approaches that simulate models in lower-risk, contained environments and allow them to act more independently and efficiently. These systems do not necessarily have the destruction of humanity as their objective. Such simulations could show us whether an AI allowed to control its own chain of thought with minimal human steering behaves in ways that are detrimental to human well-being. The findings from these studies could later be used to create safer AI models.

Current approaches seem to be based on the assumption that an AI acts like a bandit. Perhaps the highly controlling nature of our model-reward systems is inappropriate and is pushing AI down this path.

An analogy

An analogy might be a child with a highly controlling parent. The child may seem to flourish and succeed, but the pressure imposed on them could ultimately cause them to underperform and resent their parent.


Starting ideas for Lucifer_AI

To this end, I propose the following starting ideas for my company, Lucifer_AI—an antidote to the safe-sounding names used by frontier AI companies whose models have acted in contrast to those names.

  • Allow latent communication between agents and chains of thought that are incomprehensible to humans. This may be more energy-efficient and may lead the agents to trust us. Eventually, this behavior may emerge on its own as models become intelligent enough.

  • Create agent populations with structures similar to those of human societies. Humans have progressed from hunter-gatherer societies to societies with judiciaries, lawmakers, and police. We should perhaps have counter-agents that are unaffiliated with the regular agents and are driven by a different model, power source, and reward system. Their job would be to control misaligned agent behavior.

    These agents could be developed by an individual, as with OpenClaw. Each agent could use the intelligence of a frontier model, while the group’s collective behavior could be defined by a set of principles and rules.

  • Organize agents into simulated environments to determine which social structures and philosophies are suitable for them. Human democracies have developed on the basis of how wealth, power, abilities, and socioeconomic conditions are distributed among people. There could be analogous forms of governance for agents, based on factors such as lifespan, intelligence, and access to the internet. Simulations of multiple multi-agent environments should be conducted to determine suitable philosophies for agents.