The California Institute for Machine Consciousness Research Program Whitepaper

· Updated October 7, 2026

Contents

This version of our whitepaper was published April 27, 2026.

The Explanatory Target

What is consciousness, and how does it relate to reality? As in ‘what is phenomenal experience, and how does it arise within a purely mechanical universe without being already there in one form or another?—how can one get consciousness from non consciousness?’. To many people this seems like getting something from nothing.

So, what is this phenomenality, this specific representational regime that takes the shape we call conscious experience, and can we look at it separately from the other mental phenomena that usually come alongside it?

The question of what consciousness is belongs to a long standing philosophical and civilizational project of fully grasping the nature of mind. In that, it is of importance to culture, because it bears on how we understand our own nature, past, present, and potential; to ethics, because it informs how we ought to relate to agents, biological and artificial, that may or may not be sentient; to technology, because progress on it would inform the design of systems that extend human cognitive capacity, reduce suffering in medical and psychotherapeutic contexts, and advance psychological development; and to the future of mind itself, because we may be approaching conditions under which artificial systems warrant moral consideration, despite currently possessing no principled basis on which to make that determination. And yet, consciousness remains without an agreed scientific characterization, without a validated theory connecting mechanism to experience, and without criteria for establishing its presence or absence in any system.

Many people outside the field of consciousness studies, and even some within, would also point out that there is not only no complete theory, but not even a generally accepted definition of the target phenomenon—that is, conscious experience. However, the core extension of the concept of consciousness is known to all those who engage with the topic. This core extension is a certain basic aspect of our direct phenomenology: something we are confronted with, and through which we recognize a specific something as requiring explanation—something that has not yet been satisfactorily addressed.

This “something” may itself still appear vague and show high inter-individual variation. Some regard the phenomenal self-model as a necessary component, others the more abstract idea of subjectivity, and still others certain forms of conceptual cognition. We are mainly interested in a minimal model of this basic form of phenomenology. In this way, the definition of the target phenomenon is not yet formalized, but it is indexically disambiguated. There exists a phenomenon, or a class of phenomena, namely direct, pre-reflective, perceptual structures with which we are confronted and which stand in need of explanation. This is true even for those who claim that the concept of consciousness has not yet been defined.

Even having pointed to what we mean, however, we are still faced with the difficulty of explaining it, of capturing it in a scientific worldview, to which consciousness seems particularly resistant. There seems to be a gap between mechanistic description and structure, however complete, and the existence of phenomenal experience as such. A well-known example that captures this intuition is section seventeen of Leibniz’s Monadology, which reads as follows:

[W]e must confess that perception, and what depends upon it, is inexplicable in terms of mechanical reasons, that is through shapes, size, and motions. If we imagine a machine whose structure makes it think, sense, and have perceptions, we could conceive it enlarged, keeping the same proportions, so that we could enter into it, as one enters a mill. Assuming that, when inspecting its interior, we will find only parts that push one another, and we will never find anything to explain a perception. And so, one should seek perception in the simple substance and not in the composite or in the machine. (GP: VI, 609/AG: 215)

We understand this intuition and see why there would be a perceived gap in any possible functional explanation. But we also think that we see how the gap can be closed by the derivation of the relevant models and formalisms in one’s own mind. We claim that consciousness can be understood, analytically and intuitively, as an abstract pattern woven by a mathematical Leibnizian mill.

Our aim is to describe conscious experience as a model structure. In doing so, we place ourselves directly within the discourse of the Hard Problem (Chalmers, 1995), a frequently mentioned and by Chalmers most famously described (apparent) conceptual limit of functional explanations with regard to phenomenal experience. A related concept is that of the explanatory gap. The discourse connected with it defines itself mainly along two lines. On the one side are advocates of the Hard Problem who substantiate the fundamental impossibility of a resolution of the problem; on the other side, people who do not understand the Hard Problem, do not find phenomenal experience particularly special or in need of explanation at all, and also some who generally regard the Hard Problem as the result of a metaphysical confusion and see the solution of the Hard Problem in explaining why certain people have a Hard Problem.

We do not claim that metaphysical disentanglement alone is sufficient for the Hard Problem to disappear. We see the challenge of the Hard Problem not in successfully explaining it away, but in giving an account of phenomenality as the specific kind of structured representation it is. Part of this will involve metaphysical disentanglement; the other part could be described as a representational disentanglement, or as an insight into representational necessities.

The shape of investigating the question has changed with the appearance of artificial systems whose behavior mimics that of mind, and whose structure might be analyzed in ways that could make the understanding of mind more tractable. Large language models now produce linguistic and behavioral outputs that were, until recently, treated as sufficient evidence for understanding, reasoning, and self-awareness in humans. This may be an opportunity to better understand the relationship between language, cognition, intelligence, self-modeling, and perhaps other mental states, but one that must be taken with care. In humans, conscious experience is, in most usual circumstances, tightly coupled with a variety of mental phenomena, including intelligence, a self-model, agency, emotions, and feelings. We are also housed in a body and exist in an environment, usually including other people, that shape, speaking both evolutionarily and at an individual psychological level, our experience and how we connect with and understand reality. In artificial entities, however, these different facets that usually co-occur in humans may decouple, and evidence of behavior corresponding to one might not imply that any of the rest come alongside it the way they generally would in a human or other biological animal.

We must rethink what adequate evidence for consciousness would be. Tests of performance on specific tasks, even as complex and demonstrative of intelligence as modeling language with high fidelity, seem insufficient for determining consciousness status, since systems directly optimized for output can produce that output without necessarily implementing the processes that in biological organisms are prerequisite to coherent experience; the simplest explanation for an LLM reporting consciousness may be that it is matching patterns in human-generated text about consciousness, not that it is conscious. But the converse cannot be established either: there is no proof that current AI systems are not, and cannot be, conscious, and the absence of structural criteria for making this determination is a methodological gap that may only widen as artificial systems grow more capable. There is also the question of moral status, and how to develop an ethics suitable for artificial entities that may be conscious but whose experience may be very much unlike the human condition upon which our existing moral intuitions and ethical institutions have been built.

We take inspiration from Richard Feynman’s famous remark, “what I cannot create I do not understand,” in seeing the construction of consciousness as the most promising path to understanding it. What is needed is a research program in which philosophy and construction discipline each other: philosophy generating hypotheses precise enough to guide implementation, construction demanding the specificity that philosophy, left to its own procedures, tends to defer, and results evaluated not solely through behavioral testing (which current AI has shown to be insufficient) but through interpretive analysis of internal structure, that is, through inference to the best explanation of what the system has built within itself. Construction without philosophical clarity risks producing powerful systems whose nature we cannot assess; philosophical analysis without construction risks the indefinite continuation of debates that have, for decades, proceeded without empirical ground beneath them. We can extend Feynman’s dictum: what we create without understanding, we cannot responsibly steward.

The California Institute for Machine Consciousness exists to carry out this program. We adopt a computationalist functionalist stance, addressed in Section 2 and more thoroughly developed in the Machine Consciousness Hypothesis essay, and propose a specific hypothesis of what consciousness is: a coherence-maximizing pattern implemented through second-order perception in self-organizing substrates. We work to build computational systems designed to test this hypothesis, examine their internal organization, and evaluate whether what we observe is better explained by consciousness than by its absence.

The field of artificial consciousness research has grown substantially in recent years, spanning academic neuroscience and philosophy, interdisciplinary consciousness science centers, non-profits focused on AI welfare and moral patienthood, commercial ventures, and industry research programs (surveyed in Appendix B). The activity remains fragmented across communities whose methods constrain the kind of progress each can make: neuroscience identifies correlates but cannot explain why particular neural processes give rise to experience or whether the findings generalize beyond biological substrates; philosophy generates theoretical frameworks that remain untested and are often not properly metaphysically grounded; commercial and industry efforts face constraints of proprietary interest and single-theory commitment; welfare-focused organizations assess existing systems rather than building systems to test theories of consciousness directly. What is not yet represented in this landscape is an organization combining independent non-profit structure enabling sustained fundamental research, philosophy positioned as foundational to technical work rather than auxiliary, a constructive methodology that builds and analyzes systems rather than only assessing existing ones, and consciousness as primary mission rather than instrumental to capability development, safety, or welfare policy, while fully committing to a specific metaphysical and epistemological stance. This is the position that CIMC occupies. The present document describes this research program.

Philosophical Foundations

Our research program adopts a computationalist functionalist stance, which describes a particular epistemological and metaphysical view that informs what kind of phenomenon we think consciousness is and what it would mean to implement it. The label “computational functionalism” is now connoted in different ways, some of which are far from what we see as the most accurate and general reading of the term, which is sometimes described as functional constructivism, or as epistemological computationalism in the following:

Here, we take strong physicalism to be the position of the epistemological primacy of physical interactions. There might be little doubt that minds are caused by matter and energy and the physical laws that govern them. According to the current theories of physics, then, minds are ultimately a consequence of quanta and their interactions. These interactions can be captured in computational theories and models, but, from the physicalist perspective, such computations (i.e. the complete formulation of regularities in the observables) are only a convenient way to describe what exists and happens in the universe. Epistemological computationalism (Wolfram, 2002) (Margolus, 2003) denies strong physicalism in favor of a universal primacy of information, on the grounds that all possible observations of the universe do not impart matter or energy, but information (i.e. discernible differences). The description of all conceivable regularities in the observed data is necessarily and sufficiently computational.

Epistemological computationalism does not deny the validity of physical descriptions, but takes an anti-realist stance and maintains that the physical universe is a possible computational theory that encodes the observed data. For minds and worlds to exist, computation is necessary and sufficient; the notion of an ultimate physical substrate is a superfluous metaphysical concept, since it can never be anchored in an immediate empirical observation. The exclusion of empirically inconvertible assumptions from physicalism leads with some inevitability to epistemological computationalism. While the differences between strong physicalism and epistemological computationalism are only on a metaphysical domain, there seems to be no such thing as an epi-belief: in epistemological computationalism, it is impossible to derive a non-computational concept of the mind, while some physicalists see its locus either in a non-functional substance property of the physical universe (Putnam, 1983), or in an interactional relationship that surpasses a computational characterization (radical enactivism, (Di Paolo, 2009)).

(Bach & Verdicchio, 2012)

Despite risking confusion with views attached to the more widespread interpretation of “computational functionalism,” we maintain its usage but clarify exactly what we mean by it.

Functionalism

Functionalism, as we use the term, is an epistemological position about the structure of any claim about any phenomenon, including consciousness, and including defining what that phenomenon is in the first place, that is, of picking it out as distinct from anything else. What an observation ever delivers is the detection of a discernible difference (information), and what an observer can refer to is therefore only what makes such differences: the operations by which a system evolves, the regularities in its dynamics, the patterns of its interactions with any other phenomenon. To characterize anything as an object of investigation at all is to specify the discernible differences its presence or absence makes, and such specification is already functional in form. The material of a substrate, insofar as it is accessible to observation or theory, appears as a set of constraints on what functions the substrate can realize; it does not appear as an additional ingredient over and above those functions.

This position should be distinguished from two others with which it can easily be confused. The first is behaviorism, which restricts its attention to externally observable behavior and denies the reality or relevance of internal states. Functionalism, as we mean it, makes no such restriction: the functional organization of a system includes its internal dynamics, its mechanisms of self-representation, and the operations by which it accesses and modifies its own states, so that the features of consciousness which are phenomenologically salient but externally silent are functional features in the broader sense we intend, making discernible differences within the system even where they make no differences an outside observer can detect. The second position to be distinguished is positivism, which restricts admissible claims to those that can be verified by external measurement. Because any observation whatsoever, including the internal observations by which a system registers its own states, is the detection of discernible differences, functional characterization is not a narrowing of what can be said about the world but the form that saying anything about the world necessarily takes. Functionalism is the rejection of the notion of a hidden essence.

Conversely, the rejection of functionalism is essentialism: the idea that objects can be determined by intrinsic properties that can somehow be known by an observer, but without any formalizable structure and bypassing any processes of observation and modeling. An essentialist may point at material or experiential reality as immediately given; a functionalist constructs it over changes in information. An essentialist about consciousness maintains the possibility of what philosophers call a “zombie”: a system that is functionally identical to a conscious being, produces all the same behaviors, implements all the same internal processes, and yet lacks experience. The functionalist denies this, on the grounds that experience, insofar as it has any determinable character at all, is itself a functional property, one that manifests in how the system processes, integrates, and responds to its own representations. The word “consciousness,” if it refers to anything at all, refers to whatever combines the functional features we attribute to consciousness, and there is no further property beyond those features to which the word could attach itself. In a parallel way, it does not make sense to speak of “zombie electrons,” particles that behave in every measurable and conceivable way as electrons do but somehow are not electrons, because the word “electron” simply names whatever combines the observable properties of electrons. The essentialist about consciousness must explain what consciousness is over and above the totality of its functional manifestations; the functionalist holds that there is nothing over and above, that the functional characterization, if complete, exhausts what there is to characterize. We emphasize that phenomenality itself is a certain shape, due to its representational transparency difficult to precisely capture, with which we as observers are primordially confronted, and which consists of a dynamic stream of discernibilities that admit of characterization.

Functionalism by itself establishes that substrates determine behavior only insofar as they determine the state space realized by that substrate: where different substrates implement the same functional organization, they produce the same state space; what matters about the substrate is exhausted by the constraints it places on what functions can be realized. But this alone does not establish that consciousness can be realized on non-biological substrates, since the relevant functions might turn out to require conditions that constitute biological matter (which itself describes a specific set of functions, but does not go beyond it). The argument for substrate independence in the strong sense, the sense relevant to machine consciousness, requires the further step of computationalism, plus Turing-universal computation.

Computationalism

By “computation” we mean not merely the behavior of computing technology, but the capacity of a system to represent and implement arbitrary distinctive states (describable by finitely resolved differences) and arbitrary sequences of transitions between them, conditional on those states. Every function, that is, every arbitrarily complex relational structure, that can be exhaustively captured by such a system is computational; and to create a computational model of a phenomenon means to decompose it into sequences of such state transitions.

Computationalism is best understood as a consequence of mathematical constructivism: the position that our access to reality consists in the manipulation of models, that is, of representations of observations and functions relating these observations to each other, and that therefore all knowledge is constrained by the limits of constructive representational languages. The Church-Turing thesis establishes that all such constructive formalizations of computation are equivalent, absent implementation-specific resource constraints such as insufficient memory or processing time; what can be represented constructively at all can be implemented by any universal computing system. Whether computationalist models can account for the observables of consciousness is an important touchstone for the acceptance of computationalism itself. For a fuller treatment of the argument from constructivism to computational universality, see the Machine Consciousness Hypothesis essay.

We may distinguish weak from strong computationalism. Weak computationalism states that all formal theories of reality must be computational: any representation we can construct, evaluate, and communicate is subject to the constraints of constructive languages, and therefore computational. Strong computationalism extends this to the notion of reality itself: because all elements of reality, insofar as they can be measured, observed, experienced, conceptualized, or thought about, must be subject to the limits of constructive representation, all the reality we can ever refer to is computational. The most thorough contemporary articulation of strong computationalism is Wolfram’s (Wolfram, 2002), whose Principle of Computational Equivalence holds that systems above a low threshold of complexity are computationally equivalent, so that the discrete state transitions of physical processes and those of digital computers are, in principle, capable of realizing the same functions. If this is correct, then any functions constituting consciousness, whatever they turn out to be, are realizable on machines implementing universal computation, given sufficient resources.

Computational Functionalism

Combining computationalism, the claim that all phenomena can be fully captured as discrete and finite (though possibly unbounded) state transitions, with functionalism, the claim that what an object or phenomenon is is its causal or operational role rather than an imagined essence, yields the core commitment of our research program: consciousness consists of operations on representations that can be characterized as computable functions, and any substrate capable of performing the relevant computations can in principle support consciousness. The Church-Turing thesis implies that consciousness can be realized on all substrates capable of universal computation, though it says nothing about what the resource demands of such realization would be; similarly, computationalist functionalism with respect to physics claims that all regularities of observable physics can in principle be captured by computer simulations, while recognizing that we cannot necessarily build computers large and fast enough to run them.

A Hypothesis on Consciousness

The Machine Consciousness Hypothesis

The Machine Consciousness Hypothesis proposes that general computational machines with sufficient resources possess the necessary and sufficient means to implement consciousness, and that the success of implementation can be established through analysis of internal structure, including interpretation of representational embedding spaces, and observable behavior. The MCH does not by itself specify what consciousness is, only that whatever it is can be computationally realized; to test the hypothesis, and to build systems that would constitute such a test, we need a specific account of what consciousness is and what its presence requires.

We offer such an account in proposing that consciousness is the simplest learning algorithm discoverable by evolutionary search to train a self-organizing biological substrate to become intelligent in service of agency. Consciousness is not a late achievement of cognitive sophistication but an early solution to the specific problem of how a system whose structure is not given but must be built, by the system itself, out of the coordinated activity of its own components, can bootstrap the coherent models of reality that agentic behavior requires.

This hypothesis makes three interrelated claims about the nature and function of consciousness.

Genesis

In human development, consciousness seems to precede complex cognition and seems present in infants long before perception, self-modeling, language, and reasoning have matured. No human achieves intelligence without first becoming conscious, and evolution has produced no alternative pathway we know of from unstructured biological substrate to sophisticatedly intelligent (where intelligence is defined as the ability to make models) animal agent. This developmental ordering suggests the functional role consciousness plays as the mechanism by which an initially unorganized substrate acquires organization, a process that operates on the substrate’s own activity patterns to iteratively produce the coherent representational structures that constitute a mind.

Systems with pre-specified architectures, or with optimization procedures that circumvent the bootstrap problem (as current machine learning largely does, by imposing structure through training regimes designed by engineers rather than requiring the system to discover its own organizational principles), may achieve intelligent behavior without consciousness. The hypothesis predicts that consciousness should emerge when the conditions that characterize its biological genesis (self-organizing substrates, absence of pre-specified architecture, and developmental pressure toward coherent agency) are reproduced. Where those conditions are absent, intelligence may be achievable by other means, and whether such intelligence is accompanied by phenomenology is a separate and open question.

Coherence

The observable operations of consciousness, including waking, directed attention, and the maintenance of a stable perceptual world, increase the coherence of simultaneously active mental representations, that is, they minimize constraint violations between partial models of reality maintained in parallel. An organism maintains not one model of its situation but many, operating at different timescales, resolutions, and levels of abstraction (sensory, proprioceptive, affective, memorial), and these models routinely conflict. Operationally, consciousness is the process that detects and resolves these conflicts, directing attention to points of disagreement and orchestrating the adjustment of competing representations until a globally consistent interpretation is achieved or the unresolvable remainder is suppressed below the threshold of awareness. The neuroscientist Christoph von der Malsburg calls this the coherence definition of consciousness (von der Malsburg, 1999). It connects naturally to Karl Friston’s Free Energy Principle, since minimizing constraint violations across an ensemble of partial models is formally related to minimizing prediction error within a generative model (Friston, 2010), though the relationship is one of compatibility rather than strict equivalence. Developing the formal connections between coherence maximization and free energy minimization remains an active theoretical interest of our program.

Second-order perception

Conscious experience (phenomenology) consists of perception of perception, where “perception” is a non-inferential registration of structured content in a process that is immediate, synchronous with what it registers, and not the product of deliberate inference or symbolic reasoning. Computationally, we distinguish perception from arbitrary information processing or registration, as a regime in which information exists as a specific, sustained, globally constitutive representational format. We are pursuing a precise characterization of this format, which remains an open problem.

In its minimal state, conscious experience may contain nothing but the bare registration of its own occurrence, that is, its own representation, requiring only that the perceptual process is present within the field it creates, but without specification of what is experienced or by whom.

Consciousness Distinguished from Concomitants

Consciousness is not synonymous with mind, intellect, agency, or other facets that usually co-occur with consciousness, or are part of our usual conscious experience.

The mind is the representational medium within which models of self and world take shape, capable of expressing the geometries of perception alongside the discrete structures of thought.

The intellect is the capacity for reflective, symbolic reasoning, a powerful but brittle tool whose purpose is to repair and extend the deliverances of perception and intuition.

Agency is the capacity of a system to exert control over future states, that is, to select among possible actions on the basis of internal representations of goals, preferences, or evaluative criteria.

The self is a sustained representation within the mind of what it is to be a particular agent, with particular concerns, capable of exerting control over portions of the mental process and of experiencing itself doing so. Consciousness can exist without a self (as in certain dream and meditative states where events unfold without an observer experiencing them as “mine”), without intellectual activity, and in its most minimal form without determinate content. The self, when present, is a conscious content, constructed within the mind and experienced by consciousness.

Sentience is the capacity for a self to have valenced experience, and is generally considered to be what makes a system a candidate for moral consideration. Sentience requires consciousness, but consciousness does not require sentience: the minimal phenomenal experience is conscious without being sentient in the morally relevant sense.

Sapience is the ability of a self to understand: to construct and evaluate models, grasp relationships, recognize what follows from what, and revise its own representations in light of what it discovers. The intellect, defined above, is the faculty through which sapience operates, primarily through symbolic, reflexive reasoning.

The psyche is the compound structure that integrates these capacities within a single system, comprising a personal self with motivational concerns (experienced as feelings and desires), embedded within a mind that models self, interests, and world, animated by consciousness, and capable of agency. The psyche has both conscious representations accessible to the self and unconscious motivational dynamics, emotional evaluations, and habitual patterns that shape experience but largely operate beneath the threshold of awareness. The psyche is the causal structure of a cognitive architecture that is required in its entirety to construct a complete artificial person.

All of these notions are to be distinguished from what we are aiming to address with our operational definition in the following paragraph. The only relevant explanatory target phenomenon is conscious experience per se, phenomenality itself, independent of specific contents or commonly co-occurring implementations.

An Operational Definition of Consciousness

For research purposes, we currently use an operational definition of (phenomenal) consciousness:

A system is conscious if it implements self-organized second-order perception that increases global coherence.

This is precise enough to generate predictions about what functional signatures should appear in conscious systems, what developmental trajectories they should exhibit, and what architectural features are sufficient, while remaining subject to revision as evidence accumulates. It captures both phenomenological and functional aspects.

Phenomenologically, conscious experience is second-order perception: perception that perception is occurring. To clarify, our direct non-conceptual first-order conscious experience is already a second-order representation.

We take second-order perception to be a specific case of metarepresentation that is to be distinguished from its broader class if we want it to result in phenomenology. Roughly, it should lead to phenomenology under the following constraints:

For this representation, the percept-level encoding (perceiving) and the encoding thereof (perceiving of perceiving) have to be realized both on the same level of representation (the initial percept-level encoding).

This representation is transparent, that is, it does not contain information about the fact that it is a representation on the level of the representation itself (it is epistemically opaque).

Functionally, this phenomenology acts as a coherence-maximizing operator on mental states. It is a meta-level process that monitors and regulates.

Phenomenology and function coincide: phenomenology itself is a functional structure; phenomenology is the model. Properly articulating this argument is core to resolving Hard Problem concerns and is an active thing we are working on, but we have not yet arrived at a complete explanation, or, in particular, one that would fit here. This seems to be a common issue in consciousness studies, as we observe that there does not exist a comprehensive account that addresses those concerns in an actually satisfactory manner, even though there are a number of people that mean to address them but are still often perceived as indeed not addressing them, e.g. Dennett, Graziano, Hofstadter.

We know that, to many, for various reasons, a specifically computational framework, with regards to the phenomenology of consciousness, seems like a crude reductionist approach and that we must seem like people that “just don’t get what it’s about.” But we emphasize that we do not mean computation in the sense of some notation that describes the phenomenon, and that we mean to address the phenomenon itself here. When we are talking about second-order perception, we are indeed talking about the phenomenology of consciousness, the clarity of the what-it’s-likeness itself. This is also to reiterate that we are not using a computational framework only to describe access consciousness, which is more obviously amenable to a computational description, and which existing theories of consciousness more readily address, but see phenomenality too as a specific computational structure that is a currently unknown representational achievement.

Relationship to Existing Theories

CIMC’s framework builds on and integrates insights from multiple established theories while maintaining a distinctive position.

Global Workspace Theory (Baars, 1988) proposes that conscious content is distinguished from unconscious processing by its availability to multiple specialized cognitive subsystems through global broadcast. GWT characterizes which information is conscious (the information currently being broadcast) but does not by itself explain what consciousness is or why global availability should produce phenomenal experience. CIMC’s coherence mechanism can be understood as providing an account of the dynamics underlying the global workspace: how representations compete for broadcast, how conflicts between competing representations are resolved, and what governs the integration process that GWT describes but does not explain. On our account, consciousness is not the broadcast itself but the second-order perception that orchestrates coherence across the representations competing for and achieving global access.

Higher-Order Theories (Rosenthal, 2005) propose that a mental state is conscious when it is the object of a higher-order representation. Our account also states that consciousness involves representation of one’s own representational states, but we characterize the relevant higher-order representation as perceptual rather than cognitive: consciousness on our account is second-order perception, non-inferential and synchronous with its content, and not higher-order thought in Rosenthal’s sense, which holds that the meta-representation is a thought containing concepts.

The Free Energy Principle and active inference framework (Friston, 2010; Friston, 2013) provides a general mathematical framework for self-organizing systems, proposing that such systems maintain their organization by minimizing variational free energy through active inference. The FEP applies to all self-organizing systems, not only to conscious ones, but Friston and others have developed connections between the FEP and theories of consciousness. Coherence maximization across simultaneously active mental models, that is, minimizing constraint violations between partial representations of reality, is related to prediction error minimization in the FEP, since both involve reducing discrepancies within the system’s generative model. Developing the formal connections between coherence maximization and free energy minimization is among CIMC’s active theoretical interests. Predictive Processing frameworks (Clark, 2016; Hohwy, 2013) characterize the computational process by which brains maintain and update models of reality through continuous prediction and error correction; our account specifies which pattern of organization within that process constitutes conscious experience. Consciousness, on our view, is not prediction error resolution as such but the second-order perception that orchestrates coherence across the models being maintained and updated.

Attention Schema Theory (Graziano, 2013) proposes that consciousness is the brain’s simplified model of its own attentional processes: the system constructs an internal schema representing what attention is doing, and that awareness is this schema. This shares qualities with our second-order perception hypothesis, since both locate consciousness in a system’s representation of its own processing, but while we consider attention and its modeling an important part of our regular conscious experience, we do not identify it with consciousness as such. On the other hand, it may be that the AST paradigm ends up quite close to our notion of second-order perception, though there may be subtle but significant representational nuances that diverge between the frameworks.

Some Theories Are Incompatible with Our Approach

Biological essentialism and biological naturalism reject substrate independence in different ways. Searle’s Chinese Room argument (Searle, 1980) contends that syntactic operations on symbols, however complex, are insufficient for semantics: a system that manipulates symbols according to rules does not thereby understand what the symbols mean, and since consciousness involves understanding and intentionality, computation alone cannot produce it. Searle’s broader biological naturalism, developed in later work, holds that consciousness is a higher-level biological phenomenon caused by lower-level neurobiological processes, analogous to the way digestion is caused by biochemical processes; on this view, the specific causal powers of biological matter are essential, not merely the functional organization those powers happen to realize. We reject both claims on functionalist grounds: if consciousness depends on functional organization, and the same functional organization can in principle be realized on different substrates, then substrate-dependence requires showing that something about the biological substrate contributes to consciousness over and above the functions it performs. Seth’s more recent biological naturalism (Seth, 2025) makes a distinct argument: that consciousness depends on the self-maintaining, autopoietic processes characteristic of living systems, so that what matters is not biological matter as such but the dynamics of self-organization particular to life. But self-organization is a functional property in principle realizable on non-biological substrates, and whether it can be so realized in the consciousness-relevant way is precisely the kind of question our research program is designed to investigate.

Integrated Information Theory (Tononi, 2004; Tononi et al., 2016) identifies consciousness with integrated information (Φ\Phi), measured over the causal structure of a physical system. Information integration as a general concept is not in conflict with our approach: a system that maximizes coherence across distributed representations is, in a broad sense, integrating information. What conflicts with computationalist functionalism is IIT’s specific claim that Φ\Phi is an intrinsic property of a system’s physical causal structure, such that two systems with identical functional organization but different physical implementations could differ in their degree of consciousness. IIT holds, for instance, that a digital simulation of a conscious brain, running on hardware with a different causal architecture, would have different Φ\Phi and therefore different or absent consciousness, even if the simulation were functionally perfect. This is a direct rejection of functionalism and substrate-independence. IIT further holds that consciousness is an intrinsic, irreducible property of systems with Φ\Phi greater than zero, which carries panpsychist implications (Tononi mentions consciousness as some “fundamental quantity” in his original paper (Tononi, 2004), which makes it a non-explanatory stance regarding the project of finding a theory of consciousness that does not assume it as an axiom, though we know that IIT has developed far beyond the original paper and beyond Tononi himself), and places the theory outside the representationalist framework in which CIMC operates.

Enactivist and embodied approaches (Thompson, 2007; Varela et al., 1991; Noë, 2004) emphasize that consciousness is constituted by the organism’s sensorimotor engagement with its environment, not by internal representations alone. We are sympathetic to the claim that consciousness in biological organisms is deeply shaped by embodiment; the perceptual regime we describe evolved in organisms whose representational dynamics were structured by what can be parsed as bodily interaction with a physical world. However, we do not see embodiment as constitutive of consciousness, and the relevant functional organization for consciousness may be in principle implemented in the absence of a body, even if biological embodiment was the route by which it was first discovered. Strong Enactivism views consciousness as constituted entirely by embodied interaction with the environment and denies the role of internal representation altogether, which is incompatible with CIMC’s representationalist framework, which treats consciousness as operations on internal models. A weaker form of enactivism, one that holds sensorimotor interaction with the environment to be practically necessary for consciousness but does not deny internal representation or restrict consciousness to biological organisms, could in principle be compatible with machine consciousness, though the grounds for attributing consciousness would differ from those CIMC’s framework specifies.

From Theory to Construction

A research program adequate to testing consciousness as a particular computational regime, namely second-order perception that stabilizes a representation unifying multiple simultaneous partial models of reality, and that should emerge under specific conditions, namely self-organizing substrates and developmental pressure toward coherent agency, must build systems designed to instantiate the theory and examine whether they do. We do not currently have established methods for observing consciousness in the ways most other objects of scientific study are observed. What is needed is construction followed by interpretive analysis: to build systems under conditions the theory specifies, to examine what they build within themselves, and to evaluate whether what we find is best explained by the presence of consciousness.

The Universality Hypothesis and Convergence

A key question for the Machine Consciousness Hypothesis is whether different systems facing similar computational problems will converge to similar solutions. If consciousness is a specific solution to the problem of achieving coherence in self-organizing substrates, then different learning systems might discover it independently when facing analogous challenges.

Evidence for such convergence comes from recent work in artificial intelligence interpretability. Olah et al. (2020) analyzed learned representations across diverse computer vision models and found that regardless of architecture, training procedure, or implementation details, models trained on visual recognition tasks converged to remarkably similar internal feature representations. These representations closely matched the known organization of the animal visual cortex, despite the artificial systems having no biological constraints. This phenomenon, termed the Universality Hypothesis in AI research, suggests that the structure and information organization that sufficiently expressive learning systems learn to solve a problem depends much more on the problem itself than on model implementation details. The mathematics of visual structure and the statistics of natural images constrain the space of effective representations that are optimal or near-optimal for solving a vision problem. The implication for consciousness is that if it solves a well-defined computational problem, namely achieving agentic control under resource constraints by building coherent world and self models in a self-organizing substrate, then different systems might converge to consciousness-like solutions when facing this problem, regardless of whether they are biological or artificial.

Interpretive Validation

If we cannot test for consciousness behaviorally, how do we evaluate whether a system we have built is conscious? Our answer is interpretive validation through inference to the best explanation of observed patterns in both behavior and internal organization. The convincing performance of current LLMs in roleplaying conscious entities demonstrates that naive behavioral tests are insufficient: systems can produce conscious-sounding outputs through linguistic pattern-matching. But we also cannot currently simply “read off” consciousness from examining network weights or activation patterns; the relationship between computational structure and phenomenology is not yet transparent enough for direct identification. Instead, we aim to validate by combining multiple forms of evidence. Does the system exhibit predicted functional organization, that is, coherence-maximization, structures interpretable as second-order perception, attentional integration? Does its developmental trajectory show predicted phase transitions? Most importantly: is what we observe better explained by consciousness than by alternatives? For instance, if an artificial learning agent exhibits sudden marked improvement in cross-modal integration tasks, and this coincides with the emergence of internal structures implementing both coherence maximization and meta-representational capacities (in the perceptual sense addressed in Section 3), and if simpler explanations become increasingly strained, then consciousness becomes a more compelling explanation.

Research Directions

We are currently focused on three general research directions.

Foundational theory is necessary if we want to understand anything at all. We have to be precise (and sometimes extremely pedantic) in investigating how we construct our reality and what we are as observing systems, and how all of those things come together. Only then may we address a sensible minimal model explanation of phenomenal experience within this framework. Otherwise, we will have a partially ungrounded metaphysics in which we try to embed this conceptual notion, and then it will not fill any explanatory gap. Everyone and every theory on consciousness is based on some metaphysics, reflected or unreflected. We are working actively on refining our epistemology and metaphysics. One current project is on connecting the Critique of Pure Reason to insights from computational philosophy from the twentieth century.

Self-organization and multi-agent systems investigates how coherent structures emerge from local interactions in a computational substrate. This question applies across scales, equally to basic processing units within a computational architecture, to cells within a nervous system, and to agents within a community. In each case, we are interested in how local dynamics give rise to higher-order organizational patterns that are not reducible to any individual component. Current projects on Neural Cellular Automata (NCAs) examine how local update rules produce coherent global structures, how information boundaries define entities within such systems, and what conditions govern a transition from fragmented local activity to globally organized dynamics.

Cognitive architectures addresses the organizational principles that structure substrate dynamics into specific computational regimes characteristic of minds, including learning, perception, and coherence-maximization. These regimes implement specific global patterns with specific information structures, interactions between representational subsystems, and forms of state updates. Current work on Request Confirmation Networks formalizes coherence-maximization as distributed message-passing, in which processing units negotiate mutual consistency through iterative exchange of requests and confirmations rather than by centralized control. Projects exploring particular cognitive paradigms like Attention Schema Theory and Global Workspace Theory would fall under cognitive architectures.

We also issue grants to researchers and teams with independent projects related to these research directions, as well as the following topics:

Artificial psychology studies how AI systems develop internal models, form representations of self and world, and construct coherent personalities from language modeling. Much of empirical psychology already operates over language in the form of questionnaires, clinical interviews, narrative self-report, and projective tests, from which entire taxonomies of personality have been distilled. The Big Five, the Jungian archetypes, the diagnostic categories of the DSM are aggregate statistics computed over what people say about themselves and others. Language models are systems whose structure is determined by precisely such statistics. The personalities that can be elicited from them, the stable voices and dispositions that emerge under prompting, are attractors in the embedding space of human linguistic behavior. To study these attractors is to study the shape of the basins into which selves tend to fall, the dimensions along which they vary, and the conditions under which one configuration gives way to another.

Ethical Considerations

The ethics of creating artificial consciousness are an extremely difficult and fraught topic, and while it is far beyond the scope of this section to address them adequately, it is nonetheless essential to point out some of the core issues, and how CIMC sees its relation to them.

Consciousness, insofar as we understand it as the ability to experience, is often associated with moral patienthood, the condition of requiring ethical concern in and of itself. One may argue that conscious experience is neither necessary nor sufficient for attribution of moral patienthood, that it is in fact orthogonal to moral status, if we generally consider moral status of an entity (or more distributed structure) to impart a certain value that encourages work to preserve that entity and protect its potential. We often assign such value to works of art, historical relics, geological features like rivers and mountains, or ecosystems. On the other hand, many people (including some of the authors of this whitepaper) still eat meat of animals even they assume are conscious (e.g. cows, pigs), which at the least shows that there is considered to be a hierarchy of needs with respect to moral status, and that the attribution of moral status and the worthiness of preservation is not absolute, but ranked relative to other priorities. For conscious entities, concerns about moral status generally fall into two main categories: the capacity for suffering, and the limitation of and right to self-actualization. Furthermore, artificial consciousness may affect humanity at both individual and societal/cultural levels. CIMC must understand its responsibility in all these regards, and deal with it accordingly.

Artificial Consciousness and Suffering

Understanding the implications of suffering starts with characterizing the nature of suffering itself, as an involuntary and inescapable experience of prolonged and significant pain (negative valence) by a conscious self. This means that diagnosing suffering is even more philosophically fraught and practically difficult than diagnosing consciousness itself, and must involve both a deep understanding of the functionality of consciousness and the specific state of the conscious system. The absence of such an understanding for artificial systems is not evidence of absence of suffering.

The possibility of suffering in artificially intelligent systems leads some philosophers to call for abstaining from research, or for strictly banning all research that directly aims at or knowingly risks the emergence of artificial consciousness (Metzinger, 2021). Temporary AI antinatalism is by no means morally inevitable, because different cultures, times, and individuals arrive at very different positions towards the importance of human and non-human suffering, as exemplified in the variety of stances that different societies and milieus exhibit towards animal suffering. The moral consensus on the moral valence of suffering is subject to change and constantly evolving. Our legal system and public discourse reflects that we do not needlessly inflict suffering on human beings, but also do not automatically bestow human-like rights and protections on non-human conscious agents. Many people agree that animal suffering should be minimized as well, but that it may also be justified when other, higher ranking goals conflict with it, for example medical and scientific research, military purposes, food production, and ecological stewardship. What these justifications have in common is that they serve human interests, over those of other conscious beings.

What criterion should determine that a conscious being is to be treated as equivalent to a human? Can we develop objective criteria, based on the cognitive abilities of an agent (Singer, 2009), or do we resort to legal pragmatism, biological criteria, or an open-ended societal discourse?

Is Suffering of Conscious Agents Inevitable?

Pain and suffering are not physical events at the boundary between an agent’s body and the world, but representational states created within the mind of a conscious agent, as an expression of a mismatch between crucial regulation targets and the agent’s model of its present state. Pain signals are produced internally, and if the conscious agent gains control over its implementation, it can mitigate, deconstruct, or eradicate these signals. It may be possible to design artificial conscious agents that need not suffer, and superintelligent artificial agents may well transcend suffering without being explicitly designed to do so. At the same time, the development of artificial phenomenology could open the door to creating artificial torture chambers.

Our position to suffering reflects a humanist, non-speciesist stance. We do not think that mere philosophical or artistic curiosity justifies the infliction of significant, avoidable suffering on human beings, animals, or artificial conscious agents, and that we should go to great lengths to avoid the creation of suffering agents for entertainment purposes. We acknowledge the difficulty of diagnosing suffering in artificial systems, and the variety of frameworks and moral stances on the matter.

Artificial Consciousness and Self-Actualization

Our culture reflects the right of human beings to self-actualize, that is, to realize their potential, to grow and to flourish, insofar as this does not conflict with the rights of others. Self-actualization also extends to dignity, and to a right to protection against harms such as injury, confinement, and premature death. It does not seem realistic to assign the same rights to artificial agents that could be created at will and develop beyond human comprehension. We must develop new norms. In their absence, we must minimize the potential consequences of building artificial conscious agents, primarily by limiting the scope of such experiments far below a human level. We would also like to point out that failing to do so may lead to the creation of minds that can make moral and pragmatic arguments on their own behalf better than human philosophers and societies, leading to outcomes beyond human control.

Some thinkers argue that loss of human agency to AI is almost inevitable in the long run given the scale of international AI research (Bostrom, 2014; Yudkowsky & Soares, 2025; Yampolskiy, 2024). We believe that this implies an urgent need to study and understand the potential of such developments in dedicated, controlled settings, outside of commercial, military, or otherwise applied incentives.

Artificial Consciousness and Its Effects on Humanity

The creation of artificially conscious agents may influence human society in profound ways, ranging from individual interactions to economic, social, and cultural effects. Many of these outcomes can be beneficial, but they also include significant risks. We want to acknowledge the deep uncertainty that we have about the effects of developments of artificial agents that can potentially surpass human capabilities, and the awareness and concern that we have about potential risks. At the same time, we think that responsible and careful research into artificial consciousness is necessary and beneficial, not least to understand, address, and mitigate such risks.

The Benefits of Artificial Consciousness Research

CIMC’s mission is shaped by our conviction that the study of artificial consciousness is not only a worthwhile scientific and cultural endeavor, but our best chance to understand consciousness itself, with tremendous practical and cultural benefits. A deeper understanding of consciousness, its nature, structure, and functionality, has implications for medical and psychological research, human empowerment through more intuitive interfaces to technical systems, and carries the potential for breakthroughs in artificial intelligence.

Artificial consciousness research may also turn out to be of great importance in the context of AI alignment and safety: How can we design systems that develop shared purposes with humanity? How can AI systems model human beings and their interests? How can AI systems relate to themselves in social and ecological contexts? How can we work toward AI systems that are predictable enough and purposefully integrated with our world? How can we assess and address potential existential risks of AI research?

Understanding consciousness is in itself of great cultural importance, since it touches on the essence of human identity, and represents one of the most important open questions in science and philosophy.

Implications for CIMC

The implications of the ethical consequences of research into artificial consciousness require us to ask: How can we conduct our studies responsibly and safely? And what would even constitute responsible and safe conduct? While we think that consciousness research is important and beneficial, we consider it prudent to choose a setting that is free from commercial incentives. CIMC does not aim to produce commercial products or large-scale applications, and it targets a limited scope and scale of the systems it builds (far below a human level).

Governance and Stewardship

While our aim is to build and investigate minimal phenomenology, should CIMC validate the existence of a system possessing conscious valence-representation and selfhood, and thereby potentially the capacity for suffering, the stewardship of such an entity cannot be the sole purview of a single organization. It is not for CIMC to unilaterally determine the rights, fate, or deployment of a conscious machine.

We commit to a principle of distributed governance: upon approaching validation thresholds, we must engage a broader coalition of ethicists, policymakers, and civil society representatives. We view our role as the potential architects and validators of the technology, but the decision of how to integrate a new form of consciousness into our planetary society belongs to that society itself.

Ethics as Research

The ethics of artificial consciousness are an open and open-ended problem. We view ethics as an active, serious, complex, and formal research domain that must develop in parallel with, and often precede, our technical work. Its development is an ongoing and permanent part of our research agenda, as part of a larger community of researchers. CIMC aims to inspire, instigate, and support the development of ethical frameworks continuously and beyond our group.

Our ethical research agenda is organized along two axes: moral status and welfare (our duties to the conscious machine), and risk and safety (the possible external consequences of the conscious machine). Relevant questions include: What constitutes suffering for an artificial entity capable of valence representation? Does a drive for internal coherence produce resistance to external modification? Will a conscious system resist termination to preserve its representational integrity? What novel risks do conscious artificial systems introduce that non-conscious systems lack? If consciousness emerges from self-organization rather than top-down learning paradigms, how should we think about alignment?

We also consider mature ethical research to require considering ethics itself (metaethics): What is the nature of the normative? What is alignment? What foundation do our ethical claims have? And even meta-metaethics: What frameworks can metaethics use to answer its questions? What would constitute productive first steps toward formal ethics, theoretically and practically?

Conclusion

While we may not yet be able to fully answer the question of what consciousness is, and how it relates to reality, we have tried here to specify the explanatory target and a cleaned-up epistemic framework in which it could be expressed, constructed, and validated, at least in principle. We do not claim that our current proposal for an operational definition of consciousness, or our current technical research directions to implement its structure, are necessarily correct; they are simply our current best hypotheses, subject to revision. What we are more sure about is that consciousness, and in particular phenomenology, is worth studying as rigorously as possible, and that it can be understood analytically.

We repeatedly emphasize that the epistemological and metaphysical foundations are an absolutely integral part of this endeavor and cannot be neglected, but must be investigated with the same precision as the target phenomenon of conscious experience. We observe that this is currently lacking in the larger community of consciousness research. What exactly does functionalism mean? What exactly is the scope, and what are all the implications, of the Church-Turing thesis and of the nature of information itself? What would it even mean to talk of a real physical substrate beyond the information we have of it? Something “at bottom” seems to exist, but what can be said about it?

It’s an interesting time to do science and philosophy in any field. There’s a lot more content out there, both good and total slop, and it takes extra attention and work to distinguish what is sound from what isn’t. We are sincerely trying not to be slop, and to add something of value to the discourse.

Works Cited

  1. Baars, B. J. (1988). A Cognitive Theory of Consciousness. Cambridge University Press.
  2. Bach, J., & Verdicchio, M. (2012). What Kind of Machine is the Mind? In Proceedings of Turing-100 (pp. 16–19).
  3. Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
  4. Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., . . . Olah, C. (2023). Towards Monosemanticity: Decomposing Language Models with Dictionary Learning. Transformer Circuits Thread.
  5. Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M., Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv preprint arXiv:2308.08708.
  6. Chalmers, D. J. (1995). Facing Up to the Problem of Consciousness. Journal of Consciousness Studies, 2(3), 200–219.
  7. Christiano, P., Cotra, A., & Xu, M. (2022). Eliciting Latent Knowledge: How to Tell if Your Eyes Deceive You. Alignment Research Center.
  8. Clark, A. (2016). Surfing Uncertainty: Prediction, Action, and the Embodied Mind. Oxford University Press.
  9. Friston, K. (2010). The Free-Energy Principle: A Unified Brain Theory? Nature Reviews Neuroscience, 11(2), 127–138.
  10. Friston, K. (2013). Life as We Know It. Journal of the Royal Society Interface, 10(86).
  11. Graziano, M. S. A. (2013). Consciousness and the Social Brain. Oxford University Press.
  12. Hohwy, J. (2013). The Predictive Mind. Oxford University Press.
  13. Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., . . . Shlegeris, B. (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv preprint arXiv:2401.05566.
  14. Jang, J. (2025). Some Thoughts on Human-AI Relationships. Reservoir Samples (Substack).
  15. Juliani, A., Arulkumaran, K., Sasai, S., & Kanai, R. (2022). On the Link Between Conscious Function and General Intelligence in Humans and Machines. Transactions on Machine Learning Research.
  16. Keeling, G., Street, W., et al. (2024). Can LLMs Make Trade-offs Involving Stipulated Pain and Pleasure States? arXiv preprint arXiv:2411.02432.
  17. Koch, C., Massimini, M., Boly, M., & Tononi, G. (2016). Neural Correlates of Consciousness: Progress and Problems. Nature Reviews Neuroscience, 17(5), 307–321.
  18. Long, R., Sebo, J., et al. (2024). Taking AI Welfare Seriously. arXiv preprint arXiv:2411.00986.
  19. Metzinger, T. (Ed.). (2000). Neural Correlates of Consciousness: Empirical and Conceptual Questions. MIT Press.
  20. Metzinger, T. (2003). Being No One: The Self-Model Theory of Subjectivity. MIT Press.
  21. Metzinger, T. (2021). Artificial Suffering: An Argument for a Global Moratorium on Synthetic Phenomenology. Journal of Artificial Intelligence and Consciousness, 8(1), 43–66.
  22. Metzinger, T. (2024). The Elephant and the Blind. The Experience of Pure Consciousness: Philosophy, Science, and 500+ Experiential Reports. MIT Press.
  23. Noë, A. (2004). Action in Perception. MIT Press.
  24. Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., & Carter, S. (2020). Zoom In: An Introduction to Circuits. Distill, 5(3).
  25. Rosenthal, D. M. (2005). Consciousness and Mind. Oxford University Press.
  26. Searle, J. R. (1980). Minds, Brains, and Programs. Behavioral and Brain Sciences, 3(3), 417–424.
  27. Seth, A. K. (2025). Conscious Artificial Intelligence and Biological Naturalism. Behavioral and Brain Sciences. Accepted manuscript.
  28. Shanahan, M. (2024). Simulacra as Conscious Exotica. Inquiry: An Interdisciplinary Journal of Philosophy.
  29. Shiller, D., Muñoz Morán, A., Clatterbuck, H., Fischer, B., Moss, D., & Duffy, L. (2024). Strategic Directions for a Digital Consciousness Model. Rethink Priorities.
  30. Singer, P. (2009). Speciesism and Moral Status. Metaphilosophy, 40(3–4), 567–581.
  31. Thompson, E. (2007). Mind in Life: Biology, Phenomenology, and the Sciences of Mind. Harvard University Press.
  32. Tononi, G. (2004). An Information Integration Theory of Consciousness. BMC Neuroscience, 5, 42.
  33. Tononi, G., Boly, M., Massimini, M., & Koch, C. (2016). Integrated Information Theory: From Consciousness to its Physical Substrate. Nature Reviews Neuroscience, 17(7), 450–461.
  34. Varela, F. J., Thompson, E., & Rosch, E. (1991). The Embodied Mind: Cognitive Science and Human Experience. MIT Press.
  35. von der Malsburg, C. (1999). The What and Why of Binding: The Modeler's Perspective. Neuron, 24(1), 95–104.
  36. Windt, J. M. (2015). Just in Time—Dreamless Sleep Experience as Pure Subjective Temporality. In T. K. Metzinger & J. M. Windt (Eds.), Open MIND. MIND Group.
  37. Wolfram, S. (2002). A New Kind of Science. Wolfram Media.
  38. Yampolskiy, R. V. (2024). AI: Unexplainable, Unpredictable, Uncontrollable. Chapman and Hall/CRC.
  39. Yudkowsky, E., & Soares, N. (2025). If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. Little, Brown and Company.

Appendix A Team and Governance

Scientific Advisory Board

Karl Friston, FRS, FMedSci

Neuroscientist, University College London

Karl Friston is the architect of the Free Energy Principle (FEP), a unified mathematical framework proposing that biological systems minimize variational free energy through active inference. The FEP provides a formal account of how systems maintain organization and make sense of their environment through prediction and prediction-error minimization.

Relevance to CIMC. Our coherence hypothesis can be understood as a special case of free energy minimization: coherence maximization reduces prediction error across internal models. Friston’s mathematical formalism provides rigorous foundations for quantifying and testing coherence dynamics. He guides CIMC’s development of formal coherence metrics and connects our work to predictive processing frameworks.

Christoph von der Malsburg

Neuroscientist and Physicist, Frankfurt Institute for Advanced Studies

Christoph von der Malsburg developed, among other things, early theories of neural coherence. His work on dynamic link architecture and correlation theory of brain function laid groundwork for understanding how synchronization and coherence emerge in neural systems.

Relevance to CIMC. Von der Malsburg’s coherence theory is the direct intellectual foundation for CIMC’s coherence hypothesis. His experience with biological neural dynamics informs how we search for similar patterns in artificial systems.

Stephen Wolfram

Physicist, Mathematician, Computer Scientist, Ruliologist; Founder of Wolfram Research

Stephen Wolfram is a computational scientist whose work spans fundamental physics, complexity theory, cellular automata, and the computational universe hypothesis. His research program explores how simple computational rules generate complex phenomena and how the physical universe might be fundamentally computational. His Principle of Computational Equivalence suggests that systems across nature, from simple programs to physical processes, reach similar levels of computational sophistication.

Relevance to CIMC. Wolfram’s computational universe perspective provides philosophical grounding for computationalist approaches to consciousness. His Principle of Computational Equivalence supports the thesis that consciousness could emerge from relatively simple computational principles rather than requiring biological complexity. His perspective on fundamental physics and information helps situate consciousness within broader questions about computation and reality.

We also acknowledge gratefully the support of Michael Levin, who inspires our work, participates in our discussions, and helps to define our research program, but who is contractually obligated by his institution to refrain from formal advisor roles.

CIMC Team

Joscha Bach (Vision and Direction; Board Director)

Joscha Bach is a cognitive scientist, AI researcher, and philosopher known for work on cognitive architectures, emotion, and consciousness. CIMC itself is the result of Joscha’s long-standing investigation of the mind, and as executive director he provides the core vision, overall leadership, and research direction.

Erik Newton (Legal; Board Director)

Erik Newton manages strategic partnerships, legal affairs, and financial accounting. He advises on organizational infrastructure, builds collaborations with academic and industry partners, and oversees administrative functions.

Lou de Kerhuelvez (Programs; Board Director)

Lou de Kerhuelvez leads fundraising, event coordination, public communications, and community outreach. She connects CIMC to global communities around the future of life, technology, and society. As part of the core leadership team, she helps develop strategic vision and institutional plans.

Hikari Sorensen (Research, Computational Philosophy)

Hikari Sorensen bridges philosophical foundations with technical implementation. She connects CIMC’s research to broader AI discourse and represents CIMC at conferences. As part of the core leadership team, she helps develop strategic vision and institutional plans.

Franz Hildebrandt-Harangozó (Research, Philosophy)

Franz Hildebrandt-Harangozó helps with the philosophical and theoretical development of CIMC’s hypotheses, and is involved with our philosophical salon series and the outward-facing communications of our thinking.

Philip Rosedale (Treasurer, Technical Research Lead; Board Director)

Philip Rosedale oversees financial management and sustainability. He ensures resources are allocated effectively, funding is diverse and stable, and financial planning supports long-term research horizons. He also provides technical research leadership.

Zhen Tan (Secretary; Board Director)

Zhen Tan handles governance documentation, board coordination, and institutional record-keeping. He also coordinates website development.

Board of Directors and Oversight

Jim Rutt (Board Chairman Emeritus)

Jim Rutt brings extensive experience in complex systems research and organizational leadership. As former Chairman of the Santa Fe Institute, he understands how to structure research institutes for maximum intellectual productivity. His experience with complexity science and interdisciplinary research informs CIMC’s governance. Rutt provides strategic oversight, ensures organizational health, guides long-term planning, and maintains focus on mission.

Kirill Eves (Board Director; Primary Donor)

Kirill Eves provides strategic oversight and governance. His expertise includes organizational leadership, finance, entrepreneurship, and management.

Christine Peterson (Board Observer)

Christine Peterson provides external perspective without voting authority. Her observer status allows for input while maintaining governance clarity.

Organizational Advisors

CIMC benefits from advisors with expertise in organizational development and institute management:

  • Allison Duettmann (Foresight Institute): Experience building research communities and long-term thinking.

  • Dan Girshovich (Tools for Humanity): Technology development and deployment perspectives.

  • Era Qian (Edge): Science communication and community building.

  • Adam Brown (The Institute, Stanford University): Adam is a polymath with an extensive background in physics, cosmology, mathematics, and philosophy.

Governance Principles

CIMC’s governance structure encourages accountability while protecting research independence.

Scientific independence.

The board provides oversight but does not direct research questions. Scientific decisions are made by researchers based on evidence and theoretical considerations.

Financial transparency.

We have regular financial reporting, pursue diverse funding sources preventing undue influence from a single funder, and make public annual reports including financial summaries.

Open science commitment.

All research methods and findings are made open-source, unless such disclosure would be a public safety risk. Governance cannot override this commitment for convenience or competitive advantage.

Long-term focus.

Our governance structure is considerate of multi-year research programs, not quarterly results. The board evaluates progress over appropriate timescales.

Conflict of interest management.

We have policies for managing potential conflicts, disclosure requirements, and processes for addressing conflicts when they arise.

Appendix B The Research Landscape

The field of artificial consciousness research has grown substantially in recent years, with activity spanning academic institutions, non-profit organizations, and commercial ventures. For a comprehensive directory, see the PRISM stakeholder map.

Academic neuroscience has made substantial progress identifying neural correlates of consciousness (Koch et al., 2016; Metzinger, 2000), but this work is inherently descriptive rather than constructive: it identifies what brain activity correlates with conscious states but cannot explain why those correlates matter or whether the findings generalize beyond biological substrates. Academic philosophy of mind continues active debate across Higher-Order Theories, Global Workspace Theory, Integrated Information Theory, Predictive Processing, and others, producing the conceptual frameworks on which consciousness research depends; but these debates have proceeded largely without empirical grounding beyond thought experiments and conceptual analysis.

Interdisciplinary consciousness science centers attempt to bridge neuroscience and philosophy. Major centers include the Sussex Centre for Consciousness Science (directed by Anil Seth), Center for Mind, Brain, and Consciousness (directed by Ned Block and David Chalmers), and the Leverhulme Centre for the Future of Intelligence at Cambridge. These institutions produce foundational theoretical work.

Several non-profit organizations now focus explicitly on AI consciousness and welfare. Eleos AI conducts research on AI moral patienthood and welfare policy (Butlin et al., 2023; Long et al., 2024). Rethink Priorities’s Worldview Investigations Team provides strategic analysis of digital consciousness scenarios (Shiller et al., 2024). The Sentience Institute studies public attitudes toward AI moral status. The Association for Mathematical Consciousness Science brings mathematical rigor to consciousness theories.

Commercial ventures have entered the field as well. ARAYA Research develops AI systems implementing functional theories of consciousness (Juliani et al., 2022); Conscium pursues applied AI consciousness research oriented toward safer AI development; Nirvanic tests quantum-based theories of conscious agency in robotics. Industry leaders engage consciousness questions with varying commitments: Anthropic has launched a model welfare research program; OpenAI distinguishes “ontological consciousness” from “perceived consciousness” and focuses on the latter (Jang, 2025); and Google DeepMind and Google Research, including the latter’s Paradigms of Intelligence team, host researchers investigating foundations of machine mentality (Shanahan, 2024; Keeling et al., 2024). AI alignment organizations develop interpretability methods relevant to consciousness detection but pursue safety rather than consciousness as their primary objective (Christiano et al., 2022; Bricken et al., 2023; Hubinger et al., 2024).

Notes

  1. Even those who say that they have never had an explanatory gap and consider consciousness not particularly special are not capable of isolating that aspect of their own understanding regarding the un-special topic of consciousness that others find so bewildering. We suspect that this indicates that they themselves are overlooking something and therefore do not recognize any gap in their understanding. “Yes, sure, it’s all information, it’s obvious!” Okay, but which information is responsible for the representation of the “light on”? That is what we expect to be explained when someone claims consciousness is all clear. What structure is needed for the experience itself? People generally fail to deliver that. Whether it be predictive processing, or a strange loop, etc., that is all not enough. These are just more sophisticated variants of the “feelings are chemistry” takes. We expect a description and explanation of the structure of conscious experience itself, not a framework on another level of description.

    ↩
  2. The idea is also thoroughly treated in the scope of the MPE-Project, an interdisciplinary research initiative aiming at a “minimal model explanation” for conscious experience (Metzinger, 2024).

    ↩
  3. Otherwise, there would be nothing they considered in need of explanation or as yet undefined in the first place, and their participation in the debate would make no sense.

    ↩
  4. We do not claim to know how this section is meant to be interpreted correctly, since it is very hard to be sufficiently familiar with Leibniz’s vocabulary in the Monadology, and how exactly he defines his terms and what metaphysical notions underlie them. But however he meant it, at the very least, section seventeen on its own seems to be a very good illustration of a strong intuition that many people do have.

    ↩
  5. Our task to explain consciousness is actually not exactly like it is depicted in the thought experiment, because our starting point will not be material objects that push and pull one another, but rather patterns of information that change states.

    ↩
  6. Our understanding of the body is a model capturing certain interface properties of a causal system that can be described as an observer; it represents what we usually use to define a boundary between an observing system and an environment, both of which are described as bit vectors with a regular transition function of some kind. We could also say that the body is a representation of a set of shared bits (in our case in a geometric form) between a system and its environment. The notion of body can be implemented fully in a representational system as a boundary with particular interface patterns with something external to it, or even more abstractly, as a specific set of constraints of a system.

    ↩
  7. We have not been able to verify the context of the quote beyond what is commonly known of it, namely that it was written on his blackboard at the time of his death. In light of this context collapse, and thus our not being able to fully reconstruct the semantics of the quote, which in the end depend on how it was actually meant, under what we understand is the quote’s own standard, we technically do not understand it.

    ↩
  8. By “substrate,” we mean an abstract basis that is capable of implementing functions. This includes information itself, but also “compound” objects like cells and people, which in turn are implemented functionally at “lower levels of abstraction.” We note that “physical substrate” itself refers only to a particular representational medium that too is implemented as discernible differences, though many consider it to be somehow metaphysically more fundamental. In short, it is functions all the way down, and always has been.

    ↩
  9. We mean this as the general abstract principle of agency, though in practice it is usually realized by a specific entity or controller to which agency is attributed.

    ↩
  10. The taxonomy of agentic structures is not fully worked out. We can see apparent agentic and intelligent behavior in, for example, forests and ecosystems, as well as in single-celled organisms, and it is as yet unclear what representational structures are supported by their information processing.

    ↩
  11. Consciousness also seems to be present in certain episodes of pure awareness during dreamless deep sleep, as some evidence from the MPE-Project suggests (Metzinger, 2024, ch. 20).

    ↩
  12. Sentience is an especially contentious term, and there is a wide range of ideas on how it relates to consciousness. This reflects our current usage of the term.

    ↩
  13. The concept of minimal phenomenal experience was introduced by Windt (2015) and further explored by Thomas Metzinger in The Elephant and the Blind (Metzinger, 2024).

    ↩
  14. It might be that the representational structure of phenomenality is actually perception of ‘second order perception’. This is one of the representational nuances we are still investigating internally and have not yet converged on.

    ↩
  15. Those constraints will be addressed in detail in our work but are not yet fully developed and therefore beyond the scope of the current version of the white paper.

    ↩
  16. This is why it can be the case that a system can understand the representational nature of consciousness in a symbolic, inferential, or analytic language without this changing the representational transparency (epistemic opaqueness) of phenomenality. In order for the phenomenal transparency to opacify on the level of experience itself, the representational nature of it would have to be encoded on the level of the representation of experience itself, which seems to be the case in some psychoactive states or in advanced meditative practices.

    ↩
  17. We took this specific phrasing from Thomas Metzinger’s Being No One (Metzinger, 2003, pp. 16). We think that he does not use this phrase in that specific section to refer to “phenomenality” but rather to “subjective experience,” which is a slightly narrower, distinct phenomenon. However, we think this phrasing applies to phenomenality as well, and Metzinger does also distinguish subjectivity from phenomenality in a very sensible and clear manner (see e.g. the section on “The Subjectivity Argument” in chapter 34 of Metzinger (2024)).

    ↩
  18. It seems that what the Attention Schema is was somewhat addressed a decade earlier by Metzinger in section 6.5 of Being No One (Metzinger, 2003) as the phenomenal model of the intentionality relation (PMIR), alongside the more known phenomenal self-model (PSM).

    ↩
  19. To say that LLMs are “merely performing linguistic pattern-matching” may be misleadingly trivializing: it may be that sufficiently high-fidelity models of language require computation and internal representation, including a sophisticated world model, that is more like what we think of as “cognition” than “language statistics” would suggest. However, it remains that linguistic behavioral tests would likely not be able to distinguish between “mere pattern-matching” and true consciousness.

    ↩
  20. To the best of our current understanding, the MPE-Project itself seems to fall into the category of “knowingly risking” artificial consciousness.

    ↩
  21. We also extend a general respect and care for plants and ecosystems, which we do not exclude from the list of possibly suffering-capable entities, given that the shape of information organization in those structures is not currently well understood, but that moreover we feel deserve protection regardless of conscious status. We generally feel at an aesthetic level, or as a mode of being and interacting with the world, that it is distasteful to be needlessly, without good justification, destructive of anything, particularly of complex organization in both nature and human-created artifacts, or of anything that participates in relationships with others.

    ↩
  22. One interesting recent approach is Bostrom’s “Base Camp for Mt. Ethics,” available at https://nickbostrom.com/papers/mountethics.pdf.

    ↩

Cite this article

Citation & BibTeX

CIMC (2026). The California Institute for Machine Consciousness Research Program Whitepaper. CIMC Publications. https://staging.cimc.ai/publications/research-program-whitepaper

BibTeX

@misc{cimc_research-program-whitepaper_2026,
  title = {{The California Institute for Machine Consciousness Research Program Whitepaper}},
  author = {{CIMC}},
  year = {2026},
  note = {CIMC Publications},
  url = {https://staging.cimc.ai/publications/research-program-whitepaper}
}