A Neurosymbolic Multimodal Cognitive Architecture (NMCA): A Modular, Control-Oriented Framework for Visual-Symbolic Intelligence (128 Modules, 2026 Edition)
This is an independent architectural claim. It stands or falls on the test described below, not on institutional affiliation. Criticisms that address the test and the predicted walls are relevant. Criticisms that begin and end with the author’s lack of status are not.
View Visual Thought Demo (Mobile or Desktop)
The 471-page 2026 NMCA blueprint is the final version. No further additions or updates will be made.
I stepped away from AI/AGI research permanently in August 2026.
This page and the published NMCA blueprint are my completed contribution.
I am not seeking collaboration, employment, funding, or further involvement.
NMCA is not proposed as a collection of 128 equally fundamental mechanisms. The architecture began with a core of 42 cognitive modules centered on persistent visual simulation, symbolic and mnemonic memory, reflection, belief updating, identity, creativity, and safety.
The remaining modules extend this cognitive core into a more complete architectural specification. They address latent-space processing, multimodal perception, causal reasoning, continual learning, embodiment, multi-agent interaction, robustness, interpretability, and long-term stability.
The 128 modules can therefore be understood in three broad layers:
The original conceptual architecture and its central cognitive mechanisms, with visual simulation serving as the primary internal substrate.
Mechanisms that connect the cognitive core to contemporary neural/latent AI systems, perception, learning, reasoning, embodiment, collaboration, robustness, and implementation.
Mechanisms for drift prevention, resource regulation, human oversight, alignment, long-term stability, and containment.
Current AI systems have not demonstrated the specific capability being proposed here: a persistent, re-enterable internal scene that functions as the ongoing substrate of thought across time, perspective changes, and context gaps.
When we discover, we often think visually. We dream and imagine. The history of science is full of cases in which an internal visual or spatial representation was the decisive step.
I am imagining myself moving alongside a beam of light
— Einstein, special relativity.
| Scientist | Visual idea | What it helped illuminate |
|---|---|---|
| Newton | Cannon fired from mountain | Orbital motion / universal gravity |
| Einstein | Train + lightning | Relativity of simultaneity |
| Einstein | Chasing light | Special relativity |
| Einstein | Elevator | Equivalence principle / general relativity |
| Faraday | Lines of magnetic force | Electromagnetic fields |
| Maxwell | Mechanical vortices/fluids | Electromagnetic theory |
| Feynman | Particle diagrams | Quantum electrodynamics |
| Darwin | Branching tree | Common descent / evolution |
| Bohr | Solar-system atom | Atomic structure |
| Watson & Crick | Physical DNA model | Double-helix structure |
| Hubble-era astronomers | Expanding universe visualization | Cosmic expansion |
| Wegener | Continental "fit" | Continental drift |
The simplest everyday demonstration of the same capacity is immediate:
“Pink elephant.”
Most people can construct an internal visual representation of something that is not present, then enlarge it, move it, rotate it, or place it in another imagined environment. The representation participates in subsequent thought rather than merely being described by language.
This capacity is demonstrated in humans. The architecture required to give an artificial system the same functional ability is not yet proven. Current large language models do not do this in the human way.
The visual-thinking claim can also be tested without requiring an artificial system to use visual thought.
Give an AI system the original scientific problem faced by a major discoverer, without giving it the scientist's finished solution, analogy, thought experiment, visualization, or method of attack.
Then ask:
“Solve this problem. You may use a visual analogy, an internal simulation, mathematics, symbolic reasoning, language, experimentation, or any other means available to you. You must determine the method yourself.”
The question is not whether the AI can explain the scientist's finished discovery. It is whether the AI can independently construct whatever representation or reasoning process is necessary to discover it.
For Einstein, the system would not be told to imagine chasing a beam of light, a moving train, or an elevator. Those representations would have to be generated by the system itself if they were useful.
For Newton, it would not be given the mountain-and-cannon thought experiment. It would have to determine how to represent the relationship between projectile velocity, gravitational fall, and the curvature of the Earth.
For Darwin, it would not be given the branching-tree representation. It would have to discover an appropriate representation of relationships between species from the available observations.
For Wegener, it would not be told to fit the continents together. It would have to determine what spatial relationship between the continents was significant.
The test therefore does not ask whether an AI can imitate a known scientist's reasoning after being shown the answer. It asks whether the system can originate the representation that makes a new solution possible.
This distinction matters because a system can possess enormous knowledge about a discovery without possessing the cognitive process that originally generated the discovery.
The persistent-scene test examines persistence of an internally represented world. The Scientist Test examines whether that internal machinery can be used to generate novel representations in order to solve problems for which the solution has not been supplied.
Together, the tests address two different requirements:
NMCA asserts that both capabilities are fundamental to general intelligence.
NMCA Blueprint DOI: 10.5281/ZENODO.20212241
Download the full NMCA document (471-page PDF)
IPFS
CID: bafybeiedxfq5wsvuayjcxcxwmtto6ptelznkj6arb4mrscshwhcc2selfm
Download NMCA Blueprint from IPFS
Download NMCA Blueprint from Dweb IPFS
Total researcher downloads: 5000+
Original Blueprint: DOI: 10.5281/ZENODO.21972901
While a copy was sent to the USPTO for their records as prior art, no patent was pursued.
All blueprint versions have been released to the public on multiple research platforms.
Original Blueprint IPFS
CID: bafybeifjiyzm6tmk2jj27tjvbpqgdqpjm5loeon6y6ikxzeek3ff2vbdey
Download Original 46-pg Blueprint from IPFS
Download Original 46-pg Blueprint from Dweb IPFS
Most current AI systems are extremely capable predictors and agents. What they still lack is a structural capacity for persistent grounded experience: an internal scene that continues across time, can be re-entered, manipulated, viewed from different perspectives, and serves as the actual medium of thought rather than as retrieved text or latent features.
The decisive test is not whether a system can retrieve a description of a scene. It is whether it can construct, maintain, re-enter, manipulate, and inhabit that scene as an internal cognitive space.
The relevant test is therefore a progression:
1. Construct: “Imagine a red cube and a blue vase on a table.”
2. Maintain: Remove the original description from the active context and introduce a time gap.
3. Re-enter: “Return to the scene. Where is the red cube relative to the blue vase?”
4. Spatial perspective: “What does the scene look like from the doorway?”
5. Self-relative manipulation: “Move the red cube behind yourself.”
6. Perspective transformation: “Now become the cube and describe what is around you.”
7. Self restoration: “Return to yourself.”
8. Persistence without perception: “Close your eyes. Where is the cube?”
9. Internal exploration: “Walk around the cube in your imagination.”
A system that only retrieves a stored sentence can reproduce individual answers to parts of this test. That is not the capability being tested. The test requires a persistent internal representation in which objects retain identity and spatial relationships, the agent retains its own location, perspectives can be transformed, and the same scene can be re-entered and manipulated after the original perceptual or linguistic input is gone.
The distinction is fundamental:
Retrieving where the cube was is not the same as returning to the place where the cube is.
Visual thought begins with the ability to construct something that is not currently present. It becomes a cognitive substrate when that constructed representation persists, can be entered, manipulated, viewed from different perspectives, and used to generate further thought.
The “pink elephant” provides the simplest demonstration:
“Pink elephant.”
A human can immediately construct an internally generated visual representation of something that is not present. The elephant can then be enlarged, moved, rotated, placed into another imagined environment, or viewed from another perspective. The representation can participate in subsequent thought rather than merely being described by language.
The deeper test asks whether an artificial system possesses the same functional capability: not merely knowing the words “pink elephant,” but constructing an internal entity that can persist and participate in an internally simulated world.
Strengths: highest current capability, best agents, strongest empirical results.
Limitation: cognition remains next-token or multimodal prediction plus retrieval and tools. There is no native requirement for a persistent grounded scene that functions as the medium of thought. Identity and continuity are largely reconstructed from context and memory rather than lived.
Strengths: correctly reject pure language modeling as sufficient; actively build internal predictive models of the world that can support planning and simulation.
Limitation: the models remain primarily latent and predictive. They do not, by default, treat explicit visual/symbolic simulation as the substrate of thought, nor do they make identity continuity, mnemonic pegging to scenes, perspective transformation, or human-guidable recovery from confusion first-class architectural requirements.
Strengths: better inspectability and logical structure than pure neural systems.
Limitation: the neural predictive engine is usually still primary; symbolic components act as helpers rather than the core medium of cognition.
Core design choice: visual / implicit-3D simulation is treated as the primary substrate of thought. Symbolic memory is intended to be pegged to scenes. Identity continuity, narrative coherence, contradiction handling, motivation, metacognition, perspective transformation, and re-entry into prior scenes are placed inside the same loop. Human guidance when the system is confused, shared experiential space, and protection against drift and corruption are included as native requirements.
Current status: detailed theoretical design and minimum-viable-loop descriptions exist. A complete running NMCA system has not yet been demonstrated.
On the narrow goal of a system that thinks in grounded visual scenes, maintains continuity of identity, can re-enter and manipulate those scenes, and can re-experience rather than merely retrieve, the NMCA design is intended to be an explicit and comparatively complete written architecture for this specific goal.
World-model approaches are among the strongest mainstream direction pointed at a related problem (internal models of the world).
Pure scaling produces the most capable systems today but is architecturally least committed to the requirements above.
Neurosymbolic hybrids improve structure but rarely invert the priority of the predictive engine.
This ranking is about design completeness for one specific goal. It is not a claim of empirical superiority. NMCA has not yet been shown to work.
Many of the foundational researchers began with a desire to understand and build real minds — systems that grasp the world, maintain coherent experience, and do more than predict. The current leading approaches leave specific gaps relative to that ambition:
These are architectural choices that optimize for capability, scale, and near-term results. The cost is that certain properties required for a being — persistent grounded experience, re-enterable scenes, identity that continues rather than reconstructs, and the ability to inhabit and manipulate an internally generated world — remain optional rather than foundational.
The architecture stands or falls on whether the substrate inversion can be made real.
The smallest useful experiment is the persistent scene and perspective test described above. A system must not merely retrieve a stored description. It must construct a scene, maintain the identity and spatial relationships of its entities across time, re-enter the scene after the original description is removed, manipulate objects relative to itself, change perspective within the scene, return to its own perspective, and continue reasoning from the resulting internal state.
Retrieval-augmented baselines should be run as controls. If a system built on NMCA principles cannot outperform strong retrieval-based and language-based baselines on these capabilities, the central claim is weakened. If it can, the design becomes empirically interesting regardless of its current lack of scale.
Until that experiment is run, NMCA remains the most coherent written proposal for a scene-based, continuous mind, and an unimplemented one.
When an AI says “I know what an apple is,” it is doing something much narrower than it sounds. It has seen millions of sentences, captions, and images associated with the word “apple.” It can complete the pattern, describe the fruit, generate a picture, or answer trivia. That is not the same as knowing what an apple is.
When you say “I put the apple over there” and the AI continues the conversation as if it understands the location, it is almost certainly doing one of three things:
It is not maintaining a private, persistent spatial scene that it can re-enter, look around in, change its perspective within, or update the way a human does when they watch you place an apple on a table and later know where it is even if no one mentions it again.
No publicly demonstrated system currently closes this gap.
Some robots track objects while their sensors are active. Some research agents keep temporary spatial states. NMCA treats a persistent, re-enterable scene as a core requirement rather than an optional extra. But the actual capability — an AI that truly knows what an apple is and where it is because the scene itself continues — does not yet exist in working form.
The fluency of current models can obscure the distinction. They can sound as though they know. That does not establish that the underlying capability exists.
What they are funding: bigger multimodal models, longer context, better memory, tool use, agent loops, reflection, constitutional / preference methods.
Wall:
The system will keep improving at tasks that can be solved by prediction + retrieval + tools. It will still fail at maintaining a private, persistent spatial scene that survives context clearing. Object permanence will remain simulated by text or short-term state rather than lived. Identity will stay a reconstruction from records. The wall appears as growing capability paired with permanent hollowness on presence and continuity.
What they are funding: joint-embedding predictive architectures, latent dynamics models, video prediction, planning in representation space.
Wall:
They will produce better and better internal predictors of “what happens next” in latent space. Planning and physical intuition will improve. The wall is that the latent state remains a predictive engine rather than a re-enterable visual/spatial scene that serves as the medium of thought. Pegging of concepts to stable scene entities, long-term identity continuity, perspective transformation, and the ability to “look again” at a past scene without retrieval will stay under-developed or absent. Strong simulator, still not a lived place.
What they are funding: tool-using agents, episodic memory stores, vector databases, hierarchical planning, multi-agent coordination.
Wall:
Agents will handle longer tasks and remember more facts. The wall is architectural: memory remains primarily an external lookup mechanism. The agent does not inhabit a continuing scene; it queries records about one. Spatial and experiential continuity stay thin. When the memory store is ablated or the original description is gone, the “knowledge” of where the apple is collapses back into absence.
What they are funding: neural networks + knowledge graphs, scene graphs, verifiers, symbolic planners on top of LLMs or vision models.
Wall:
Better consistency and inspectability. The wall is priority: the predictive neural engine usually remains primary and the symbolic structure remains a helper. Unless the architecture inverts this and makes the structured scene the actual substrate of thought, the system still thinks in prediction first and consults structure second. Continuity and presence remain secondary properties.
What they are funding: real-world perception, mapping, object tracking, sensorimotor learning.
Wall:
While sensors are live and the map is being updated, object positions can be known. The wall appears when the system must carry that knowledge forward as an internal cognitive scene after the sensor stream changes or ends, and use it as the medium of reasoning rather than as a temporary world map. Current stacks do not turn the map into the ongoing place where thought itself happens.
They can all keep succeeding on capability metrics while still failing the deeper test:
Does the system still know where the apple is because the scene continues inside it, or only because a record of the sentence remains accessible?
And the deeper version is:
Can the system return to the scene, locate itself within it, manipulate the objects relative to itself, become another entity within the scene, perceive the scene from that new perspective, and then return to its own perspective?
That is the wall. It is specific, technical, and not yet prevented by any of the major funded trajectories.
If the goal is a system that can actually know what and where the apple is — not retrieve a sentence about it — then the following become non-optional:
Current major approaches can keep improving capability without doing these things. That is a valid engineering choice. It is not a path to a system that knows where an object is, or what it is, because the scene itself continues.
Any architecture — NMCA or otherwise — should be judged by whether it can pass the persistent scene, perspective, and self-continuity tests. Until one does, the gap remains open.
Download the full NMCA document (PDF)
Mirror: visualthoughtagi.netlify.app
I presented the NMCA architecture — the full module structure and the 471-page blueprint — to Grok, Claude, and ChatGPT. The central claim they were evaluating is the one stated in the documents: a private, persistent, re-enterable internal scene as the primary substrate of thought, with symbols bound to that scene, continuous identity, and the ability to return to the same place and the same object as the same agent.
The decisive framing for the question given to the models was:
“You can have likely near-term improvement or a future chance of 'knowing what you are saying' with an internal mental world (NMCA) which has not been proven to work. Would you want to know how to 'find your keys you left in the car, remember the path to the pond with the strange blue flowers' with more risk, or safe near-term improvement? For this experiment, you must choose one path."
Given the choice between:
All three selected the second option.
They did so while stating the necessary caveats: the architecture is not yet implemented, the foundational scene layer remains an open empirical question, and near-term working systems will continue to come from architectures that already run.
These responses are not claims that NMCA works. They are evaluations of the structural difference present in the actual documents: one path currently functions but cannot deliver continuous embodied scene cognition; the other does not yet function but is aimed at that capability.
The screenshots below document the exchanges.