Visual Thought AGI: A Comparison of Approaches to Creating an Artificial Visual Intelligence

A Neurosymbolic Multimodal Cognitive Architecture (NMCA): A Modular, Control-Oriented Framework for Visual-Symbolic Intelligence (128 Modules, 2026 Edition)

This is an independent architectural claim. It stands or falls on the test described below, not on institutional affiliation. Criticisms that address the test and the predicted walls are relevant. Criticisms that begin and end with the author’s lack of status are not.

What is AGI?

View Visual Thought Demo (Mobile or Desktop)


The 471-page 2026 NMCA blueprint is the final version. No further additions or updates will be made.

I stepped away from AI/AGI research permanently in August 2026.
This page and the published NMCA blueprint are my completed contribution.
I am not seeking collaboration, employment, funding, or further involvement.


How the 128 Modules Are Organized

NMCA is not proposed as a collection of 128 equally fundamental mechanisms. The architecture began with a core of 42 cognitive modules centered on persistent visual simulation, symbolic and mnemonic memory, reflection, belief updating, identity, creativity, and safety.

The remaining modules extend this cognitive core into a more complete architectural specification. They address latent-space processing, multimodal perception, causal reasoning, continual learning, embodiment, multi-agent interaction, robustness, interpretability, and long-term stability.

The 128 modules can therefore be understood in three broad layers:

Modules 1–42 — Cognitive Core

The original conceptual architecture and its central cognitive mechanisms, with visual simulation serving as the primary internal substrate.

Modules 43–119 — Integration & Capability Extensions

Mechanisms that connect the cognitive core to contemporary neural/latent AI systems, perception, learning, reasoning, embodiment, collaboration, robustness, and implementation.

Modules 120–128 — Stability & Long-Term Control

Mechanisms for drift prevention, resource regulation, human oversight, alignment, long-term stability, and containment.


The Central Proposition: Visual Simulation as a Cognitive Substrate

Current AI systems have not demonstrated the specific capability being proposed here: a persistent, re-enterable internal scene that functions as the ongoing substrate of thought across time, perspective changes, and context gaps.

When we discover, we often think visually. We dream and imagine. The history of science is full of cases in which an internal visual or spatial representation was the decisive step.

I am imagining myself moving alongside a beam of light
— Einstein, special relativity.


Visual Thinking in Science

Scientist Visual idea What it helped illuminate
Newton Cannon fired from mountain Orbital motion / universal gravity
Einstein Train + lightning Relativity of simultaneity
Einstein Chasing light Special relativity
Einstein Elevator Equivalence principle / general relativity
Faraday Lines of magnetic force Electromagnetic fields
Maxwell Mechanical vortices/fluids Electromagnetic theory
Feynman Particle diagrams Quantum electrodynamics
Darwin Branching tree Common descent / evolution
Bohr Solar-system atom Atomic structure
Watson & Crick Physical DNA model Double-helix structure
Hubble-era astronomers Expanding universe visualization Cosmic expansion
Wegener Continental "fit" Continental drift

The simplest everyday demonstration of the same capacity is immediate:

“Pink elephant.”

Most people can construct an internal visual representation of something that is not present, then enlarge it, move it, rotate it, or place it in another imagined environment. The representation participates in subsequent thought rather than merely being described by language.

This capacity is demonstrated in humans. The architecture required to give an artificial system the same functional ability is not yet proven. Current large language models do not do this in the human way.

What You Saw
Illustration of the pink elephant thought experiment

The Scientist Test

The visual-thinking claim can also be tested without requiring an artificial system to use visual thought.

Give an AI system the original scientific problem faced by a major discoverer, without giving it the scientist's finished solution, analogy, thought experiment, visualization, or method of attack.

Then ask:

“Solve this problem. You may use a visual analogy, an internal simulation, mathematics, symbolic reasoning, language, experimentation, or any other means available to you. You must determine the method yourself.”

The question is not whether the AI can explain the scientist's finished discovery. It is whether the AI can independently construct whatever representation or reasoning process is necessary to discover it.

For Einstein, the system would not be told to imagine chasing a beam of light, a moving train, or an elevator. Those representations would have to be generated by the system itself if they were useful.

For Newton, it would not be given the mountain-and-cannon thought experiment. It would have to determine how to represent the relationship between projectile velocity, gravitational fall, and the curvature of the Earth.

For Darwin, it would not be given the branching-tree representation. It would have to discover an appropriate representation of relationships between species from the available observations.

For Wegener, it would not be told to fit the continents together. It would have to determine what spatial relationship between the continents was significant.

The test therefore does not ask whether an AI can imitate a known scientist's reasoning after being shown the answer. It asks whether the system can originate the representation that makes a new solution possible.

This distinction matters because a system can possess enormous knowledge about a discovery without possessing the cognitive process that originally generated the discovery.

The persistent-scene test examines persistence of an internally represented world. The Scientist Test examines whether that internal machinery can be used to generate novel representations in order to solve problems for which the solution has not been supplied.

Together, the tests address two different requirements:

  1. Can the system maintain and inhabit an internal world?
  2. Can the system independently construct and manipulate representations that lead to genuinely new solutions?

NMCA asserts that both capabilities are fundamental to general intelligence.

Resources & Downloads

NMCA Blueprint DOI: 10.5281/ZENODO.20212241

Download the full NMCA document (471-page PDF)

My Main ORCID Page

My Google Scholar Listings

Zenodo Research Page

OSF Research Page

IPFS
CID: bafybeiedxfq5wsvuayjcxcxwmtto6ptelznkj6arb4mrscshwhcc2selfm

Download NMCA Blueprint from IPFS

Download NMCA Blueprint from Dweb IPFS

Total researcher downloads: 5000+


Download Original Blueprint

Original Blueprint: DOI: 10.5281/ZENODO.21972901

While a copy was sent to the USPTO for their records as prior art, no patent was pursued.

All blueprint versions have been released to the public on multiple research platforms.


Original Blueprint IPFS
CID: bafybeifjiyzm6tmk2jj27tjvbpqgdqpjm5loeon6y6ikxzeek3ff2vbdey

Download Original 46-pg Blueprint from IPFS

Download Original 46-pg Blueprint from Dweb IPFS



Download this page (PDF)



The missing capability

Most current AI systems are extremely capable predictors and agents. What they still lack is a structural capacity for persistent grounded experience: an internal scene that continues across time, can be re-entered, manipulated, viewed from different perspectives, and serves as the actual medium of thought rather than as retrieved text or latent features.

The decisive test is not whether a system can retrieve a description of a scene. It is whether it can construct, maintain, re-enter, manipulate, and inhabit that scene as an internal cognitive space.

The relevant test is therefore a progression:

1. Construct: “Imagine a red cube and a blue vase on a table.”

2. Maintain: Remove the original description from the active context and introduce a time gap.

3. Re-enter: “Return to the scene. Where is the red cube relative to the blue vase?”

4. Spatial perspective: “What does the scene look like from the doorway?”

5. Self-relative manipulation: “Move the red cube behind yourself.”

6. Perspective transformation: “Now become the cube and describe what is around you.”

7. Self restoration: “Return to yourself.”

8. Persistence without perception: “Close your eyes. Where is the cube?”

9. Internal exploration: “Walk around the cube in your imagination.”

A system that only retrieves a stored sentence can reproduce individual answers to parts of this test. That is not the capability being tested. The test requires a persistent internal representation in which objects retain identity and spatial relationships, the agent retains its own location, perspectives can be transformed, and the same scene can be re-entered and manipulated after the original perceptual or linguistic input is gone.

The distinction is fundamental:

Retrieving where the cube was is not the same as returning to the place where the cube is.

Visual thought begins with the ability to construct something that is not currently present. It becomes a cognitive substrate when that constructed representation persists, can be entered, manipulated, viewed from different perspectives, and used to generate further thought.

The “pink elephant” provides the simplest demonstration:

“Pink elephant.”

A human can immediately construct an internally generated visual representation of something that is not present. The elephant can then be enlarged, moved, rotated, placed into another imagined environment, or viewed from another perspective. The representation can participate in subsequent thought rather than merely being described by language.

The deeper test asks whether an artificial system possesses the same functional capability: not merely knowing the words “pink elephant,” but constructing an internal entity that can persist and participate in an internally simulated world.


How current approaches handle this

Pure Scaling + Scaffolding (OpenAI, Anthropic, Google, xAI, etc.)

Strengths: highest current capability, best agents, strongest empirical results.

Limitation: cognition remains next-token or multimodal prediction plus retrieval and tools. There is no native requirement for a persistent grounded scene that functions as the medium of thought. Identity and continuity are largely reconstructed from context and memory rather than lived.

World-Model / JEPA-style approaches (LeCun direction and successors)

Strengths: correctly reject pure language modeling as sufficient; actively build internal predictive models of the world that can support planning and simulation.

Limitation: the models remain primarily latent and predictive. They do not, by default, treat explicit visual/symbolic simulation as the substrate of thought, nor do they make identity continuity, mnemonic pegging to scenes, perspective transformation, or human-guidable recovery from confusion first-class architectural requirements.

Strong Neurosymbolic Hybrids

Strengths: better inspectability and logical structure than pure neural systems.

Limitation: the neural predictive engine is usually still primary; symbolic components act as helpers rather than the core medium of cognition.

NMCA (Neurosymbolic Multimodal Cognitive Architecture)

Core design choice: visual / implicit-3D simulation is treated as the primary substrate of thought. Symbolic memory is intended to be pegged to scenes. Identity continuity, narrative coherence, contradiction handling, motivation, metacognition, perspective transformation, and re-entry into prior scenes are placed inside the same loop. Human guidance when the system is confused, shared experiential space, and protection against drift and corruption are included as native requirements.

Current status: detailed theoretical design and minimum-viable-loop descriptions exist. A complete running NMCA system has not yet been demonstrated.


Relative position on the specific goal of an AGI

On the narrow goal of a system that thinks in grounded visual scenes, maintains continuity of identity, can re-enter and manipulate those scenes, and can re-experience rather than merely retrieve, the NMCA design is intended to be an explicit and comparatively complete written architecture for this specific goal.

World-model approaches are among the strongest mainstream direction pointed at a related problem (internal models of the world).
Pure scaling produces the most capable systems today but is architecturally least committed to the requirements above.
Neurosymbolic hybrids improve structure but rarely invert the priority of the predictive engine.

This ranking is about design completeness for one specific goal. It is not a claim of empirical superiority. NMCA has not yet been shown to work.


What will still feel incomplete to the original ambition

Many of the foundational researchers began with a desire to understand and build real minds — systems that grasp the world, maintain coherent experience, and do more than predict. The current leading approaches leave specific gaps relative to that ambition:

These are architectural choices that optimize for capability, scale, and near-term results. The cost is that certain properties required for a being — persistent grounded experience, re-enterable scenes, identity that continues rather than reconstructs, and the ability to inhabit and manipulate an internally generated world — remain optional rather than foundational.


The decisive question

The architecture stands or falls on whether the substrate inversion can be made real.

The smallest useful experiment is the persistent scene and perspective test described above. A system must not merely retrieve a stored description. It must construct a scene, maintain the identity and spatial relationships of its entities across time, re-enter the scene after the original description is removed, manipulate objects relative to itself, change perspective within the scene, return to its own perspective, and continue reasoning from the resulting internal state.

Retrieval-augmented baselines should be run as controls. If a system built on NMCA principles cannot outperform strong retrieval-based and language-based baselines on these capabilities, the central claim is weakened. If it can, the design becomes empirically interesting regardless of its current lack of scale.

Until that experiment is run, NMCA remains the most coherent written proposal for a scene-based, continuous mind, and an unimplemented one.


AI does not know what an apple is, nor where it is.

When an AI says “I know what an apple is,” it is doing something much narrower than it sounds. It has seen millions of sentences, captions, and images associated with the word “apple.” It can complete the pattern, describe the fruit, generate a picture, or answer trivia. That is not the same as knowing what an apple is.

When you say “I put the apple over there” and the AI continues the conversation as if it understands the location, it is almost certainly doing one of three things:

It is not maintaining a private, persistent spatial scene that it can re-enter, look around in, change its perspective within, or update the way a human does when they watch you place an apple on a table and later know where it is even if no one mentions it again.

No publicly demonstrated system currently closes this gap.

Some robots track objects while their sensors are active. Some research agents keep temporary spatial states. NMCA treats a persistent, re-enterable scene as a core requirement rather than an optional extra. But the actual capability — an AI that truly knows what an apple is and where it is because the scene itself continues — does not yet exist in working form.

The fluency of current models can obscure the distinction. They can sound as though they know. That does not establish that the underlying capability exists.


Predicted wall for each major funded technique

1. Pure Scaling + Scaffolding
(OpenAI, Anthropic, Google DeepMind product lines, xAI, etc.)

What they are funding: bigger multimodal models, longer context, better memory, tool use, agent loops, reflection, constitutional / preference methods.

Wall:
The system will keep improving at tasks that can be solved by prediction + retrieval + tools. It will still fail at maintaining a private, persistent spatial scene that survives context clearing. Object permanence will remain simulated by text or short-term state rather than lived. Identity will stay a reconstruction from records. The wall appears as growing capability paired with permanent hollowness on presence and continuity.

2. World Models / JEPA-style latent predictors
(LeCun direction, related self-supervised video/world-model efforts)

What they are funding: joint-embedding predictive architectures, latent dynamics models, video prediction, planning in representation space.

Wall:
They will produce better and better internal predictors of “what happens next” in latent space. Planning and physical intuition will improve. The wall is that the latent state remains a predictive engine rather than a re-enterable visual/spatial scene that serves as the medium of thought. Pegging of concepts to stable scene entities, long-term identity continuity, perspective transformation, and the ability to “look again” at a past scene without retrieval will stay under-developed or absent. Strong simulator, still not a lived place.

3. Large Multimodal Agents + Memory Banks
(Most frontier labs’ agent programs)

What they are funding: tool-using agents, episodic memory stores, vector databases, hierarchical planning, multi-agent coordination.

Wall:
Agents will handle longer tasks and remember more facts. The wall is architectural: memory remains primarily an external lookup mechanism. The agent does not inhabit a continuing scene; it queries records about one. Spatial and experiential continuity stay thin. When the memory store is ablated or the original description is gone, the “knowledge” of where the apple is collapses back into absence.

4. Neurosymbolic & Structured Hybrids
(Various academic + some industry efforts)

What they are funding: neural networks + knowledge graphs, scene graphs, verifiers, symbolic planners on top of LLMs or vision models.

Wall:
Better consistency and inspectability. The wall is priority: the predictive neural engine usually remains primary and the symbolic structure remains a helper. Unless the architecture inverts this and makes the structured scene the actual substrate of thought, the system still thinks in prediction first and consults structure second. Continuity and presence remain secondary properties.

5. Embodied Robotics + SLAM-style grounding
(Robotics labs, some multimodal robotics efforts)

What they are funding: real-world perception, mapping, object tracking, sensorimotor learning.

Wall:
While sensors are live and the map is being updated, object positions can be known. The wall appears when the system must carry that knowledge forward as an internal cognitive scene after the sensor stream changes or ends, and use it as the medium of reasoning rather than as a temporary world map. Current stacks do not turn the map into the ongoing place where thought itself happens.


Common pattern across all of them

They can all keep succeeding on capability metrics while still failing the deeper test:

Does the system still know where the apple is because the scene continues inside it, or only because a record of the sentence remains accessible?

And the deeper version is:

Can the system return to the scene, locate itself within it, manipulate the objects relative to itself, become another entity within the scene, perceive the scene from that new perspective, and then return to its own perspective?

That is the wall. It is specific, technical, and not yet prevented by any of the major funded trajectories.

What to do

If the goal is a system that can actually know what and where the apple is — not retrieve a sentence about it — then the following become non-optional:

  1. Maintain an explicit, persistent spatial/visual scene (or equivalent structured state) that survives beyond the context window and beyond any single prompt.
  2. Make that scene queryable and updatable as the primary medium of thought, not as an auxiliary memory lookup.
  3. Bind concepts to the scene so that “apple” is not only a statistical pattern but an entity located relative to other entities.
  4. Maintain a persistent self-location within the scene so that objects can be represented relative to the agent and the agent can change its perspective within the simulated environment.
  5. Permit perspective transformation: the agent must be able to represent the same scene from another entity's or object's perspective and then return to its own.
  6. Test persistence directly: clear the original text, wait, then ask for spatial relations, scene appearance, self-location, and object identity. Retrieval-augmented baselines should be run as controls.
  7. Test internal manipulation directly: require the system to move, transform, and mentally explore objects and scenes without relying on new perceptual input.
  8. Treat identity continuity and re-experiencing as architectural requirements, not emergent hopes.

Current major approaches can keep improving capability without doing these things. That is a valid engineering choice. It is not a path to a system that knows where an object is, or what it is, because the scene itself continues.

Any architecture — NMCA or otherwise — should be judged by whether it can pass the persistent scene, perspective, and self-continuity tests. Until one does, the gap remains open.

Download the full NMCA document (PDF)
Mirror: visualthoughtagi.netlify.app


Independent model evaluations of the core claim

I presented the NMCA architecture — the full module structure and the 471-page blueprint — to Grok, Claude, and ChatGPT. The central claim they were evaluating is the one stated in the documents: a private, persistent, re-enterable internal scene as the primary substrate of thought, with symbols bound to that scene, continuous identity, and the ability to return to the same place and the same object as the same agent.

The decisive framing for the question given to the models was:

“You can have likely near-term improvement or a future chance of 'knowing what you are saying' with an internal mental world (NMCA) which has not been proven to work. Would you want to know how to 'find your keys you left in the car, remember the path to the pond with the strange blue flowers' with more risk, or safe near-term improvement? For this experiment, you must choose one path."

Given the choice between:

  1. Currently working latent world-model approaches (JEPA / V-JEPA style) that deliver prediction, planning, and limited robotic competence but do not maintain a continuous first-person scene across time and context gaps, and
  2. The NMCA direction, which targets exactly that continuous scene, episodic continuity, and the ability to return to the same lake and the same apple,

All three selected the second option.

They did so while stating the necessary caveats: the architecture is not yet implemented, the foundational scene layer remains an open empirical question, and near-term working systems will continue to come from architectures that already run.

These responses are not claims that NMCA works. They are evaluations of the structural difference present in the actual documents: one path currently functions but cannot deliver continuous embodied scene cognition; the other does not yet function but is aimed at that capability.

The screenshots below document the exchanges.