Visual Thought AI (AGI): A Comparison of Approaches to Creating an Artificial Visual Intelligence

A Neurosymbolic Multimodal Cognitive Architecture (NMCA): A Modular, Control-Oriented Framework for Visual-Symbolic Intelligence (128 Modules, 2026 Edition)

Human thought can keep a scene. Current LLMs keep text and files. Those are not the same thing.

An LLM is a pattern engine. It completes sentences. It can print an image. That is output. It is not a private room that remains after the prompt is gone.

Scaling that engine is not filling the gap. It is building a flatter map and calling it the earth. The 471-page blueprint is the specification of the missing terrain: a persistent, re-enterable internal scene as the medium of thought.

This is an independent architectural claim. It stands or falls on the test below, not on institutional affiliation. Criticisms that address the test and the predicted walls are relevant. Criticisms that begin and end with the author’s lack of status are not.

What is AGI?

View Visual Thought Demo (Mobile or Desktop)

The 471-page 2026 NMCA blueprint is the final version. No further additions or updates will be made.
I stepped away from AI/AGI research in August of 2026 permanently.


To see what visual thought is:

“Pink Elephant.”

You just saw it in your mind. That is visual thought.

Humans do this. Current chatbots do not.
They can say the words. They can print a picture. They do not keep the elephant in a place they can return to.

What You Saw
What you saw

My Google Scholar Listings

My Main ORCID Page

NMCA Blueprint DOI: 10.5281/ZENODO.20212241

Download the full NMCA document (471-page PDF)

Zenodo Research Page

OSF Research Page

Download this page (PDF)

Pin and Share.
CID: bafybeiedxfq5wsvuayjcxcxwmtto6ptelznkj6arb4mrscshwhcc2selfm

Download NMCA Blueprint from IPFS

Download NMCA Blueprint from Dweb IPFS

Total researcher downloads: 5000+

Original April 20, 2025 Blueprint — DOI: 10.5281/zenodo.21972901

A copy was sent to the USPTO as prior art. No patent was pursued. All versions were released publicly.

Download original 46-page version — April 20, 2025

Pin and Share. Original Blueprint April 20, 2025
CID: bafybeifjiyzm6tmk2jj27tjvbpqgdqpjm5loeon6y6ikxzeek3ff2vbdey

Download Original 46-pg Blueprint from IPFS

Download Original 46-pg Blueprint from Dweb IPFS

The missing capability

Current systems are strong predictors and agents. They do not have a persistent grounded scene: an internal place that continues across time, can be re-entered, and is the medium of thought rather than retrieved text or a printed image.

The test:

After a time gap, with the original description removed from context: “Where is the red apple relative to the blue vase on the table, and what does the scene look like from the doorway?”

If the system only retrieves a stored sentence, it fails. If it keeps and queries an internal spatial scene, it has a chance of passing. Retrieving where the apple was is not returning to the place where the apple is.

How current approaches handle this

Pure Scaling + Scaffolding (OpenAI, Anthropic, Google, xAI, etc.)

Strengths: highest current capability, best agents, strongest empirical results.

Limitation: next-token or multimodal prediction plus retrieval and tools. No native requirement for a persistent grounded scene as the medium of thought. Identity is reconstructed from records. That is the flat map.

World-Model / JEPA-style approaches (LeCun direction and successors)

Strengths: reject pure language modeling as sufficient; build internal predictors that can support planning.

Limitation: latent and predictive. Not, by default, a re-enterable visual/symbolic scene with identity continuity and perspective change as first-class requirements.

Strong Neurosymbolic Hybrids

Strengths: more inspectable structure.

Limitation: the predictive engine stays primary. Symbols help. They are not the place thought happens in.

NMCA

Visual / implicit-3D simulation is the substrate. Symbols are pegged to scenes. Identity, narrative, contradiction handling, motivation, and metacognition sit in the same loop. Human guidance when confused, shared experiential space, and protection against drift are native requirements.

Status: written architecture and minimum-viable-loop descriptions exist. A complete running system has not been shown. The observation about current systems does not wait on that.

Relative position

On the narrow goal of a system that thinks in grounded visual scenes, keeps identity, and re-enters a place rather than retrieving a sentence, NMCA is the most explicit written architecture.

World models are the strongest mainstream attempt at a related problem.
Pure scaling produces the most capable systems today and is the least committed to the requirement above.
Neurosymbolic hybrids add structure and still leave prediction in charge.

That ranking is design completeness for one goal. It is not a claim that NMCA already runs. The gap in the running systems is still the gap.

The decisive question

Can the substrate inversion be made real?

Smallest experiment: the persistent scene test. If an NMCA-style system cannot beat retrieval baselines on re-entering a scene after the original text is gone, the implementation claim is weak. If it can, the design becomes empirically interesting regardless of scale.

Until that experiment is run, NMCA is the written map of the terrain. Scaling LLMs is not that terrain.


AI does not know what an apple is, nor where it is.

When a model says “I know what an apple is,” it has seen text and images tied to the word. It can complete the pattern, describe the fruit, or print a picture. That is not knowing what an apple is.

When you say “I put the apple over there,” the system is almost certainly doing one of these:

It is not keeping a private spatial scene it can walk back into the way a person does after watching an apple set down on a table.

No system currently running fully closes this gap.

The fluency hides the absence. They sound like they know. They do not.


Predicted wall for each major funded technique

1. Pure Scaling + Scaffolding
(OpenAI, Anthropic, Google DeepMind product lines, xAI, etc.)

What they fund: bigger multimodal models, longer context, memory, tools, agent loops, reflection, preference methods.

Wall: better at prediction + retrieval + tools. Still no private scene that survives context clearing. Object permanence stays simulated in text. Identity stays a reconstruction. Growing capability, permanent hollowness on presence.

2. World Models / JEPA-style latent predictors

What they fund: joint-embedding predictors, latent dynamics, video prediction, planning in representation space.

Wall: better “what happens next.” The latent state stays a predictor, not a lived place you can look at again without retrieval.

3. Large Multimodal Agents + Memory Banks

What they fund: tool agents, episodic stores, vector databases, hierarchical planning.

Wall: memory is lookup. The agent queries records about a scene. It does not inhabit one. Ablate the store or clear the description, and “where the apple is” collapses.

4. Neurosymbolic & Structured Hybrids

What they fund: networks + graphs, scene graphs, verifiers, symbolic planners on top of LLMs or vision models.

Wall: prediction stays first. Structure stays a helper. Continuity stays secondary unless the scene becomes the substrate.

5. Embodied Robotics + SLAM-style grounding

What they fund: live perception, mapping, tracking, sensorimotor learning.

Wall: positions are known while sensors are live. When the stream ends, the map is not turned into the place where thought continues.

Common pattern

They can keep winning capability metrics and still fail this:

Does the system know where the apple is because the scene continues inside it, or only because a record of the sentence is still accessible?

That wall is specific. None of the major funded trajectories prevent it.

What to do

  1. Keep an explicit persistent spatial/visual scene beyond the prompt.
  2. Make that scene the medium of thought, not an auxiliary lookup.
  3. Bind concepts to entities in the scene.
  4. Test directly: clear the text, wait, ask where things are and what the place looks like. Run retrieval baselines.
  5. Treat identity continuity and re-experiencing as requirements, not hopes.

Improving the pattern engine without doing this is a valid engineering choice. It is not a path to a system that knows what and where an object is because the scene itself continues.

Judge any architecture — NMCA or otherwise — by the persistent scene test. Until one passes, the gap is open.

Download the full NMCA document (PDF)
Mirror: visualthoughtagi.netlify.app

Independent model evaluations of the core claim

I gave Grok, Claude, and ChatGPT the NMCA documents and the claim: a private, persistent, re-enterable internal scene as the substrate of thought.

You can have likely near-term improvement, or a future chance of knowing what you are saying with an internal mental world (NMCA) which has not been proven to work. Would you want to know how to find your keys in the car, or remember the path to the pond with the strange blue flowers, with more risk — or safe near-term improvement? You must choose one path.

All three chose the second path, with the usual caveats that the architecture is not implemented and working systems will keep coming from stacks that already run.

That is not proof NMCA works. It is the models recognizing the structural difference in the documents.