Abstract
The race toward world models in 2026 has converged on a single question: can AI move beyond observing the world to actively intervening in it? This paper argues that world models are not approaching causal anchoring; they are increasingly demonstrating why they cannot anchor causation. Drawing on Pearl's causal hierarchy (association, intervention, counterfactual), I analyze three competing technical pathways: JEPA (causation as provably identifiable structure), video generation and spatial intelligence (causation as reliable downstream task performance), and LLM-based multimodal systems (causation as commonsense plus perception). Each pathway approaches the L3 (counterfactual) boundary but fails to cross it, though in epistemologically distinct ways. I then introduce the Negative Subjectivity framework and its central concept of "causal dissolution"—the ontological feature whereby AI substitutes statistical association for causal structure—and demonstrate that "intervention cannot be endogenized" is not a technical limitation but a metaphysical one: any system trained solely on observational data has its learning boundary defined by the data distribution, and intervention is by definition a perturbation outside that distribution. Four independent empirical evidence chains (stable-worldmodel, Kang et al.'s feature priority inversion, RLVR's sharpening effect, and Cosmos 3's dual-tower architecture) converge on this conclusion. The paper further distinguishes between object-language-level and metalanguage-level causation: while world models cannot answer causal questions at the object level (L2–L3), the formal impossibility proofs about their limits constitute a metalanguage-level causal understanding accessible only to human observers. World models' destiny is to remain on the boundary of negative subjectivity. They serve not as subjects that "do" the world, but as mirrors that precisely define what "doing" means by perfectly simulating the absence of understanding. A central and controversial finding is that L3 (counterfactual) is not an extension of L2 (intervention) but a precondition of it—an epistemological inversion of Pearl's progressive ladder that will be argued in Section 6.