Context-Dependent Affordance Computation in Vision-Language Models
A Functional Affordance Architecture for Visual Understanding
Abstract
We propose that the standard vision pipeline inverts the actual structure of biological spatial cognition. Spatial awareness is grounded in functional semantics: geometric structure becomes accessible to reasoning only through prior contextual and affordance encoding. We formalize this as a semantic-first spatial architecture and derive testable predictions distinguishing it from geometry-first alternatives, evaluated against the behaviour of contemporary vision-language models.
Suggested citation
Murad Farzulla (2025). Context-Dependent Affordance Computation in Vision-Language Models. Dissensus Working Paper DAI-2518. DOI: 10.5281/zenodo.18091089
Methodology
Affordance theory
Ecological psychology
Phenomenology
Action-distance metrics
Topics
Philosophy
Cognitive Science