Showing posts with label Visual Perception. Show all posts
Showing posts with label Visual Perception. Show all posts

Friday, July 17, 2009

Language, Thought, and Space (III)



In the second chapter of his book, Stephen Levinson discusses a concept that has been crucial to this blog: frames of reference. (see e.g. these posts) The term as it is used today was coined by Gestalt theorists of perception in the 1920s and was used to signify the steady and constant background against which other objects could be made out and identified. It can be defined as “‘a unit or organization of units that collectively serve to identify a coordinate system with respect to which certain properties of objects, including the phenomenal self, are gauged’ (Rock 1992: 404, emphasis in Levinson 2003: 24).

Frames of references seem to be highly similar across modalities such as vision, touch, gesture, and language. Without these structural similarities (or ‘isomorphisms’) “we could not reach to what we see, or talk about what we feel with our hands, or give route descriptions in language and gesture.” (Levinson 2003: 25). There are, however, also differences: vision is viewer-centred, and touch and grasp are object-centred.

In general, frames of references can be classified by the following distinctions.
Absolute vs. Relative. Psychologically, the received view is that we organize our spatial thinking in relation to objects and ourselves. The frame of reference is thus relative to our own ego-centric bodily coordinates. An absolute frame of reference, on the other hand would consist of fixed angles with coordinates that do not depend on our personal egos as anchoring. And as we have seen, contrary to the received view, both kinds of frames of references are employed in the world’s languages. (Levinson 2003: 27f.).
Similar, but not completely identical is the differentiation between egocentric and allocentric frames of reference. This designates a difference
"between coordinate systems with origins within the subjective body frame of the organism, versus coordinate systems centred elsewhere (often unspecified).” (Levinson 2003: 28).

Our mental maps of our environment and our place are either egocentric or allocentric and landmark-based, including the relations, distances and angles between different landmarks, or allocentric and based on “fixed bearings.” These distinctions can not only be found in the world’s languages, but are also used by neuroscientists when they look at the mental map-building capacities of animals.
In studies of conceptual development it was also argued, following Jean Piaget, that for a long time ‘egocentric’ frames of reference are primary and that children switched to ego-centric frames of reference only to a much later date.

In studies of the visual system we often find a distinction between viewer-centred vs. object-centred. If we identify an object we are also able to mentally rotate it and imagine how it would look from another angle. This means that the retinal impression of the viewer gets interpreted and classified in a more abstract object-centred frame of reference during perception.
Another distinction made when looking at visual imagery and visual perception is that between orientation-bound vs. orientation-free. Orientation-bound information changes with perspective and change of location, whereas orientation-free information does not change. For example, when we rotate a d it can become a b, the information changes. But a ball looks the same from all perspectives and the information is thus orientation-free.
The most important distinction for psychology and language however, is the difference between
“viewer-centred frames, object-centred frames, and environment-centred frames of reference.Ina viewer-centred frame, objects are represented in a retinocentric, head-centric or body-centric coordinate system based on the perceiver’s perspective of the world. In an object-centred frame, objects are coded with respect to their intrinsic axes. In an environment-centred frame, objects are represented with respect to salient features of the environment, such as gravity or prominent visual landmarks. “ (Carlson-Radvansky & Irwin 1993: 224).
Levinson called these the relative, intrinsict and absolute frames of reference. (Levinson 2003: 33).

The distinctions made in various disciplines at times are quite confusing and there are many conflicting positions. However, a broad differentation such as this seems valid.
Next, we have to distinguish between three levels on which different frames of references can be constructed: perceptual, conceptual, and linguistic. There is especially much diversity on the linguistic level, which will be discussed in my next post. As I'm going home tomorrow I don't really know when I'll have access to the internet again, but I hope it wont't be too long.

Reference:
Carlson-Radvansky, L.A. and Irwin, D.A. (1993): Frames of reference in vision language: Where is above? Cognition
46: 223-244

Levinson, Stephen C. (2003) Space in Language and Cognition : Explorations in Cognitive Diversity. West Nyack, NY, USA: Cambridge University Press.

Rock, I. (1990), The frame of reference, in I. Rock (ed.), The legacy of Soloman Asch, pp. 243– 268. Hillsdale, NJ: Lawrence Erlbaum.

Monday, January 28, 2008

The Cognitive Foundations of Perspective III

In this post I continue to elaborate on my inquiry into the cognitive structure of the shared systemic space. As we have seen Discourse Representation Theory, File Change Semantics, the Theory of Visual Indexes, and the theory of object files together construe a useful methodology for a research program interested in the structure of mental representation.

Hurford himself uses Kamp & Reyle’s (1993) box-notation of Discourse Representation Structures to describe the mental representations of non-human animals, because the observations I discussed in my last post led him to conclude that
“it is natural to assume that human language evolved by building upon pre-existing representational schemes in animals” (Hurford 2007: 140).
There is one thing we have to make clear, before we can make the findings I describe in my last post fruitful for research into the properties of the systemic space: the contributions of global and local attention. Basically, when we perceive a visual scene
“An initial rapid pass through the visual hierarchy provides the global framework and gist of the scene and primes competing identities through the features that are detected. Attention is then focused back to early areas to allow a serial check of the initial rough bindings and to form the representations of objects and events that are consciously experienced.” (Treisman 2005: 541)

To give you an example adapted from Hurford (2007: 152), if we want to represent the results of global and local scans toward a scene in Hurford’s adapted notation of Kamp & Reyle’s (1993) Discourse Representation Structures, the result of a quick global scan would look like this:
but if the result of focal attention to the individual scene would have the following mental representation:
This process is closely related to profiling, that is the distinction between figure, the
“integrated visual experience that ‘stands out’ in the center of attention” (Coren et al. 1999: 564),
and 'ground',
“the background against which figures appear” (Coren et al. 1999: 565).
The most famous illustration of this principle is Rubin’s reversible face-vase figure, where we can either see a white vase as the figure which stands in front of a black ground or two black faces that are in front of a white background (Goldstein 1999: 187).

As there is additional “evidence that imagery engages brain mechanisms that are used in perception and action“ (Kosslyn et al. 2001: 635), and the fact that “perceptual representations are routinely activated during comprehension” (Zwaan 2004), (and as we have already established that we can indeed we can draw an analogy between these two areas), we are able to make an analogy between this “primitive example of perceptual organization” (Coren et al. 1999: 296) and language comprehension.

Thus we can say that the mental representations underlying language comprehension, i.e. the structure of the systemic space, probably underlies the same principle of global and local attention/figure and ground. What I mean by this is that when we create a discourse universe, we create layers of meaning on a variety of planes, such as temporal layers (Before I started studying, I… But now… etc.), Theory of Mind layers (I thought that he knew that she knew that I…), degrees of relevance, and so on.
In such a structured systemic space, some aspects are more important than others, are ‘background knowledge’ so to speak, whereas other aspects are in the foreground and are brought to the spotlight of our attention. They represent the ‘figure’ of the message against the ‘ground’ of context. (Köller 2004: 442f)
Köller (2004: 442ff.), following the German linguist Harald Weinrich, calls this power of language to establish such layered and meta-structured systemic spaces its ability to create reliefs.

Summarizing these considerations, there are now some additional properties of the systemic space, and we have gained additional insight into how language can locate things in the coordinate system of subjective orientation:



In my next post I'll speculate a bit about the ontogenetic as well as phylogenetic pathway that may have led to our modern ability to take perspectives.

P.S:
On a related note, the admin of the language evolution blog has posted the first post on "Major Language Evolution Papers", this time about Tomasello et al.'s (2005) great paper on the "Origins of human cognition" - go check it out!


References:

Coren, Stanley, Lawrence M. Ward and James T. Enns. Sensation and Perception. 5th
ed. Fort Worth: Harcourt Brace, 1999.

Goldstein, E. Bruce. Sensation & Perception. 5th ed. Pacific Grove: Brooks/Cole, 1999.

Hurford, James M. 2007. The Origins of Meaning: Language in the Light of Evolution. Oxford: OUP.

Kosslyn, Stephen M., Giorgio Ganis and William L. Thompson. “Neural Foundations
of Imagery.” Nature Reviews Neuroscience 2 (2001): 635-642.

Kamp, Hans and Uwe Reyle. 1993. From Discourse to Logic: Introduction to Modeltheoretic Semantics of Natural Language, Formal Logic and Discourse Representation Theory. Dordrecht, Holland: Kluwer Academic.

Köller, Wilhelm. 2004. Perspektivität und Sprache. Zur Struktur von Objektivierungsformen in Bildern, im Denken und in der Sprache. Berlin/ New York: de Gruyter.

Treisman, Anne (2005). Psychological issues in selective attention. In Michael A. Gazzaniga (Ed.), The Cognitive Neurosciences, III,. Cambridge, MA: MIT Press.: 529–544.

Zwaan, Rolf A. (2004). The immersed experiencer: toward an embodied theory of language comprehension. In: B.H. Ross (Ed.), The Psychology of Learning and Motivation, Vol. 44. New York: Academic Press.

Thursday, January 24, 2008

The Cognitive Foundations of Perspective II

So in my last post I summarized a bunch of overlapping scientific research programs which can be combined under the notion of mental representation as a creation of a virtual “systemic space”, which can be is exemplified by the idea of short-time memory as a “workbench”, where you can store and manipulate virtual objects, which is used frequently in the cognitive sciences.

Quite strikingkly, when we use Wittgenstein’s (1953) notion of shared activities as ‘games’ and the creation of a shared systemic space as a ‘language game’, there are also some interesting implication of Pinker et al.’s (2008) statement that language serves as a as a
“reference point in coordination games.” (Pinker et al. 2008).
Thus we can also locate their definition of “focal point” i.e.,
"a salient location that two rational agents can agree on when they would be better off coordinating their behavior than acting independently.” (Pinker et al. 2008: 837)
on Bühler’s coordinate system of subjective orientation, or our notion of a shared systemic space.
On this view we can regard a discourse as the negation of the exact focal position of a proposition in the coordinate system of the shared systemic space, (what Pinker et al. call the ‘problem space’). According to Pinker et al., such negotiation is mostly used
“to negotiate the type of relationship holding between speaker and hearer (in particular, dominance, communality, or reciprocity)” (Pinker et al. 2008: 833)
Hurford (2007) reviews further approaches which can be subsumed under this notion:
The first he mentions is Discourse Representation Theory (DRT: Kamp and Reyle 1993). DRT is concerned with describing from a semantic point of view how in discourse we build up a universe consisting of the things we mention. In Kamp and Reyle’s notation this ‘discourse universe’ is represented by a box, which they call a Discourse Representation Structure. The most simple kind of structure in such a DRS would look something like this, with the set of ‘discourse referents’ (x, y, z, etc.) at the top of the box:

This, of course, is basically what I, following Köller (2004), would call a systemic space.
According to Kamp and Reyle, the process of semantic representation is the following: On hearing a sentence (S1), we create a DRS. When the next sentence (S2) is uttered in discourse, this sentence contributes new information to the already constructed DRS, or in my notation, new propositions are transferred into the systemic space or old ones are transformed. This process goes on and on with every new sentence.
Thus, the hearer relates the new sentence to the informational structure already obtained, thereby dynamically construing and manipulating the shared systemic space. (Kamp & Reyle 1993: 59). We could say that with every sentence we change from one mental model of the discourse to another (i.e. M1 ->M2 ->M3, etc.) (Kamp & Reyle 1993: 96).
Kamp & Reyle also have something very interesting to say about the cognitive and attentional underpinnings of the representation of shared systemic spaces: When you hear a proper name in discourse, you assign it an index (like x, y, z, etc.) to make temporary reference possible and to keep track of the discourse referent. They call this an ‘external anchor’ for a discourse referent (x) which maps x onto some real individual (like say, Zombie-Scientist George if he really existed) (Kamp & Reyle 1993: 248)

This proposal is closely related to some other theories of cognition. According to Zenon Pylyshyn’s theory of FINST, regarding the ability to keep track of moving objects in a visual scene,
“a small number of visual objects can be preattentively indexed or tagged and thereby accessed more rapidly by a subsequent attentional process (e.g., the traditional "spotlight of attention") (Sears & Pylyshyn 2000: 1)
These visual indices can be seen as mental labels that can be attached to objects in order to keep track of them. (Hurford 2007: 92).

Another related theory, highlighted by Hurford (2007: 139), is ‘File Change Semantics”, according to which
“A listener’s task of understanding what is being said in the course of a conversation bears relevant similarities to a file clerk’s task. Speaking metaphorically, let me say that to understand an utterance is to keep a file which, at every time in the course of the utterance, contains the information that has so far been conveyed by the utterance.” (Heim, 1983:167)
Again, there is a psychological theory which closely echoes this assessment in the visual domain. According to Kahneman & Treisman (1992) set up ‘object files’
“as a temporary episodic representation, within which successive states of an object are linked and integrated” (Kahneman & Treisman 1992: 175)
These files are constantly updated by new information about the target’s features or location.
The attentional limit of things we can consciously be aware of seems to lie at 4 target objects, (Hurford 2007: 93, Cowan 2001) and seems to hold true for the perceptual space as well as for the systemic space.

This research about visual indexes as
“a means of setting attentional priorities when multiple stimuli compete for attention” (Sears and Pylyshyn 2000: 2) )
,as well as the idea of information ‘files’ goes very well with our notion of language as a means to pilot attention toward certain propositions in the systemic space.

In sum, we see that there are overlapping theories concerning mental representations of the perceptual space as of the virtual systemic space. Hurford (2007) argues that this independent convergence of several areas of research indicates that:
“one bit of language-processing machinery has been co-opted (and probably adapted somewhat) from pre-existing visual scene processing machinery.” (Hurford 2007: 140)

References:


Cowan, Nelson. 2000. The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences 24(1), 87–114.

Heim, Irene. 1983. File change semantics and the familiarity theory of definiteness. In R. Bäuerle, C. Schwarze, and A. von Stechow (Eds.), Meaning, Use, and Interpretation of Language. Berlin: Walter de Gruyter: 164– 189.

Hurford, James M. 2007. The Origins of Meaning: Language in the Light of Evolution. Oxford: OUP.

Kahneman, Daniel. and Anne. Treisman.1992. The reviewing of object files: object-specific integration of information. Cognitive Psychology 24, 175–219.

Kamp, Hans and Uwe Reyle. 1993. From Discourse to Logic: Introduction to Modeltheoretic Semantics of Natural Language, Formal Logic and Discourse Representation Theory. Dordrecht, Holland: Kluwer Academic.

Köller, Wilhelm. 2004. Perspektivität und Sprache. Zur Struktur von Objektivierungsformen in Bildern, im Denken und in der Sprache. Berlin/ New York: de Gruyter.

Pinker, Steven, Martin A. Nowak and James L. Lee. 2008. The logic of indirect speech. Proceedings of the National Academy of Sciences 105.3: 833–838.

Sears, Christopher R. and Zenon W. Pylyshyn. 2000. “Multiple object tracking and attentional processing.” Canadian Journal of Experimental Psychology 54(1), 1–14.

Wittgenstein, Ludwig. 1953. Philosophical Investigations. Oxford: Basil Blackwell. (Translated by G. E. M. Anscombe)