# A Philosophical Introduction to Language Models \subtitlefontPart II: The Way Forward

## Metadata
- Author: [[arxiv.org]]
- Full Title: A Philosophical Introduction to Language Models \subtitlefontPart II: The Way Forward
- Category: #articles
- Summary: insert summary
- My notes:
- Summary: The shift to distributed representation models complicates how we interpret neural network behaviors, as disabling certain nodes affects the entire system. Researchers found that these networks can perform addition tasks through a method called 'Fourier multiplication.' While current language models are not perfect representations of human cognition, they can still provide insights into how humans generate language.
- URL: https://arxiv.org/html/2405.03207v1
## LLM Chats
## NotebookLM
## LLM Audio
## Highlights
> Butlin et al. ([2023](https://arxiv.org/html/2405.03207v1#bib.bib17)) survey prominent neuroscientific theories of consciousness that are compatible with computational functionalism. From these theories, they derive a list of ‘indicator properties’ – features or mechanisms that the theories associate with consciousness. The more indicator properties an AI system exhibits, the more likely it is to be conscious. Prominent theories of consciousness include, among others: recurrent processing theory, which proposes that recurrence within neural networks distinguishes conscious from unconscious processing (Lamme [2006](https://arxiv.org/html/2405.03207v1#bib.bib76)); global workspace theory, which claims that consciousness arises when information is broadcast to widespread networks from a limited-capacity workspace (Baars [1993](https://arxiv.org/html/2405.03207v1#bib.bib6), Dehaene & Naccache [2001](https://arxiv.org/html/2405.03207v1#bib.bib34)); and higher-order theories, which hold that consciousness depends on higher-order representation of lower-level activity (Carruthers & Gennaro [2023](https://arxiv.org/html/2405.03207v1#bib.bib21), Brown et al. [2019](https://arxiv.org/html/2405.03207v1#bib.bib12)). ([View Highlight](https://read.readwise.io/read/01jr30xnc3ds85she2dp7gvz05))
> A refrain that has recently appeared in many popular science articles, blogs, and social media posts holds that LLMs may reflect the discovery of an entirely new kind of ‘alien’ intelligence that is fundamentally unlike our own. This may be seen as a mere statement of fact – an optimistic appraisal that the linguistic behavior already exhibited by state-of-the-art LLMs is obviously intelligent, coupled with a pessimistic appraisal that the underlying processes producing this behavior are much at all like ours. With respect to most of the philosophical questions raised above, however, it reflects an almost complete dodge and a return to a naive behaviorism about intelligence by use of an eye-catching label. In any particular case, we need to ensure that before we accept the use of such a label as a descriptor of some scientifically-important LLM behavior that we have scientifically-respectable methods to assess that behavior along the important dimensions reviewed in Part I. ([View Highlight](https://read.readwise.io/read/01jr31f7mx9ma746edxf9axdmt))
> As Nickles ([2020](https://arxiv.org/html/2405.03207v1#bib.bib102)) points out, this language first began to appear as a descriptor of deep learning performance well before the current boom in LLMs, especially regarding new machine-learning-based methods to analyze data in high-complexity sciences like biology and fundamental physics. There, the language of ‘alien’ intelligence was conceptually tied to older philosophical dreams of rationalists like Descartes to create a science that was free from the shackles of human perspective and bias – an objective science free from the foibles of human ‘powers and faculties’, as Hume put it. The hope kindled by ‘alien intelligence’ here is that by automating the process of data collection and analysis and relaxing requirements on transparency and intelligibility, “nonhuman sensors and neutral algorithms can replace these human ‘powers and faculties’ ” with alien agents that are wholly “data-driven and increasingly are able to construct their own models and search strategies”. Indeed, there is some reason to think that breakthrough systems like AlphaFold, which have nearly solved the highly-complex protein folding problem that eluded human microbiologists for more than a century, achieved their results precisely by discovering inscrutable features that are too complex for humans to understand (Buckner 2020, Kieval 2023). Whether one buys the talk of alien intelligence and inscrutable features in this context, however, the concept derives its utility from pragmatic scientific goals that humans largely accept – the ability to predict and perhaps control natural phenomena. ([View Highlight](https://read.readwise.io/read/01jr31jestg56b6zn5qk4dfxbn))
> In Part II, we turned to newer philosophical questions raised by the current state of the art in language modeling research. We reviewed philosophically-grounded interventionist methods that aim to uncover the causal mechanisms underlying LLM performance, such as the existence of algorithmically meaningful internal representations and computations. This growing body of work suggests that state-of-the-art LLMs implement important aspects of the abstractness, systematicity, and generalizability of human cognitive processes, although they still fall short in terms of efficiency, completeness, and agency. We also discussed emerging trends in LLM research, including the development of multimodal models and the integration of LLMs into broader ‘agent’ architectures. These approaches attempt to address some of the weaknesses of traditional LLMs by creating AI systems that can learn from self-exploration, maintain internal memories, and formulate plans based on their interactions with real or virtual environments. While it remains unclear whether current LLMs satisfy proposed computational markers of consciousness, these advancements open up the possibility of future systems exhibiting increasingly sophisticated forms of intelligence. ([View Highlight](https://read.readwise.io/read/01jr321ae5vfafh744g71nhb4r))