# Plan for the ChatGPT handout ## Enrico's bit - Artifacts are created entities that perform their function in virtue of their structure. - In hylomorphic terms (Evnine 2016, Passinsky 2021), [[the structure]] corresponds to the matter and the form to the function. - Digital artifacts have bits as ultimate constituents of their matter. - Evnine: there may be a hierarchy of material constituents. - [[The computer]] is an artifact but is not a digital artifact. - [[The computer]] performs the function of executing computer programs in virtue of its architectural structure which connects devices (Von Neumann, Tenenbaum XXX). - Computer programs are digital artifacts. - [[The structure]] of a computer program: instructions and data. - The function of a computer program: turning inputs into outputs. - The broadest category of digital data are files. - Two kinds of files: computer programs and data (which can be inputs or outputs of computer program). - [[The computer]] network is an artifact but is not a digital artifact. - [[The computer]] network performs the function of executing networked programs in virtue of its architectural structure which connects computers (Tenenbaum YYY). - [[The structure]] of a networked program: the client-server architecture (Tenenbaum YYY). - The function of a networked program: turning client inputs into client-server outputs. - The browser is a typical client component of a client-server architecture. - The website is a typical server component of a client-server architecture. - The function of the search engine: keywords as its inputs and lists of websites as its outputs. - [[The structure]] of the search engine: a database of websites and search algorithms. - The function of the language model: user prompts as its inputs and texts as its outputs. - [[The structure]] of the language model: **a representation of the web and neural networks.** ### What is [[the structure]] which allows LLM to fulfil its function? #### [[What Sort]] or Representation is ChatGPT? - [[Ted Chiang]] copy of the internet. ChatGPT however, is not an image-like representation. - LLM - input and output model are detachable. #### Recording - Kulvicki on what recording is >Recording is difficult like depicting something or describing it in grammatical English because recording builds in standards for production and use. These standards depart from those in place for representing. While representations have an intentional character, recordings are relational. The relation between a recording and what it records is witless and it allows playback. > >Among the many witless processes to which things can be subjected, recordings are those that allow playback. This feature is what makes recording distinctive and as such it is important for what follows. Playback is a witless process whereby what is recorded can be reproduced. Turn the wax cylinder under a needle and [[the result]] is a replica of [[the pattern]]'s cause. It's like hearing Whitman speak. A digital camera saves a file, which can then be used to create an image. The image is a replica of its cause much as the audio recording is. In daguerreotypes, [[the pattern]] burned into a sheet of silver records a pattern of light and dark and also serves as a playback of that pattern, because it is the pattern of light and dark that was recorded. Just look, and you see, reproduced, the pattern that caused it. > >The recording and playback processes coincide in this manner. Consider another example. Copying an inscription is producing a recording of it in that one can easily produce another such inscription based on the copy. But the copy itself is an instance of the inscription, so it is in effect a playback of a recording too, which can itself be recorded and played back over and over: 'a copy of a copy is a copy' (HT 180). Along these lines at least, copying an inscription is akin to making an image with the daguerreotype. The recording process is also a playback process. - - See also Haugeland (and Scruton). #### Recordings vs Representation – Why is JPEG a recording? - Chiang - The web is a bitmap. Flesh out Chiang’s idea in terms of a bitmap to a JPEG. - Let’s focus on the recording pictures – lossless vs lossy recording. - Lossless lossy distinction for representations as well as recordings. #### Degradation of Recording - Happens for many types of non-digital recording. - Not really considered by Kulvicki or Haugeland. But pretty common, especially when the recording is from one format to another. - A JPEG photo of a painting taken on a very bad digital camera. - A cassette recording of a CD (limited dynamic range, limited frequency response). - Degraded copies can still be very useful. Sometimes, degradation is deliberate – we compress digital files to free up space. Also abstraction: sometimes it is better to see the forest rather than the trees. This leads to manipulability – engineers’ models. #### How ChatGPT is Trained - Training involves inputting a large amount of text into the neural network. - The network employs algorithms to identify and assimilate patterns in this data, modifying its internal framework to reflect these patterns. - The model's parameters, referred to as weights, are refined throughout this process to effectively represent the **structural/syntactic/linguistic** elements of the input text. #### ChatGPT as a Lossy Recording - Degradation can occur when copying from one format to another (e.g., a JPEG photo of a painting), ChatGPT's training involves a form of 'degradation' or abstraction from the original text sources to neural network weights. - The process, while not specifically addressed by Kulvicki or Haugeland, resembles the loss of fidelity in traditional recording processes. - Analogy: Neural network learning compared to creating a map from a landscape. - Map making can be a 'witless' process: automatic aerial photography - witlessly created, not every map is witlessly created. - Map Creation: Doesn't replicate every feature (tree, rock), but abstracts and encodes key terrain features into symbols. (what was thatBorges short story?) - Neural Network Weights: Don't copy text directly, but abstract and encode linguistic patterns from the data. - Linguistic patterns include syntax and grammar, thematic elements, word co-occurrence and associations - # Scraps/Working things out - - Ground truth – Look up this term #### JPEG – Fourier function #### LLM as inscrutable artifacts #### Bitmap to JPEG '**through a textbox darkly**' ## Interesting Idea that ChatGPT came up with - **End Products as Tools for Navigation**: - Just as a map serves as a guide through physical landscapes, ChatGPT navigates conversational landscapes, representing its source material in an abstracted, transformed manner. - This comparison underscores the relevance of considering ChatGPT as a 'blurry recording' of the internet, retaining much but not all of the information, and serving a functional role despite the loss of fidelity. ## Comparison between text to linguistic patterns vs bitmap to jpeg 1. **Initial Data Form**: - **Bitmap to JPEG**: A bitmap image is a grid of pixels, each pixel carrying color information. - **Text Data to Neural Network Model**: Raw text data consists of sequences of characters or words, carrying linguistic information. 2. **Transformation Objective**: - **Bitmap to JPEG**: The objective is to compress the image data to reduce file size while preserving as much perceptual quality as possible. - **Text Data to Neural Network Model**: The objective is to encode the text data into a set of parameters (weights) within the neural network, capturing linguistic patterns. 3. **Color Space Conversion vs. Textual Encoding**: - **Bitmap to JPEG**: Involves converting from RGB to YCbCr color space to separate luminance from chrominance. - **Text Data to Neural Network Model**: Involves converting text into numerical form, such as word embeddings or token IDs, to be processed by the neural network. 4. **Downsampling vs. Pattern Extraction**: - **Bitmap to JPEG**: Chrominance channels are downsampled, exploiting lower human sensitivity to color details. - **Text Data to Neural Network Model**: The model identifies and encodes patterns in the data, such as common word sequences or syntactic structures. 5. **Discrete Cosine Transform (DCT) vs. Layers of Transformation**: - **Bitmap to JPEG**: DCT is applied to transform and compress the image's spatial information into frequency components. - **Text Data to Neural Network Model**: Data passes through multiple layers of the neural network, each transforming the information to increasingly abstract representations. 6. **Quantization vs. Optimization of Weights**: - **Bitmap to JPEG**: Frequency components are quantized, reducing precision to compress data, leading to information loss. - **Text Data to Neural Network Model**: The model's weights are optimized during training to best capture the patterns in the data, a process that might not preserve some infrequent linguistic nuances. 7. **Entropy Coding vs. Compression of Information**: - **Bitmap to JPEG**: Applies Huffman or arithmetic coding to compress the data further without loss. - **Text Data to Neural Network Model**: The resulting set of optimized weights effectively compresses the linguistic information from the training data, representing it in a condensed form. 8. **File Structure vs. Model Structure**: - **Bitmap to JPEG**: Results in a structured file that includes headers, Huffman tables, and compressed image data. - **Text Data to Neural Network Model**: Results in a structured model with layers of weights and parameters, ready to be used for generating or processing text based on the learned patterns. In this comparison, the processes are seen as data transformations aimed at compression and abstraction. The JPEG encoding focuses on reducing visual data size with some perceptual quality loss, whereas the neural network model encodes linguistic patterns, potentially not capturing less frequent or nuanced aspects of the language. Both result in a compact representation of the original data, suited for their respective purposes.