# Deep Learning and Synthetic Media

## Metadata
- Author: [[Raphaël Millière]]
- Full Title: Deep Learning and Synthetic Media
- Category: #articles
- Summary: insert summary
- My notes:
- Document Tags: [[toread]]
- Summary: Deep learning is transforming how synthetic audiovisual media, often called "deepfakes," are created, making production easier and more realistic. This technology challenges traditional distinctions between different types of media synthesis and opens up new creative possibilities. As deep learning advances, it blurs the lines between real and synthetic media, allowing for innovative forms of content that were previously impossible.
- URL: https://readwise.io/reader/document_raw_content/272633001
- Source File: Deep learning and synthetic media by Raphaël Millière.pdf
## LLM Chats
## NotebookLM
## LLM Audio
## Highlights
> Deep learning and synthetic media ([View Highlight](https://read.readwise.io/read/01jmpvvjytw6cp17d2ryqvrebb))
# Deep Learning and Synthetic Media

## Metadata
- Author: [[SpringerLink]]
- Full Title: Deep Learning and Synthetic Media
- Category: #articles
- Summary: insert summary
- My notes:
- Document Tags: [[toread]]
- Summary: Deep learning has transformed the creation of synthetic audiovisual media, making it easier to produce realistic sounds and images. This new form of media, often called "deepfakes," challenges traditional boundaries and definitions of media synthesis. As a result, we are seeing the emergence of novel kinds of interactive and controllable media that were not possible before.
- URL: https://link.springer.com/article/10.1007/s11229-022-03739-2
- Source File: Deep learning and synthetic media by Raphaël Millière.pdf
## LLM Chats
## NotebookLM
## LLM Audio
## Highlights
> An impressive and salient example of this progress can be found in so-called “deepfakes”, a portmanteau word formed from “deep learning” and “fake” (Tolosana et al., [2020](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR66)). This term originated in 2017, from the name of a Reddit user who developed a method based on DL to substitute the face of an actor or actress in pornographic videos with the face of a celebrity. However, since its introduction, the term “deepfake” has been generically applied to videos in which faces have been replaced or otherwise digitally altered with the help of DL algorithms, and even more broadly to any DL-based manipulations of sound, image and video. ([View Highlight](https://read.readwise.io/read/01jmpx1v3nfb2fvhtdjprp5km6))
> undermining the epistemic and testimonial value of photographic media and videos (Fallis, [2020](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR16); Rini, [2020](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR54)), ([View Highlight](https://read.readwise.io/read/01jmpx7m4y1cmth4brdask2jrr))
> Given enough training data and computational power, DL methods have proven remarkably effective at classification tasks, such as labeling images using many predetermined classes like “African elephant” or “burrito” (Krizhevsky et al., [2017](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR36)). However, the recent progress of DL has also expanded to the manipulation and synthesis of sound, image, and video, with so-called “deepfakes” (Tolosana et al., [2020](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR66)). As mentioned at the outset, I will mostly leave this label aside to focus on the more precise category of *DL-based synthetic audiovisual media*, or DLSAM for short. Most DLSAM are produced by a subclass of unsupervised DL algorithms called *deep generative models*. Any kind of observed data, such as speech or images, can be thought of as finite set of samples from an underlying probability distribution in a (typically high-dimensional) space. For example, the space of possible color images made of 512 \(\times \) 512 pixels has no less than 786,432 dimensions—three dimensions per pixel, one for each of the three channels of the RGB color space. Any given 512 \(\times \) 512 image can be thought of as a point within that high-dimensional space. Thus, all 512 \(\times \) 512 images of a given class, such as dog photographs, or real-world photographs in general, have a specific probability distribution within \(\mathbb {R}^{786432}\). Any specific 512 \(\times \) 512 image can be treated as a sample from an underlying distribution in high-dimensional pixel space. The same applies to other kinds of high-dimensional data, such as sounds or videos. ([View Highlight](https://read.readwise.io/read/01jmpxx6tdfkehjtf47wr1qn73))
- Tags: [[generating philosophy paper]] [[appreciation of ai]]
- Note: latent space
> The progress of image generation has been remarkable since the introduction of GANs in 2014. State-of-the-art GANs trained on domain-specific datasets, such as human faces, can now generate high-resolution photorealistic images of non-existent people, scenes, and objects.[Footnote 3](https://link.springer.com/article/10.1007/s11229-022-03739-2#Fn3) The resulting images are increasingly difficult to discriminate from real photographs, even for human faces, on which we are well-attuned to detecting anomalies.[Footnote 4](https://link.springer.com/article/10.1007/s11229-022-03739-2#Fn4) Other methods now achieve equally impressive results for more varied classes or higher-resolution outputs, such as diffusion models (Dhariwal & Nichol, [2021](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR14)) and Transformer models (Esser & Ommer, [2021](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR15)). ([View Highlight](https://read.readwise.io/read/01jmpysby9c311gzkjesccj8ar))
> The way in which deep generative models learn from data has deeper implications for the Continuity Question. According to the *manifold hypothesis*, real-world high-dimensional data tend to be concentrated in the vicinity of low-dimensional manifolds embedded in a high-dimensional space (Tenenbaum et al., [2000](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR63); Carlsson, [2009](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR6); Fefferman et al., [2016](https://link.springer.com/article/10.1007/s11229-022-03739-2#ref-CR17)). Mathematically, a manifold is a topological space that locally resembles Euclidean space; that is, any given point on the manifold has a neighborhood within which it appears to be Euclidean. A sphere is an example of a manifold in three-dimensional space: from any given point, it locally appears to be a two-dimensional plane, which is why it has taken humans so long to figure out that the earth is spherical rather than flat. ([View Highlight](https://read.readwise.io/read/01jmq5e6y1sm938fvgzd24agft))
## New highlights added May 6, 2025 at 10:34 AM
> These manipulations are tailored to a particular desired output. By contrast, DLSAM are sampled from a continuous latent space that has not been shaped by the desiderata of a single specific output, but by manifold learning. ([View Highlight](https://read.readwise.io/read/01jtjagy7jjjyh8885aaa5qsc8))