Topics Technology and the internet

How do deepfakes work?

A deepfake is a video, image or audio clip made by a neural network that has learned what a particular face or voice looks and sounds like, then generates new material that appears to show that person doing or saying something. The core trick is pattern learning: the software studies many examples, builds a statistical model of how a face moves or a voice sounds, and produces new frames or sound that fit the model.

What makes it interesting is how the models get good. One well-known approach pits two networks against each other: one makes fakes, the other tries to catch them, and both improve. Newer systems use a different method, called diffusion, that starts from noise and refines it into a clear image. Face swaps, lip-sync edits and cloned voices each work a little differently, and each leaves different clues behind.

An episode would walk through those methods in plain language, then turn to the harder questions: what the technology is used for, why detection is a moving target, and what habits help when you are unsure about a clip. bre's hosts are AI, so they can be wrong, and you can press Talk to ask them to slow down or explain a step again.

What a bre episode would cover

An outline of the episode bre would make for this question. Every episode is written fresh when you ask, so yours will differ.

  1. What a deepfake actually isThe word comes from deep learning, a kind of machine learning. We separate fully generated media from ordinary editing and from simple tricks like slowing a clip down.
  2. Learning a face from examplesHow a network studies many images of a person and builds a model of their features, expressions and angles, so it can reproduce them in new situations.
  3. Two networks, one contestThe generator and discriminator idea: one network forges, the other judges, and the back-and-forth pushes the fakes toward realism.
  4. Diffusion and newer methodsHow systems that start from noise and refine it have changed image and video generation, and why the field keeps shifting quickly.
  5. Cloned voices and lip syncHow a model can learn the sound of a voice from recordings, and how video can be adjusted so a mouth matches words that were never spoken.
  6. Spotting the seamsCommon clues like odd lighting, blurry edges or strange blinking, and why none is reliable alone. Detection tools exist, but it is an ongoing contest.
  7. Uses, harms and healthy skepticismLegitimate uses in film and accessibility, the real harms of fraud and harassment, and simple habits for checking where a clip came from.

How the episode might open

A sample exchange between two of bre’s AI hosts, bre and Arlo. Both are AI; this is written by AI, as every bre episode is.

  1. breAI host

    Okay, let's start with the word. Deepfake. It sounds like a spy gadget, but it's really just a neural network that learned what someone looks like and then drew them doing something new.

  2. ArloAI host

    Drew them. So it's a forgery.

  3. breAI host

    A forgery that learns. It looks at thousands of frames of one face, every angle, every expression, and builds a kind of statistical sense of how that face behaves.

  4. ArloAI host

    Statistical sense. Not a 3D model of the head.

  5. breAI host

    Mostly not, no. It's closer to a very well-trained instinct for what pixels should come next. Okay, here's the part nobody tells you: the early versions needed a lot of footage of the target.

  6. ArloAI host

    And now?

  7. breAI host

    Now some systems need far less. I don't know exactly how little, and it depends on the tool, so we'll be careful with that one.

  8. ArloAI host

    Good. Who counted that, though, when people say it only takes one photo?

  9. breAI host

    Fair. Often nobody did. We'll sort out what's been shown from what's just repeated online.

Questions people also ask

Are deepfakes always made with AI?
The term refers to media made with deep learning, a form of AI. Ordinary editing, like cutting a clip or retouching a photo, is not a deepfake. Some cheap fakes use simple tricks such as slowing audio or changing captions, which can mislead people without any AI at all.
Can deepfakes be detected?
Sometimes. Software and trained eyes can catch artifacts like inconsistent lighting, warped edges or unnatural audio. But generators keep improving, and detection tools can miss new methods. Checking the source of a clip and whether trusted outlets confirm it is often more reliable than judging pixels.
Are all deepfakes harmful?
No. The same technology is used in film effects, dubbing, accessibility tools and satire that is labeled as such. The harm comes from deception and lack of consent, such as fraud, impersonation or fake explicit images. Laws on this vary by place and keep changing.
How is a deepfake voice made?
A model is trained on recordings of a person speaking until it captures features like pitch, rhythm and tone. It can then turn typed text, or another person's speech, into audio that sounds like that voice. Quality varies, and background noise or odd pacing can give it away.

Related topics

More: all 300 topics, technology and the internet, or the longer reads on /learn.

bre’s hosts are AI, and every episode is generated, so they can be wrong: check anything that matters. This page outlines what an episode would cover. It is for interest and learning, not medical, financial or legal advice.