Skip Navigation

Microsoft’s VASA-1 can deepfake a person with one photo and one audio track

arstechnica.com

Microsoft’s VASA-1 can deepfake a person with one photo and one audio track

On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track. In the future, it could power virtual avatars that render locally and don't require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

17 comments
  • We're going to need strong digital signatures on everything, and we need it fast, else we won't be able to believe anything we see. It will be Steve Bannon's "flood the zone with shit" dream come true.

    • We’re going to need strong digital signatures on everything

      That won't help anything considering how easy it is to strip metadata.

  • That lip sync is scary good. It's still a little off, the teeth are weirdly stretchy, but nobody would notice it's a deepfake on first glance.

    Seems very similar to Nvidia's idea of only having a moving photo for video calls to reduce bandwidth needed. Very nice.

    • We'd need better optimization and more powerful processing on ye average laputopu for that to happen.

  • Revenge porn machine go brrrrr.

    Parents need to learn this stuff and teach their kids about it. Rumored nudes were enough to ruin kids lives at my highschool, nevermind "real" ones.

  • No. No, they can't. This shit still takes lots and lots of training data.

    It's just like any job. You can't just fully fake something in one day. At best, you might get 60% of the way there, maybe 80% after adding on some generic experience. But, you're not going to fully mimic anything without lots of training and experience.

  • This is the best summary I could come up with:


    On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track.

    In the future, it could power virtual avatars that render locally and don't require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

    To show off the model, Microsoft created a VASA-1 research page featuring many sample videos of the tool in action, including people singing and speaking in sync with pre-recorded audio tracks.

    The examples also include some more fanciful generations, such as Mona Lisa rapping to an audio track of Anne Hathaway performing a "Paparazzi" song on Conan O'Brien.

    While the Microsoft researchers tout potential positive applications like enhancing educational equity, improving accessibility, and providing therapeutic companionship, the technology could also easily be misused.

    "We are opposed to any behavior to create misleading or harmful contents of real persons, and are interested in applying our technique for advancing forgery detection," write the researchers.


    The original article contains 797 words, the summary contains 183 words. Saved 77%. I'm a bot and I'm open source!

17 comments