Evento
#Seminarios DaSCI

Image and Video Generation using Deep Learning

Ponente: Stéphane Lathuilière es profesor asociado en Telecom París, Francia, en el equipo multimedia. Hasta octubre de 2019, fue becario de postdoctorado en la Universidad de Trento (Italia) en el Grupo de Multimedia y Comprensión Humana, dirigido por el Prof. Nicu Sebe y la Prof. Elisa Ricci. Recibió el título de Master en Matemáticas Aplicadas e Informática de la ENSIMAG, Instituto de Tecnología de Grenoble (Grenoble INP), Francia, en 2014. Realizó su tesis doctoral en el Instituto Internacional de Investigación MICA (Hanoi, Vietnam). Trabajó para obtener su doctorado en matemáticas e informática en el Equipo de Percepción de Inria bajo la supervisión del Dr. Radu Horaud, y lo obtuvo en la Universidad de Grenoble Alpes (Francia) en 2018. Sus intereses de investigación abarcan el aprendizaje de máquinas para problemas de visión por ordenador (por ejemplo, adaptación de dominios, aprendizaje continuo) y modelos profundos para la generación de imágenes y vídeos. Publica regularmente artículos en las conferencias más prestigiosas sobre visión por computador (CVPR, ICCV, ECCV, NeurIPS) y en las revistas más importantes (IEEE TPAMI).

Resumen (en inglés): Generating realistic images and videos has countless applications in different areas, ranging from photography technologies to e-commerce business. Recently, deep generative approaches have emerged as effective techniques for generation tasks. In this talk, we will first present the problem of pose-guided person image generation. Specifically, given an image of a person and a target pose, a new image of that person in the target pose is synthesized. We will show that important body-pose changes affect generation quality and that specific feature map deformations lead to better images. Then, we will present our recent framework for video generation. More precisely, our approach generates videos where an object in a source image is animated according to the motion of a driving video. In this task, we employ a motion representation based on keypoints that are learned in a self-supervised fashion. Therefore, our approach can animate any arbitrary object without using annotation or prior information about the specific object to animate.

GrabaciónImage and Video Generation using Deep Learning -Recording (in English)

Image and Video Generation using Deep Learning (Slides)