Abstract: The human species is engaged in a major scientific quest: to understand the neural mechanisms of human (primate) intelligence. Recent advances in multiple subfields of brain research suggest that the next key steps in this quest will be the construction of real-world, systemic-level network models that aim to abstract, emulate, and explain the primate neural mechanisms underlying natural intelligent behavior. In this talk, I will describe the history of how neuroscience, cognitive science, and computer science converged to create image-specific, computationally computable, deep neural network models aimed at abstracting, emulating, and adequately explaining the mechanisms of central visual object recognition in primates. Based on a wealth of primate neurophysiological and behavioral data, some of these network models are currently the most advanced (i.e., the most accurate) scientific theories of the internal mechanisms of primate ventral visual flow and how these mechanisms underpin the ability of humans and other primates to rapidly and accurately infer latent world content (e.g., object identity, position, pose, etc.) from the pixel set of most natural images. Although still far from complete, these cutting-edge scientific models already have many uses in brain science and beyond, and I will describe three recent examples from our team. First, I will describe recent neural and behavioral work comparing and contrasting object perception by primates with object perception by models in the context of an adversarial attack. Second, I will describe recent behavioral work comparing and contrasting object perception by primates with object perception by models in the context of rapid object learning. Third, if time permits, I will highlight the use of leading models to design patterns of light energy in the retina (i.e., personalized synthetic images) to precisely modulate neural activity deep in the brain. In our view, this is an exciting new avenue of potential clinical benefit to humans.
Speaker: Jim DiCarlo is Professor of Systems and Computational Neuroscience at the Massachusetts Institute of Technology. The main goal of his research team is to discover and artificially emulate the brain mechanisms underlying human visual intelligence. Over the past 20 years, DiCarlo and his collaborators have helped develop, using the non-human primate animal model organism, our contemporary engineering-level understanding of the neural mechanisms underlying the processing of visual information in the ventral visual stream – a complex series of interconnected brain areas – and how that processing underpins basic cognitive abilities such as object and face recognition. He and his collaborators aim to use these new scientific insights to guide the development of more robust computer vision (“AI”) systems, reveal new ways to beneficially modulate brain activity through modulating the images that reach our eyes, expose new methods to accelerate visual learning, lay the groundwork for new neural prostheses (brain-machine interfaces) to restore lost senses, and provide a scientific basis for understanding how sensory processing is altered in conditions such as agnosia, autism, and dyslexia. DiCarlo trained in biomedical engineering, medicine, systems neurophysiology and computer science at Northwestern (BSE), Johns Hopkins (MD/PhD) and Baylor College of Medicine (Postdoc). He was Director of the MIT Department of Cognitive and Brain Sciences from 2012-2021, and is currently Director of the MIT Quest for Intelligence (2021-present), where he and his leadership team work to advance interdisciplinary research at the interface of natural and artificial intelligence. DiCarlo is an Alfred P. Sloan Research Fellow, a Pew Scholar in Biomedical Sciences, a McKnight Scholar in Neuroscience, and an elected member of the American Academy of Arts & Sciences.
Recording: Reverse engineering of human brain mechanisms of object perception and learning