Speaker Abstracts

Abstracts of speaker talks are listed in alphabetical order. For a full schedule of presentations, please visit our Program & Schedule page.


Evolutionary Optimization Reveals Structural Constraints on Reservoir Architecture for Spatiotemporal Chaos

Nima Dehghani, Massachusetts Institute of Technology

Biological systems maintain function in fluctuating environments by transforming past stimulation into internal dynamical states that support future-oriented responses. Reservoir computing provides a computational analogue, but standard formulations often treat the recurrent substrate as a fixed random network and train only the readout. Here we ask how the substrate itself changes when reservoir architecture is placed under evolutionary selection for prediction. Using the Kuramoto–Sivashinsky equation as a testbed for spatiotemporal chaos, we evolved reservoirs over five construction hyperparameters: size, connectivity degree, spectral radius, input scaling, and readout regularization. Evolution reduced prediction error at the population level, extended the low-error forecast horizon, and organized the design space along a diminishing-return size–efficiency frontier. Structural analyses showed that evolved reservoirs remained within a conserved stochastic-block-model-like spectral envelope while refining low-eigenvalue modes, locking modularity to an intermediate band, and pruning connection cost within that band. Pareto analysis showed that elite reservoirs occupied a horizontal floor in the cost–modularity plane, indicating that accuracy and efficiency were achieved jointly rather than through a simple trade-off. These findings show that evolutionary optimization does not merely improve prediction, but exposes interpretable structural constraints on the recurrent substrate: it stabilizes a task-suitable dynamical class and refines the architectural degrees of freedom most relevant for prediction. Evolutionary reservoir computing therefore provides a bio-inspired framework for studying how predictive demands shape adaptive dynamical networks.


Distinct mechanisms underlying in-context learning in transformers

Cole Gibson, Princeton University

Modern distributed networks, notably transformers, acquire a remarkable ability (termed `in-context learning’) to adapt their computation to input statistics, such that a fixed network can be applied to data from a broad range of systems. Here, we provide a complete mechanistic characterization of this behavior in transformers trained on a finite set $\mathcal{S}$ of discrete Markov chains. The transformer displays four algorithmic phases, characterized by whether the network memorizes and generalizes, and whether it uses 1-point or 2-point statistics. We show that the four phases are implemented by multi-layer subcircuits that exemplify two qualitatively distinct mechanisms for implementing context-adaptive computations. Minimal models isolate the key features of both motifs. Memorization and generalization phases are delineated by two boundaries that depend on data diversity, $K = |\mathcal{S}|$. The first ($K_1^\ast$) is set by a kinetic competition between subcircuits and the second ($K_2^\ast$) is set by a representational bottleneck. A symmetry-constrained theory of a transformer’s training dynamics explains the sharp transition from 1-point to 2-point generalization and identifies key features of the loss landscape that allow the network to generalize. Put together, we show that transformers develop distinct subcircuits to implement in-context learning and identify conditions that favor certain mechanisms over others.


How evolution shapes the capacity to learn

Erin Hecht, Harvard University

Many of the behaviors that are most important to an organism’s ecological success are not innate: they must be acquired through experience. Yet the ability to acquire them can itself evolve. How does selection acting across generations alter the neural and developmental processes through which individuals learn?

I will use domestic dogs as a model for investigating this question. Dog lineages have undergone intense and recent selection for behaviors that nevertheless depend on learning, including herding, hunting, scent detection, communication with humans, and trained service work. Comparative neuroimaging across breeds, populations, and individual dogs provides evidence that this selection has altered brain systems involved in learning, motivation, fear, and social communication. In particular, recent work suggests that adaptation to the human environment has been accompanied by changes in cortical systems associated with trainability, while a white-matter pathway involved in processing human communication shows signatures of both evolutionary change and experience-dependent plasticity.

These findings suggest that evolution need not specify complex behaviors directly. Instead, selection can shape how organisms attend to, engage with, and learn from particular features of their environments, thereby making some developmental outcomes more likely than others. I will discuss how this perspective may help us understand the evolution of efficient and reliable learning, as well as the tradeoffs that accompany increased developmental plasticity.


An inductive bias for generalization in mouse olfactory learning

Venkatesh Murthy, Harvard University

Animals must generalize from limited experience, yet behavioral experiments in the laboratory setting rarely assess whether or how rapidly they generalize. This contrasts machine learning systems, where generalization is considered a fundamental test of learning, and emphasizes performance evaluation with new in-distribution or out-of-distribution examples. Here, we used an olfactory categorization task in mice to investigate rules of generalization. We trained mice to discriminate between two target odorants mixed with a variable number (0-14) of background odors. There are 32766 possible mixture stimuli to be classified in two categories, yet mice learn to generalize from as few as 8 unique mixtures. Analysis of individual variability revealed features in learning dynamics during training that predict performance in the generalization phase. A linear supervised learning algorithm could describe the generalization from few exemplars well, whereas nonlinear classifiers were necessary to explain memorization. Our experiments suggest that mice have an inductive bias towards generalization, and will memorize only when forced to do so. We are currently investigating neural population dynamics in the olfactory cortex during and after learning.


Using neuro-biomechanical simulations to probe neural control of learned skills

Bence P. Ölveczky, Harvard University

The goal of my lab is to decipher how the brain learns and controls motor skills. The standard approach dissects the underlying circuits, brain area by brain area, inferring function by relating neural recordings and perturbations within each area to behavior. This reductionist approach encounters fundamental problems in highly recurrent systems, where activity in any one node is shaped by the dynamics of the whole. It is further complicated by embodiment: motor circuit function is defined relative to the body it controls, not measurable aspects of behavior. I will describe how neuro-biomechanical simulation, accelerated by advances in AI and physics simulation, allows us to more readily relate neural activity to the computations motor circuits perform. I will also show that we can recapitulate motor skill learning in virtual rats, yielding a digital twin that learns what our rats learn but in a system where every circuit element is observable and configurable. While this does not, in itself, resolve the problem of how recurrent circuits control a body, it provides a substrate for generating and testing hypotheses about motor learning that would otherwise be out of reach.


How dendritic trees shape local learning and route credit

Houman Safaai, Harvard University

Biological neurons receive input through thousands of synapses on branching dendrites, yet artificial networks typically use point neurons trained by global backpropagation. Biological learning must instead rely on local signals. We ask how conductance and topology shape credit assignment, from local rules to anatomical routes. We first derive exact gradients for conductance-based dendritic networks with excitatory and inhibitory synapses. Each gradient factors into local eligibility presynaptic activity, driving force, and input resistance and a compartment error transported from the soma through the tree. Local learning thus becomes a feedback problem: can a limited somatic teaching signal approximate the dendritic error field? Simulations and model perturbations show that shunting changes path gains and can improve feedback fidelity in specific regimes; supplying exact transported errors narrows the gap to backpropagation. We then test these predictions in reconstructed anatomy. Across eight MICrONS reconstructions, sparse ancestry-defined routes captured nearly half of the modeled credit energy with less than 7% of dense-feedback wiring. Their capture approached a dense oracle and exceeded random paths, depth bins, and ancestry-shuffled controls in every cell. Across more than 100 modeled sites, focal shunts concentrated gradient changes among descendant synapses more than current-matched additive perturbations did, with somatic voltage and output error fixed. However, in a preliminary analysis, ancestry-defined routes did not outperform ancestry-shuffled routes when credit was derived from measured visual responses. Together, these studies suggest that topology can provide a sparse substrate for routing credit and conductance can regulate route gain, but anatomy alone does not ensure alignment with a particular learning problem.


Viability determines the structure and efficiency of motor practice

Nidhi Seethapathi, Massachusetts Institute of Technology

When learning a new motor skill, biological learners do not practice continuously. Instead, they break practice into attempts of variable duration, frequently abandoning a deteriorating movement well before it fails and starting afresh. This structure is universal across motor learning, seen in infants learning to walk, adults adapting to a force field, and adults manipulating objects, yet no existing account of motor learning explains it. Standard reinforcement learning and error-correction accounts treat failure as a cost to be minimized, not as a principle that organizes practice itself. We propose instead that biological learners structure practice by treating failure as a constraint: they continuously monitor whether task success remains reachable from their current state, and terminate and reset the attempt when it is not. We formalize this as viability-aware termination, and test it in physics-based simulations using a reinforcement learning modeling framework, against an alternative representing standard accounts, which has full access to the same information but treats failure only as a cost, without using it to structure practice. Across three motor tasks–manipulation, reaching, and locomotion–viability-aware termination is necessary to reproduce the empirical structure and timecourse of human practice, while the cost-only alternative is not. Strikingly, while the learning rule was only hypothesized to explain behavioral structure, it also yields substantially faster learning and more energy-efficient solutions than the alternative, an emergent functional benefit, without being designed for either. These results identify a task-general computational principle by which biological learners actively shape their own learning experience, and suggest that state-dependent termination, not just reward, is an under-explored lever for efficient learning.


Emergent Traveling Waves in Neural Circuits

Navid Shervani-Tabar, Massachusetts Institute of Technology

Artificial neural networks achieve striking performance but do not capture a prominent neural property. Cortical activity is organized by structured spatiotemporal dynamics, traveling waves (TW), which have been implicated in a wide range of functions. Existing computational models often rely on hand-crafted connectivity and imposed dynamics, offering insight into their impact but less into how the waves naturally emerge in biological circuits. Here, we found that TWs emerge in models under biologically plausible constraints (spatially organized, directionally biased connectivity). Under an empirical neural manifold constraint, these wiring principles naturally emerge to support traveling wave dynamics in the recurrent model. We further show that wave propagation provides a robust mechanism for maintaining working memory in the presence of visual distractors. We compared these model predictions to non-human primate prefrontal cortex recordings, revealing a similar mechanism. Together, these results advance our understanding of traveling waves as a substrate for cognition and offer a framework for mechanistic accounts of cortical computation.


Learning long-ranged structure in English: Code length and entropy dynamics during LLM training

Lindsay Smith, Princeton University

We use large language models (LLMs) to uncover long-ranged structure in English texts from a variety of sources. We test both open-weight models and a 1.7B-parameter LLM that we train ourselves, controlling the exact data the model sees during training. The conditional entropy and code length in many cases continue to decrease with context length at least to $N\sim 10^4$ characters, implying that there are direct dependencies or interactions across these distances. A corollary is that there are small but significant correlations between characters at these separations, as we show from the data independent of models. The distribution of code lengths reveals that the model becomes nearly certain of a growing fraction of characters at large $N$. Training reduces code length over the full range of context lengths, with only comparatively modest changes in the shape of its context-length dependence. Our results constrain efforts to build statistical physics models of LLMs and of language itself. Ongoing work analyzes the effect of training data, data ordering, and model initialization on the learning dynamics of the LLM during training.


Degraded early visual input shapes holistic perceptual organization in humans and deep networks

Lukas Vogelsang, Massachusetts Institute of Technology

While contemporary deep networks rely primarily on local image features, humans effortlessly integrate local elements into global configurations. Such perceptual organization is critical for structuring complex visual environments into gestalts. How do humans acquire this capacity? Here, we present converging experimental and computational evidence that the spatial and chromatic limitations of a newborn’s visual input may bias learning toward holistic organization, and that imposing analogous constraints confers similar benefits to deep networks. Experimental evidence derives from our work with a unique group of children from rural India who had been born blind and surgically gained sight late in childhood. As their retinas have matured by the time of surgery, they receive higher-fidelity input from sight onset, largely skipping the initial degradations of typical development. We longitudinally tracked the performance of 16 newly-sighted children on a Gestalt grouping task and found that, despite marked gains in other visual abilities, performance remained persistently poor and substantially below matched controls across all cues tested. Computational modeling paralleled this pattern. Convolutional neural networks trained on developmentally inspired degraded-to-nondegraded input trajectories and evaluated on a custom 4,200-image Gestalt stimulus set formed representations that more strongly encoded holistic organization in their final layers. Networks trained without this progression, akin to the experience of the late-sighted, failed to develop such organization, showing significantly lower silhouette scores for Gestalt-defined clusters. Together, the results on Gestalt-like organization presented here complement past tests of the hypothesis that initial degradations may be adaptive. Beyond advancing a developmental account, training with initially degraded inputs also offers a biologically inspired route to stronger global processing in machine vision systems.


Developmental temporal degradation as a scaffold for robust video classification

Marin Vogelsang, Massachusetts Institute of Technology

In marked contrast to deep neural networks, the human visual system is able to integrate information across extended temporal sequences. Here, we draw inspiration from visual development to examine the potential origin of this capacity and its role in conferring resilience against temporal perturbations. Infants begin visual experience with spatially and temporally degraded inputs, and recent work suggests that early spatial degradations may be adaptive, promoting larger receptive fields and improved generalization. Here, we test whether analogous immaturities in temporal vision similarly scaffold the emergence of robust representations along the time axis. We trained 3D convolutional neural networks on the temporally meaningful ‘Something-Something’ V2 action classification dataset, manipulating temporal fidelity via Gaussian temporal blur, with spatial blur as a control. Inspired by human development, we also included protocols in which degraded inputs were later followed by high-fidelity inputs. Initial exposure to temporal blur consistently shaped learned representations. First-layer receptive fields became tuned to markedly lower temporal frequencies, indicative of longer-range temporal integration. Notably, these properties persisted even after subsequent training on high temporal resolution inputs. The networks trained with initial temporal blur also generalized better across testing conditions and showed increased robustness to Gaussian noise uncorrelated in time and to local perturbations, including short-range frame-order flips. In contrast, spatial blur without temporal low-passing yielded minimal benefits. Together, these findings extend the ‘Adaptive Initial Degradation’ hypothesis of perceptual development into the time domain. Early degradations act as inductive biases toward holistic representations resilient to temporal perturbations, and low-to-high training offers a simple, biologically inspired procedure for improving robustness in deep networks.


What the connectome teaches us about lottery-ticket signs

Qingyang (Alice) Wang, Howard Hughes Medical Institute

The signs of connections shape the capacity of a network [Wang et al., ICCV 2023], how efficiently it learns [Wang et al., ICML 2023], and which sparse subnetworks—lottery tickets—can train [Zhou et al. 2019; Gadhikar & Burkholz 2024; Oh et al. 2025]. Yet finding winning tickets—and their signs—remains hard. Since the signs are what make a ticket train, we seek the sign configuration that gives a strong ticket high representation capacity. Prior work ties the optimal excitatory fraction to stimulus tuning [Rubin, Abbott & Sompolinsky 2017], but no method quantifies the capacity a sign configuration confers, independent of task or tuning. We introduce one: functional complexity, which counts the non-linearly-separable functions it can represent, isolating sign where participation ratio and spectral norm cannot. It is how we read out the signs a general lottery ticket would need—which no other measure can.
Biological connectomes are sparse yet capable, and keep their signs fixed for life. Shuffling the signs while holding the wiring fixed, we measure the functional complexity of each configuration: across the larval and adult Drosophila connectomes, capacity is maximized by abundant excitation—75–81% of neurons—with broadly connected inhibition. The arrangement, not just the proportion, drives this: assign E/I independently of degree, and the optimum collapses to ~50%. The real connectomes reach high functional complexity—0.66 (larva) and 0.85 (adult), among the highest sampled.
We applied the design principle to sparse networks. At 99% sparsity on Fashion-MNIST, a sparse CNN with fixed wiring—80% excitation, 4×-more-connected inhibition—reaches 75% accuracy in ~40% fewer epochs than unconstrained pruning; in the narrowest, deepest nets (width 16, depth 4) it holds ~75% validation accuracy where unconstrained pruning collapses to ~10%.
All in all, the connectome has much to teach about signing a lottery ticket.


Resource constraints externalize memory in brains and machines

Ruiyi Zhang, Carnegie Mellon University

Intelligent systems do not operate in isolation from the body, yet how physical actions support internal computation remains unclear. Here we show that primates and artificial agents recruit eye movements to maintain beliefs about hidden goals when internal computation is constrained. In a naturalistic navigation task, both systems spontaneously tracked invisible goals with their eyes, despite the absence of visual cues. This strategy emerged in artificial agents only when internal memory capacity was limited, mirroring resource constraints in biological neural systems. Restricting gaze caused a dramatic collapse in performance in both systems, while natural behavioral variability and causal perturbations established that oculomotor state directly shapes internal beliefs about hidden goals. Neural recordings from macaque parietal cortex revealed that goal representations are distributed across embodied and internal components, which together form a shared latent representational space across brains and artificial networks. Together, these findings identify eye movements as a functional external memory and reveal a convergent principle of intelligence: when internal computation is constrained, intelligent systems recruit the body to support cognition.