Publication: Ghost mechanism: an analytical model of Abrupt learning in recurrent networks
| dc.contributor.coauthor | Dinc, F. | |
| dc.contributor.coauthor | Cirakman, E. | |
| dc.contributor.coauthor | Kurtkaya, B. | |
| dc.contributor.coauthor | Yuksekgonul, M. | |
| dc.contributor.coauthor | Jiang, Y. | |
| dc.contributor.coauthor | Schnitzer, M. J. | |
| dc.contributor.coauthor | Tanaka, H. | |
| dc.date.accessioned | 2026-08-14T11:23:59Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Abrupt learning, long performance plateaus followed by rapid convergence, is a common phenomenon in recurrent neural networks (RNNs) trained on working-memory tasks. In such cases, the networks develop transient slow regions in state space that extend the effective timescales of computation. However, the mechanisms driving sudden performance improvements and their causal role remain unclear, largely because we lack an analytical dynamical-systems framework. To address this gap, we introduce the ghost mechanism, a general process by which finite-dimensional continuous-time dynamical systems exhibit transient slowdown near the remnant of a saddle-node bifurcation. By reducing the high-dimensional dynamics near ghost points, we derive a one-dimensional canonical form that analytically captures learning as a process controlled by a single scale parameter. Using this model, we study a form of abrupt learning emerging from ghost points and identify a critical learning rate that scales as an inverse power law with the timescale of the learned computation. Beyond this rate, learning collapses through two interacting modes: (i) vanishing gradients and (ii) oscillatory gradients near minima. These features can lock the system into high confidence but incorrect predictions when parameter updates trigger a no-learning zone, a region of parameter space where gradients vanish. We validate these predictions in low-rank RNNs, where ghost points precede abrupt transitions and further demonstrate their generality in full-rank RNNs trained on canonical working-memory tasks. Our theory offers two approaches to address these learning difficulties: Increasing trainable ranks stabilizes learning trajectories, while reducing output confidence mitigates entrapment in no-learning zones. Overall, the ghost mechanism reveals how the computational demands of a task constrain the optimization landscape, demonstrating that well-known learning difficulties in RNNs partly arise from the dynamical systems they must learn to implement. | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | WOS | |
| dc.description.indexedby | Scopus | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | N/A | |
| dc.description.sponsorship | We would like to thank Dr. Marta Blanco-Pozo, Dr. Saeed Ahmed Khan, Dr. Sarah Kushner, and Dr. Daniel Kunin for their insightful comments on the manuscript; Dr. Nina Miolane, Abby Bertics, and members of the Geometric Intelligence Laboratory at UC Santa Barbara for their helpful feedback on an earlier version of the project; and Dr. Boris Shraiman for fruitful discussions on the loosely constrained parameters and learning in high-dimensional dynamical systems. M. J. S. gratefully acknowledges funding from the Simons Collaboration on the Global Brain, the Vannevar Bush Faculty Fellowship Program of the U.S. Department of Defense, and Howard Hughes Medical Institute. F. D. receives funding from Stanford University's Mind, Brain, Computation and Technology program, which is supported by the Stanford Wu Tsai Neuroscience Institute. Internships of E. C. and B. K. were supported in part by a grant from the Feldman-McClelland Open-a-Door fund of the Pittsburgh Foundation. B. K. thanks the Impact Scholarship and Research Scholarship for Critical Thinkers programs from the Bridge to Tuerkiye Funds for supporting his visit to Stanford University. Some of the computing for this project was performed on the Sherlock cluster. We would like to thank Stanford University and the Stanford Research Computing Center for providing computational resources and support that contributed to these research results. This research was supported in part by Grant No. NSF PHY-2309135 and the Gordon and Betty Moore Foundation | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | 97 | |
| dc.identifier.ScopusQuartile | Q1 | |
| dc.identifier.WoSPercentile | 96 | |
| dc.identifier.WoSQuartile | Q1 | |
| dc.identifier.doi | 10.1103/mjcl-lb4x | |
| dc.identifier.embargo | N/A | |
| dc.identifier.grantno | PHY-2309135 | |
| dc.identifier.grantno | 2919.02 | |
| dc.identifier.issn | 2160-3308 | |
| dc.identifier.issue | 2 | |
| dc.identifier.scopus | 2-s2.0-105043512203 | |
| dc.identifier.uri | http://doi.org/10.1103/mjcl-lb4x | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34457 | |
| dc.identifier.volume | 16 | |
| dc.identifier.wos | 001816914700001 | |
| dc.keywords | Mechanism (biology) | |
| dc.keywords | Computer science | |
| dc.keywords | Epistemology | |
| dc.keywords | Philosophy | |
| dc.language | eng | |
| dc.publisher | American Physical Society | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | Physical Review X | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Physical sciences | |
| dc.subject | Physics and astronomy | |
| dc.subject | Statistical and nonlinear physics | |
| dc.subject | Computer science | |
| dc.subject | Artificial intelligence | |
| dc.title | Ghost mechanism: an analytical model of Abrupt learning in recurrent networks | |
| dc.type | Journal Article | |
| dspace.entity.type | Publication |
