Sparsely Distributed Representations
what?
Sparse Distributed Representations (SDRs) are high-dimensional binary vectors where only a very small fraction of the bits are active (set to 1) at any given time, and the semantic meaning of the data is distributed across these active bits. [1, 2, 3]
Inspired by how the mammalian brain and neocortex encode information via populations of neurons, SDRs serve as a foundational data structure in frameworks like Hierarchical Temporal Memory (HTM)
Core Properties of SDRs
Semantic Meaning: Unlike traditional computer codes (like ASCII) where bit positions are arbitrary, each bit in an SDR represents a specific elemental feature or attribute of the data.
Built-in Similarity (Overlap): Comparing two SDRs via bitwise overlap (counting shared active bits) instantly reveals how semantically similar two concepts are. https://discourse.numenta.org/t/sparse-distributed-representations/2150
High Capacity: Despite low activity levels (e.g., 40 active bits out of 2,048), the combinatorial math yields an astronomically massive number of unique representable patterns.
Robustness & Noise Tolerance: If an SDR is partially corrupted, missing bits, or superimposed with noise, the system can still accurately recognize the underlying pattern. see the Wikipedia article on SDRs
Union Operations: Combining multiple SDRs using a bitwise OR operation allows a system to represent multiple simultaneous predictions or composite concepts efficiently.
Benefits Over Dense Embeddings
Enhanced Interpretability: The sparse, structured nature of active bits makes it easier to inspect and understand what features trigger a representation. https://seanpedersen.github.io/posts/sparse-distributed-representations/
Energy and Storage Efficiency: Storing only the indices of active nonzero components significantly reduces memory footprints and processing overhead during associative retrieval. www.sciencedirect.com/science/article/pii/S0925231226005655