Citations
Abstract
Representation learning affects the performance of many machine learning systems. In problems such as blind source separation (BSS) and self-supervised learning (SSL), the goal is often to learn latent features that are useful for downstream tasks. These features may need to separate mixed signals, reduce dependence, capture hidden structure, or remain stable under changes in noise, sampling, or viewpoint. Existing methods often address these goals through specialized model architectures, strong assumptions about the data, or computationally expensive optimization. This dissertation develops a batching-based strategy for representation learning and source separation. While batches are often treated mainly as tools for computational efficiency, this work treats them as statistical objects that can reveal useful latent structure. By organizing samples into structured, random, or task-specific batches, the proposed approach extracts local statistical summaries that guide learning without requiring substantial changes to the underlying model architecture. The dissertation studies this idea across several applications. In BSS, structured batches are used to form local covariance matrices that support source recovery. In SSL, random sub-batches are used to reduce redundancy among learned features. In federated analysis, batch-based representations support source separation when data are distributed across sites. In hyperspectral reconstruction, task-specific material batches allow regularization to act on learned spectral representations and improve material reconstruction. Although these applications differ, they share the same principle: the way data or latent representations are grouped can shape the statistical information available to the learning method. Overall, this dissertation shows that batching can be more than a practical tool for computation. It can also serve as a mechanism for introducing structure into representation learning. By changing how data are grouped, compared, and summarized, the proposed strategy improves learned representations across multiple tasks without requiring more complex models.
