Mathematical Foundations of Polyphonic Music Generation via Structural Inductive Bias and Factorized Representations: A Case Study on Beethoven Piano Sonatas
Seo, Joonwon
Citations
Abstract
This dissertation addresses a fundamental challenge in AI music generation: the inability to produce coherent musical phrases that bridge local note patterns and global form—what I term the “Missing Middle” problem. Using Beethoven’s structurally complex piano sonatas as our testbed, I present a dual contribution that bridges rigorous mathematical theory with empirical innovation.
Our applied contribution begins with statistical verification of functional independence between Pitch and Hand attributes in musical data (NMI=0.1663). This empirical finding motivates Smart Embedding, a factorized architecture that reduces the number of embedding parameters by 48.30% (from 176 to 91) while achieving a 9.47% improvement in validation loss.
Our theoretical contribution provides mathematical justification for these empirical gains. Using Information Theory, I prove factorization results in negligible information loss (0.153 bits). Through Rademacher Complexity analysis, I establish a 28.09% tighter generalization bound. I further formalize the architecture as a structure-preserving functor using Category Theory. SVD analysis confirms that performance improvements stem from structural inductive bias rather than compression efficiency alone.
The dissertation includes a comprehensive data pipeline addressing positional bias (from 73.91% to 49.81% balanced distribution) and presents a designed expert listening study (N=53) for perceptual validation. This work demonstrates how principled mathematical frameworks can drive practical advances in AI music generation.
