TNSPC: Learning from Partially Observed Data Using Tensor Network Structured Probabilistic Circuits
IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP), Istanbul
DOI: 10.1109/MLSP62443.2025.11204307 · Code
Abstract
Partially observed and heterogeneous data pose significant challenges in machine learning, often leading to biased results or reduced performance. Existing imputation methods typically struggle with high-dimensional, mixed data types. We introduce TNSPC, a framework combining tensor network structures with probabilistic circuits to give an analytically tractable, scalable solution for imputing such data. We contrast the canonical polyadic (CP), tensor train (TT), and tensor tree network (TTN) structures against prominent probabilistic circuits, finding TNSPCs generally comparable and, on some datasets, best performing.
The problem
Missing, mixed-type data can be found everywhere. Deletion assumes MCAR and adds bias; multiple imputation by regression can need up to \(2^M\) models (one per missingness pattern over \(M\) variables); deep methods (GAIN, VAEs, MIDAS) handle heterogeneity but are heavy and opaque.
The idea
Tensor decompositions and probabilistic circuits (PCs) are two views of the same sum–product object. TNSPC encodes the joint distribution as an order-\(M\) tensor (one mode per feature) and factorizes it into a valid PC. The payoff is exact, efficient marginalization: integrating out a variable touches only its leaf. That one property makes learning from missing data, forming conditionals, and imputing cheap once trained — with no separate model per missingness pattern.
The model splits into leaves (independent univariate distributions — Normal, Beta, Poisson, Bernoulli, Categorical such that one model handles heterogeneous features) and a core that combines them via sum/product nodes following a tensor-network structure. Parameters are fit with Expectation–Maximization; missing values are simply marginalized during training. Imputation fills each gap with the expectation of \(p(x_i \mid x_{\neg i})\).
Core structures
| Structure | Idea | Core params |
|---|---|---|
| CP | Flat mixture: one product node per channel then a sum node. | \(\propto C\) |
| TT | Alternating chain of sum/product nodes (matrix product state). | \(\propto C \cdot V\) |
| TTN | Binary tree built bottom-up over pairs of nodes. | \(\propto C^2 (V - 1)\) |
Results
On the 20-dataset discrete density-estimation benchmark (test log-likelihood, higher is better), CP and TT clearly beat TTN and Einsum. The simple CP is state of the art on bbc, c20ng, and nltcs; TT wins on dna competitive with dedicated PCs (LSPN, CNET, SoftLearn). On three real heterogeneous datasets (bike, adult, mozilla) imputation degrades gracefully as the missing fraction grows, though accuracy is feature dependent.
Conclusions
- One tractable joint model also enables density estimation, surrogate data generation, and supervised prediction.
- Imputations come from a full generative model, so each filled value is a principled conditional expectation rather than an ad-hoc guess.
- Best fit for MCAR/MAR tabular data; MNAR, images, and sequences are out of scope.
- TT/TTN are sensitive to feature ordering (unlike CP).
Cite
@INPROCEEDINGS{11204307,
author={Ciolli, Niccolò and Mørup, Morten and Schmidt, Mikkel N.},
booktitle={2025 IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP)},
title={TNSPC: Learning from Partially Observed Data Using Tensor Network Structured Probabilistic Circuits},
year={2025},
volume={},
number={},
pages={1-6},
keywords={Analytical models;Tensors;Statistical analysis;Circuits;Estimation;Machine learning;Probabilistic logic;Imputation;Data models;Integrated circuit modeling;tensor network;probabilistic circuits;imputation;heterogeneous;datasets;density},
doi={10.1109/MLSP62443.2025.11204307}}