BTNR: Bayesian Tensor Networks for Regression
Transactions on Machine Learning Research (TMLR)
Abstract
Universal function approximation via basis expansions such as Fourier or polynomial bases is resource intensive. While the number of coefficients grows moderately in low dimensions, it can scale exponentially with the input size, leading to high complexity and potential overfitting. Tensor networks compress the coefficient space allowing for linear rather than exponential scaling. Existing tensor networks methods derive their learning equations for one specific network topology and few studies explore Bayesian learning methods. Probabilistic frameworks naturally derive confidence estimates for model predictions, allowing real-world applications to account for uncertainties in the decision process. We propose Bayesian Tensor Networks for Regression (BTNR), a variational Bayesian inference framework whose analytical update equations are topology agnostic. A single derivation applies to a large class of tensor network structures, independent of the feature map. BTNR also recovers standard ridge regularized alternating least squares (ALS) inference as a special case. Within the framework, automatic relevance determination (ARD) can prune irrelevant structure, removing the need for a part of the costly ablation over hyperparameters that classical model estimation requires. The Bayesian formulation further yields predictive uncertainty directly from the inferred posterior, without recourse to a separate estimation procedure. We benchmark BTNR on multivariate polynomial regression across four tensor network topologies and contrast it against ridge regularized ALS as well as non tensor network based standard Bayesian regression baselines. BTNR attains point estimates comparable to hyperparameter tuned ALS while reducing the hyperparameter search, and delivers well calibrated predictive uncertainty, reaching the lowest calibration error on several datasets with respect to the Bayesian baselines. Importantly, BTNR provides a universal foundation for exploring tensor network structures in a supervised learning context.
Introduction
Any function can be written as a linear combination of a suitable functional basis, but the space of coefficients grows exponentially with the input dimension. Regression with a degree-\(L\) polynomial feature map \(\Phi(\mathbf{x})\) seeks a coefficient tensor \(\mathbf{A}\) such that
\[ y_s = \langle\, \Phi(\mathbf{x}_s),\, \mathbf{A}\, \rangle + \epsilon_s , \]
where \(\langle\cdot,\cdot\rangle\) is a full tensor contraction and \(\epsilon_s\) is i.i.d. noise. The coefficient tensor \(\mathbf{A}\) has \(\mathcal{O}\!\left((d+1)^{L}\right)\) entries. Tensor networks can compress that tensor into a subspace that can grow only linearly on the degree depending on the structure. In this paper we want to overcome two limitations respect to traditional methods. First, non-probabilistic estimation (ALS, DMRG) demands an expensive hyperparameter search over bond dimensions and regularization strengths. Second, the learning equations are almost always derived for one chosen topology and do not trivially transfer to others. BTNR approaches both obstacles with a single topology-agnostic Bayesian derivation.
Methods
BTNR assumes a mean-field variational posterior over the network and derives update equations without structural assumptions about how the nodes are wired. Each node \(\mathbf{A}^{i}\) gets a multivariate Gaussian posterior, each edge a Gamma-distributed relevance parameter \(\lambda\), and the likelihood noise precision \(\tau\) is Gamma as well:
\[ q(\mathbf{A}^{i}) = \mathcal{N}\!\left(\boldsymbol{\mu}^{i}, \boldsymbol{\Sigma}^{i}\right), \qquad q(\lambda_{jk}) = \mathrm{Ga}\!\left(\alpha^{q}_{jk}, \beta^{q}_{jk}\right), \qquad q(\tau) = \mathrm{Ga}\!\left(\alpha^{q}_{\tau}, \beta^{q}_{\tau}\right). \]
Maximizing the evidence lower bound (ELBO) by moment matching gives closed-form updates that reference only the embedded input tensor \(\mathbf{X}_s\) and two derived networks: the mean network \(\boldsymbol{\mu}\) and the second moment network \(\mathbf{P}\). Because nothing depends on the topology, the same equations apply unchanged to any structure and any feature map.
The mean and second-moment networks
The whole derivation runs on two networks built from the weight network \(\mathbf{A}\). The mean network \(\boldsymbol{\mu}\) inherits \(\mathbf{A}\)'s topology exactly and contracts with the input \(\mathbf{X}_s\). The second-moment network \(\mathbf{P}\) has the same node structure but doubled edges, so each block carries two bond legs and two input legs; it contracts with \(\mathbf{X}_s \otimes \mathbf{X}_s\). Its nodes are simply
\[ \mathbf{P}^{i} = \mathbb{E}_{q}\!\left[\mathbf{A}^{i} \otimes \mathbf{A}^{i}\right] = \boldsymbol{\Sigma}^{i} + \boldsymbol{\mu}^{i} \otimes \boldsymbol{\mu}^{i}. \]
Update equations
Moment matching yields one closed-form update per group of parameters, applied in sweeps until the ELBO converges (it is guaranteed non-decreasing). For a node's covariance and mean,
\[ \boldsymbol{\Sigma}^{i\star} = \Bigl(\mathbb{E}[\tau]\textstyle\sum_{s} \hat{\mathbf{P}}^{i}\mathbf{X}^{P}_{s} + \tilde{\boldsymbol{\Gamma}}_{E(i)}\Bigr)^{-1}, \qquad \boldsymbol{\mu}^{i\star} = \mathbb{E}[\tau]\,\boldsymbol{\Sigma}^{i\star}\textstyle\sum_{s} y_s\,\hat{\boldsymbol{\mu}}^{i}\mathbf{X}^{\mu}_{s}. \]
Here \(\hat{\mathbf{P}}^{i}\) and \(\hat{\boldsymbol{\mu}}^{i}\) are the environments (the network with node \(i\) removed), and \(\tilde{\boldsymbol{\Gamma}}\) is the diagonal tensor of expected edge relevances that drives the ARD regularization. The edge and noise parameters have equally compact updates, e.g. the noise precision:
\[ \alpha^{q\star}_{\tau} = \alpha^{p}_{\tau} + \tfrac{S}{2}, \qquad \beta^{q\star}_{\tau} = \beta^{p}_{\tau} + \tfrac{1}{2}\sum_{s}\Bigl( \lVert y_s - \langle\boldsymbol{\mu}\,\mathbf{X}^{\mu}_{s}\rangle\rVert^{2} + \langle(\mathbf{P} - \boldsymbol{\mu}\otimes\boldsymbol{\mu})\,\mathbf{X}^{P}_{s}\rangle \Bigr). \]
Setting \(\boldsymbol{\Sigma} = 0\) and using a single uniform edge value collapses these into ridge-regularized ALS — the classical algorithm drops out of the same derivation as a special case.
Predictive uncertainty
The posterior gives predictive variance directly, split into an epistemic term (uncertainty about the model parameters) and an aleatoric term (the inferred data noise):
\[ \sigma^{2}_{E,s} = \langle(\mathbf{P} - \boldsymbol{\mu}\otimes\boldsymbol{\mu})\,\mathbf{X}^{P}_{s}\rangle, \qquad \sigma^{2}_{A} = \mathbb{E}\!\left[\tau^{-1}\right], \qquad \sigma^{2}_{s} = \sigma^{2}_{E,s} + \sigma^{2}_{A}. \]
No separate calibration procedure is needed: the intervals come straight from the inferred posterior.
Results
BTNR is benchmarked on multivariate polynomial regression across four tensor-network topologies including the binary tensor tree (BTT), matrix product state (MPS), canonic polyadic decomposition (CPD) and tensor ring (TR) against ridge-regularized ALS and non tensor networks Bayesian baselines (Gaussian processes, Bayesian deep ensembles, horseshoe BNNs, BASS).
- Point estimates on par with tuned ALS. With the same structure and basis, BTNR generally performs well against ALS \(R^2\) seemingly reaching well generalizing parameters automatically, without the ablation over bond dimensions and regularization strengths that ALS requires.
- Structure learned, not searched. Automatic relevance determination prunes irrelevant bond dimensions during training, removing the costly grid search over network bonds.
- Well-calibrated uncertainty. BTNR attains comparable expected calibration error (ECE) and stays competitive with non tensor networks models elsewhere, with predictive intervals of reliable size.
- Loop structures (the tensor ring) hold up well against loopless ones as a consequence of Mean Field approximation.
Conclusions
BTNR provides a single, topology agnostic Bayesian inference framework for tensor network regression. In our work we show this property by analyzing CPD, TT/MPS, MPOs, trees and rings. We recover ridge regularized ALS as a special case, while hyperparameter search of bond dimensions and regularization is removed by ARD-driven pruning. The model yields predictive uncertainty decomposed into epistemic and aleatoric parts from the posterior.
Limitations. The objective is non-convex, so there is no ELBO/optimality guarantee and hence no guarantee of reaching the optimal pruned structure; the full experimentations are currently restricted to a polynomial feature map.
Cite
@article{ciolli2026btnr,
title = {{BTNR}: Bayesian Tensor Networks for Regression},
author = {Niccol{\`o} Ciolli and Jesper L{\o}ve Hinrich and Morten M{\o}rup},
year = {2026},
journal = {Transactions on Machine Learning Research},
issn = {2835-8856},
url = {https://openreview.net/forum?id=5yPMm3hq7k},
}