Normal Approximation of Random Gaussian Neural Networks
Abstract
In this paper, we provide explicit upper bounds on some distances between the (law of the) output of a random Gaussian neural network and (the law of) a random Gaussian vector. Our main results concern deep random Gaussian neural networks with a rather general activation function. The upper bounds show how the widths of the layers, the activation function, and other architecture parameters affect the Gaussian approximation of the output. Our techniques, relying on Stein’s method and integration by parts formulas for the Gaussian law, yield estimates on distances that are indeed integral probability metrics and include the convex distance. This latter metric is defined by testing against indicator functions of measurable convex sets and so allows for accurate estimates of the probability that the output is localized in some region of the space, which is an aspect of a significant interest both from a practitioner’s and a theorist’s perspective. We illustrated our results by some numerical examples.
Funding: This research was supported by the European Union’s Horizon 2020 research project WARIFA under grant agreement no. 101017385, by the PRIN project 2022 “Variational Analysis of Complex Systems in Materials Science, Physics and Biology” (CUP B53D23009290006), and by the INdAM project “Modelli ed Algoritmi per dati ad elevata dimensionalità” (CUP E53C23001670001).
1. Introduction
This work is part of the literature studying random neural networks (NNs for short), that is, NNs whose biases and weights are random variables. In the context of modern deep learning, the interest in these types of networks is twofold; on the one hand, they naturally constitute a prior in a Bayesian approach, and on the other hand they may represent the initialization of gradient flows in empirical risk minimization. See Roberts et al. (2022) for a general reference on the subject.
Within the boundaries of this topic, many contributions in the literature have been handling the asymptotic Gaussianity of random NNs, as the number of neurons in the hidden layers tends to infinity. A seminal paper is Neal (1996), where the output of a shallow (i.e., having a single hidden layer) random NN, viewed as a stochastic process on the sphere, is shown to converge to a Gaussian process, as the number of neurons in the hidden layer grows large. From that point onward, many sophisticated results have been published for deep (i.e., having more than one hidden layer) random NNs at the large width limit. Early contributions in this direction were from Lee et al. (2018), Matthews et al. (2018), Yaida (2019), and Yang (2019). These achievements were extended in Hanin (2023), where it was proved that the output of a deep random NN, with Gaussian biases and weights, viewed as a random element on the space of continuous functions on a compact set, converges to a Gaussian process as the number of neurons in all the hidden layers tends to infinity. Large and moderate deviations of the output of a deep random NN, with Gaussian biases and weights, were studied in Macci et al. (2024).
Recently, the problem of the quantitative Gaussian approximation of the output of a random NN has received a lot of attention. For instance, Eldan et al. (2021), exploiting Wasserstein distances, provided quantitative versions of the results in Neal (1996) when the activation function was polynomial, ReLU, and hyperbolic tangent. We emphasize that the shallow random NN model considered in Neal (1996) and Eldan et al. (2021) has weights on the outer layer given by Rademacher random variables and weights on the inner layer distributed according to the Gaussian law (see Remark 3.4). Slightly different models were investigated in Klukowski (2022) and Cammarota et al. (2024). Indeed, as the number of neurons on the hidden layer grew large, Klukowski (2022) provided a quantitative functional central limit theorem for a shallow random NN model with input variables still on the sphere and weights on the outer layer still given by Rademacher random variables but weights on the inner layer uniformly distributed on the sphere, whereas as the number of neurons on the hidden layer grew large, Cammarota et al. (2024) gave quantitative functional central limit theorems for a shallow random NN model with weights on outer and inner layers distributed according to the Gaussian law and again input variables on the sphere. A quantitative functional central limit theorem for deep random NNs, with input variables on the sphere and Lipschitz continuous activation functions, was proved in Balasubramanian et al. (2024); this paper shares with our work the idea to apply Stein’s method for the Gaussian approximation in the context of deep NNs, albeit in a different mathematical setting.
A significant achievement is provided in Basteri and Trevisan (2024), where, for the first time in the literature, a quantitative proof of the Gaussian behavior of the output of a deep random Gaussian NN (i.e., a random NN whose biases and weights are Gaussian distributed) with a Lipschitz continuous activation function was given; in Basteri and Trevisan (2024), the distance from Gaussianity was measured by means of the 2-Wasserstein metric, which comes from the Monge-Kantorovich problem with quadratic cost. As far as shallow random Gaussian NNs with univariate output is concerned, we mention Bordino et al. (2024), which provided quantitative bounds on the Kolmogorov, the total variation and the 1-Wasserstein distances between the output and a Gaussian random variable, when the activation function was sufficiently smooth and had a sub-polynomial growth. A special mention is deserved for the independently written paper Favaro et al. (2023), where Stein’s method was used to obtain tight probabilistic bounds for various distances between the output (and its derivatives) of a deep random Gaussian NN and a Gaussian random vector. We refer the reader to Remarks 4.2, 6.2, and 6.4 for comparisons between our results and the corresponding achievements in Favaro et al. (2023).
The main contribution of our paper concerns the Gaussian approximation of the output of deep random Gaussian NNs in the convex and 1-Wasserstein distances under mild assumptions on the activation function (which, differently from Basteri and Trevisan (2024), can be non-Lipschitz). A specialization of these results clearly provides approximations for shallow random Gaussian NN with single input and real valued (i.e., univariate) output. However, in this specific case we furnish direct proofs, which (for various technical reasons) give the same rates under more general assumptions on the activation function.
For shallow random Gaussian NNs with univariate output, combining the Stein method for the Gaussian approximation with the integration by parts formula for the Gaussian law, we provide explicit bounds for the Kolmogorov, the total variation and the 1-Wasserstein distances between the output and a Gaussian random variable, under a minimal assumption on the activation function (see Theorem 4.1). Remarkably, we obtain the same rate of convergence as in Bordino et al. (2024) as the number of neurons in the hidden layer grows large, our constants being presumably better than the ones in Bordino et al. (2024) (see Table 1). For deep random Gaussian NNs, the novelty of our results is that we measure the error in the Gaussian approximation of the output in terms of the convex distance (see Theorem 6.1) and of the 1-Wasserstein distance (see Theorem 6.3) for a class of activation functions that strictly includes the family of Lipschitz continuous functions if either the biases are non null or the activation function vanishes at 0 (see Proposition 5.2). Remarkably, for both the convex and the 1-Wasserstein distances the rate of convergence that we obtain is of the same order as the one in Basteri and Trevisan (2024) because the number of neurons in all the hidden layers tend to infinity. The proofs of Theorems 6.1 and 6.3 are based on the Stein method for the multivariate Gaussian approximation and the integration by parts formula for the multivariate Gaussian law. The presence of more than one hidden layer complicates the derivations of the bounds, which rely on a key estimate for the L2-distance between the so-called collective observables and their limiting values (see Theorem 5.1). We emphasize that, when considering the convex distance, an expedient tool is provided by a smoothing lemma that we borrow from Schulte and Yukich (2019).
|
Table 1. Values of the Constants Given in Theorems 3.1 and 4.1
| Theorem 3.1 | 5.05 | 2.52 | 2.01 |
| Theorem 4.1 | 1.68 | 0.84 | 0.67 |
Note. Here, , γ = 3, , L = 1, , and x = 1.
It is well known that localizing the output of a random NN, that is, having a control over the probability that the output lies in a region of the space (belonging to a large class of measurable sets), is of valuable interest for practitioners. From a theorist’s point of view, the output distribution of a NN is often analytically untractable, and computing the probability that the output belongs to some measurable set results in performing a “heroic” mathematical integration (see, e.g., Roberts et al. 2022, p. 49). Our Theorems 4.1(ii) and 6.1 offer some insights into the localization problem in a simple and efficient way for both the univariate output of a shallow random Gaussian NN and the output of a deep random Gaussian NN, respectively. We refer the reader to Section 7 for some numerical illustrations of this issue.
The paper is organized as follows. In Section 2, we introduce our toolkit such as the integral probability metrics considered in the paper and some preliminaries on the Stein method. In Section 3, first we introduce all of the NNs considered in this work and then give a brief overview of the main results in Basteri and Trevisan (2024) and Bordino et al. (2024), comparing them with our achievements. In Section 4, we give upper bounds for the Kolmogorov, the total variation and the 1-Wasserstein distances between the (univariate) output of a shallow random Gaussian NN and a Gaussian random variable. In Section 5, we prove the aforementioned key estimate on the L2-distance between the collective observables and their limiting values. In Section 6, we furnish explicit upper bounds on the convex and the 1-Wasserstein distances between the output of a deep random Gaussian NN and a Gaussian random vector. Finally, in Section 7, we present some numerical illustrations concerning the above-mentioned issue of the output localization.
2. Preliminaries
In the present section, we introduce some notation, and we recall some results that will be of use throughout the paper.
2.1. Distances Between Probability Measures
In this paper, we consider various distances between probability measures on : the total variation distance, the convex distance, the Komogorov distance, and the p-Wasserstein distances. Hereon, the symbol denotes the Euclidean norm on .
The total variation distance between the laws of two -valued random vectors and , written , is given by
The convex distance between the laws of two -valued random vectors and , written , is given by
The Kolmogorov distance between the laws of two -valued random vectors and , written , is given by
For , the p-Wasserstein distance between the laws of two -valued random vectors and , written , is given by
Clearly, by Jensen’s inequality, , and it follows directly by the definitions that . Furthermore, for all , if , as , where , and are random vectors with values in , then converges in law to , as (see, e.g., Villani 2009 and Nourdin and Peccati 2012).
In view of the Kantorovich-Rubinstein duality (see theorem 5.10 and equation (5.11) in Villani 2009), the 1-Wasserstein distance between the laws of two -valued random vectors and such that satisfies the relation
Because dc is defined by testing against indicator functions of Borel convex sets rather than arbitrary Borel sets, the convex distance can be expected to be estimated more easily than the total variation distance; moreover, the convex distance looks more flexible than the Kolmogorov distance; for example, it enjoys a number of invariance properties not satisfied by dK (see Benktus 2003).
As for the relation between the convex distance and the optimal transport metric , it turns out that the convex distance to a fixed centered Gaussian law is bounded from above by a multiple of the square root of the 1-Wasserstein distance. More precisely, one has the following Proposition 2.5, which is proved in Nourdin et al. (2022).
Here and henceforth, we denote by , a centered Gaussian vector with invertible covariance matrix .
For any d-dimensional random vector , we have
We note that is an isoperimetric constant that satisfies the following relation (see Nazarov 2004):
2.2. The One-Dimensional Stein Equation
Throughout this paper, we denote by the one-dimensional Gaussian law with mean and variance and let .
The celebrated Stein equation for the one-dimensional Normal approximation Stein (1972) is given by
Hereon, for a Lipschitz continuous function , we denote by the Lipschitz constant of g, and for a function we denote by the supremum norm of f.
The following claims hold:
2.3. The Multidimensional Stein Equation
Throughout this paper, given a sufficiently smooth function , we define
Let , be the set of d × d real matrices. For a function , we denote by the Hessian matrix of f at and by the operator norm on , that is, for any . We consider the Hilbert-Schmidt inner product and the Hilbert-Schmidt norm on , which are defined, respectively, by
The Stein equation for multivariate Normal approximation is defined as
The following lemmas provide solutions to Stein’s Equation (5) under different assumptions on g.
Let . Then, the function
For measurable and bounded, define the smoothed function
For any , the function
is such that solves (5) with gt in place of g, and
(8)For any d-dimensional random vector , it holds
See proposition 4.3.2 in Nourdin and Peccati (2012) for Lemma 2.7; in particular, in Nourdin and Peccati (2012) it is noticed that
2.4. The Smoothing Lemma and the Integration by Parts Formula for Gaussian Random Vectors
We state a remarkable smoothing lemma for the convex distance proved in Schulte and Yukich (2019); see lemma 2.2 therein. It plays a crucial role in the proof of the Normal approximation of the output of a deep random Gaussian NN in the metric dc; see Theorem 6.1.
Let be a d-dimensional random vector. Then, for any ,
We recall the Gaussian integration by parts formula (we refer the reader to exercise 3.1.4 in Nourdin and Peccati (2012) for the Part (i) of Lemma 2.10 and to exercise 3.1.5 in Nourdin and Peccati (2012) for the Part (ii) of Lemma 2.10).
The following claims hold:
, if and only if, for any differentiable function such that , we have .
Let with bounded first partial derivatives. Then,
This relation holds true even if Σ is not positive definite.
3. Random Neural Networks
We let , we take L + 2 positive integers , and we fix a function . A fully connected NN of depth L with input dimension n0, output dimension , hidden layer widths , and nonlinearity σ is a mapping
We say that the neural network is a (fully connected and) deep random Gaussian neural network denoted by
NNs of depth L = 1 are called shallow NNs. We will denote shallow NNs (respectively, shallow random Gaussian NNs) by (respectively, by ).
Throughout this paper, we will also consider NNs with univariate output, that is, NNs with .
Consider a deep random Gaussian neural network . It turns out that the random variables , are independent and identically distributed with
For , let be the σ-field generated by the random variables
By construction, for any fixed , given , the random variables , are independent and Gaussian (as linear combination of independent Gaussian random variables). A straightforward computation yields
Setting , we define the quantities
The random variable is often referred to as collective observable at layer ; see, for example, Roberts et al. (2022) and Hanin (2023).
3.1. Some Related Literature
Consider the output of a deep random Gaussian neural network , and let
It follows from theorem 1.2 in Hanin (2023) (which indeed, more generally, establishes a functional weak convergence) that, if σ is continuous and polynomially bounded, then,
The following result for shallow random Gaussian NNs was proved in Bordino et al. (2024); see theorem 3.2 therein.
Let be a shallow random Gaussian NN with univariate output. If
In Section 4, we will give bounds on the quantities , of order , as , under a minimal assumption on the activation function (note that Condition (14) excludes the important case of the ReLU function, that is, ); see Theorem 4.1. In Section 6, we will give two general bounds on and for deep random Gaussian NNs; see Theorems 6.1 and 6.3. When specialized to shallow random Gaussian neural networks , they provide computable bounds, respectively, on and of order , as , see theorem 3.3 in Bordino et al. (2024) for a related result.
The first result in the literature that quantifies the convergence in distribution (13) with is given in Basteri and Trevisan (2024), where the following theorem has been proven.
Let be a deep random Gaussian NN, and suppose that the activation function σ is Lipschitz continuous. Then,
The next corollary is an immediate consequence of Theorem 3.2, the fact that , Proposition 2.5, and (3).
Let the assumptions and notation of Theorem 3.2 prevail. Then,
Theorems 6.1 and 6.3 will provide, under alternate assumptions on σ (see by Proposition 5.2(ii)), explicit bounds, respectively, on and that are of the same order of the bound in (16), as . In particular, the bound on the convex distance of Theorem 6.1 considerably improves the one in (17). We emphasize that to compare the results in Basteri and Trevisan (2024) with our achievements, we specialized those results to the case of a single input . In fact, the bound (15) was proven in Basteri and Trevisan (2024) for multiple inputs, and consequently the bound (17) can be stated for multiple inputs too.
Although not strictly related to our results, the pioneering papers Eldan et al. (2021) and Neal (1996) deserve a special mention. In Neal (1996) the author considered a random shallow with univariate output, , for all , , independent with law and , independent, identically distributed with law and independent of the random variables . It was proven in Neal (1996) that there exists a Gaussian process G on (the unit sphere in ) such that the process converges in distribution to G, as . Quantitative versions of this result (for various Wasserstein metrics and some specific choices of σ) are provided in Eldan et al. (2021).
4. Normal Approximation of Shallow Random Gaussian NNs with Univariate Output
The following theorem holds.
Let be a shallow random Gaussian NN with univariate output, and assume that the activation function σ is such that
Then:
Note that the bound on the Kolmogorov distance is better than the one that can be obtained using the relation .
Remarkably, theorem 3.3 in Favaro et al. (2023) shows that if is a shallow random Gaussian NN with univariate output and the activation function σ is polynomially bounded to order (see definition 2.1 in Favaro et al. (2023)), then there exist two constants such that
Although this inequality shows the optimality of the rate , because the constants are not provided in closed form, it cannot be directly used for the purpose of output localization (see Section 7). Moreover, the assumption (18) on σ does not require any regularity of the activation function.
We prove the three bounds (i), (ii), and (iii) separately by the Stein method.
Consider the Stein Equation (4) with
Let fg be the unique solution of the Stein equation (see Lemma 2.6(i)). Then, for any ,
Taking the expectation, we have
By Lemma 2.10(i), we have
Setting
Taking the modulus in (21) and then using that uniformly in (see Lemma 2.6(i)) and that , we have
Consider the Stein Equation (4) with , where is a Borel set. Let fg be the unique solution of the Stein equation (see Lemma 2.6(ii)). Then,
Taking the expectation and arguing as in (19), we have
Along similar computations as for (21), we have
Taking the modulus on this relation and then using that (see Lemma 2.6(ii)) and that , we have
Consider the Stein Equation (4) with , where is Lipschitz continuous with . Let fg be the unique solution of the Stein equation (see Lemma 2.6(iii)). Then,
Taking the expectation and arguing as in (19), we have
Along similar computations as for (21), we have
Taking the modulus on this relation and then using that (see Lemma 2.6(iii)) and that , we have
Note that both Theorem 3.1 and Theorem 4.1 provide bounds on , with a common rate , but different constants. In Table 1, we compare those constants in a special case. We observe that the constants given by Theorem 4.1 are more effective than those in Bordino et al. (2024). We also note that Condition (18) is satisfied by the ReLU activation function, whereas the assumptions of Theorem 3.1 do not hold for the ReLU.
5. A Key Estimate for the Collective Observables
The next theorem provides an estimate for the L2-norm of the random variable . In Section 6, such an estimate will play a crucial role in the proofs of the results on the Normal approximation of the output of a deep random Gaussian NN both in the convex and in the 1-Wasserstein distances (see Theorems 6.1 and 6.3).
Hereon, we denote by the L2-norm of a real-valued random variable Y.
Let be a deep random Gaussian NN. Suppose that the activation function σ is such that:
For any , there exists a polynomial
with nonnegative coefficients dependent only on and degree independent of , such that
(22)For any .
Then, for any , we have
The proof of the theorem is given later on in this section. We proceed stating a proposition and a remark, which clarify the generality of our assumptions on the activation function σ.
The following statements hold:
If σ is the perceptron function, that is, , then it satisfies Conditions (i) and (ii) of Theorem 5.1.
If σ is Lipschitz continuous and , then σ satisfies Conditions (i) and (ii) of Theorem 5.1. In particular, Condition (i) holds with
(24)If σ is Lipschitz continuous and , then σ satisfies Conditions (i) and (ii) of Theorem 5.1. In particular, Condition (i) holds with
(25)If σ is such that is Lipschitz continuous and , then σ satisfies Conditions (i) and (ii) of Theorem 5.1. In particular, Condition (i) holds with
(26)
The proof of the proposition is given later on in this section.
As a consequence of Proposition 5.2, we have that the most common activation functions satisfy the assumptions of Theorem 5.1. Indeed, one can easily prove that the ReLU function , the hyperbolic tangent function , the function , and the function are Lipschitz continuous and equal to zero in zero. Moreover, if , then also the sigmoid function and the softplus function satisfy the assumptions of Theorem 5.1, and indeed they are Lipschitz continuous. We emphasize that, if either the biases are nonnull or the activation function vanishes at 0, then the conditions on the activation function of Theorem 5.1 are more general than the one required in Basteri and Trevisan (2024), where is assumed Lipschitz continuous (see Theorem 3.2 and Proposition 5.2 parts (ii) and (iii)). We also emphasize that the conditions on the activation function of Theorem 5.1 are satisfied by the perceptron function that is noncontinuous and therefore non-Lipschitz (see Proposition 5.2(i). Another non-Lipschitz function that satisfies the conditions of Theorem 5.1 when the biases are nonnull I, for example, . Indeed, is the ReLU function and therefore Lipschitz continuous (see Proposition 5.2(iv)).
We consider separately the cases of the first hidden layer and that of the following ones.
Case .
Because the random variables , are independent and identically distributed with law , we have
Note that this latter term is finite because of the assumption (ii).
Case .
Take . We have already noticed that, given , the random variables are independent with Gaussian law with mean zero and variance
Therefore, letting denote a standard Gaussian random variable, independent of , we have
By assumption (i) and the fact that is independent of , we have
Therefore,
Note that
By (29), we have
By assumption (i), we have
Letting denote the law of a random variable X and again using (29), we have that the random variable has the same law as the random variable
Therefore, by (33) and Jensen’s inequality, we have
By assumption (i), it then follows that
Iterating this inequality, we have
Consequently,
Combining this inequality with (30) and using the elementary relation , for any , we have
Now, we use the relation (35) to prove (36) by induction on . Taking in (35), we have
Then, by (35), we have
The claim immediately follows noticing that, because of the nonnegativity of the quantities and , we have , for any .
Because σ is Lipschitz continuous, we have
Multiplying and dividing the term in (37) by , we easily have that
The proof is similar to the proof of (ii) and therefore is omitted.
If is Lipschitz continuous, then
Therefore, Condition (ii) of Theorem 5.1 is an immediate consequence of this latter relation and the fact that the absolute moments of Z are finite. □
6. Normal Approximation of Deep Random Gaussian NNs
6.1. Normal Approximation of Deep Random Gaussian NNs in the Convex Distance
The following theorem holds.
Let be a deep random Gaussian NN, and let the notation and assumptions of Theorem 5.1 prevail. Then,
Here, , and clearly the upper bound with constant C2 holds under the additional assumption that the biases are nonnull.
Let be a deep random Gaussian NN. Theorem 3.5 in Favaro et al. (2023) shows that if
We stress that the bounds in Theorem 6.1 provide (under different assumptions on σ and without any condition on the widths of the hidden layers) a detailed description of the analytical dependence of the upper estimate on the parameters of the model.
Throughout this proof, for ease of notation, we put
We preliminarily note that it suffices to prove that
Indeed, the claim then follows by Theorem 5.1, noticing that
If , then the inequality (41) holds because (which follows by the definition of the convex distance) and
From now on, we assume . Let , where C is an arbitrarily fixed measurable convex set in , and let
For any , by Lemma 2.8(i),
By Lemma 2.9, it follows that
Without loss of generality, hereafter we assume that is independent of . Therefore, by (43), we have
By Lemma 2.8(i), for any , the mapping is in and has bounded first-order derivatives. Then, because is a centered Gaussian random vector with covariance matrix
On combining this relation with (45) and taking the expectation, we have
By Lemma 2.8(ii), for any , we have
Moreover,
On combining these relations and using the elementary inequality , we have
On combining this latter inequality with (44), we have
Because , by this relation, we have
Setting in this latter inequality (note that this choice of the parameter t is admissible because ), we have
Taking the square root and multiplying by , we have
Because , we have
By (50) with , (52), and (53), we finally have (41); indeed,
The proof is completed. □
6.2. Normal Approximation of Deep Random Gaussian NNs in the 1-Wasserstein Distance
The following theorem holds.
Let be a deep random Gaussian NN, and let the notation and assumptions of Theorem 5.1 prevail. Then,
Here again, the upper estimate with the constant K2 holds if .
Let be a deep random Gassian NN with univariate output, widths of the hidden layers satisfying (40), and a polynomially bounded to order activation function σ. Then, by theorem 3.3 in Favaro et al. (2023), we have that there exist two constants such that
Clearly, this inequality shows the optimality of the rate . Here again, we note that the corresponding bound provided by Theorem 6.3 gives (under different assumptions on σ and without any condition on the widths of the hidden layers) a detailed description of the analytical dependence of the upper estimate on the parameters of the model.
Let be arbitrarily fixed. Without loss of generality, we assume that is independent of . By Lemma 2.7, we then have
Again, by Lemma 2.7, we have that the mapping is in and has bounded first-order derivatives. Applying Lemma 2.10(ii) exactly as in the proof of Theorem 6.1 (see a few lines before Equation (47)), we have
By Lemma 2.7, we have
On combining these relations with (49), we have
The claim follows taking the supremum over g on this inequality and then using Theorem 5.1, relation (42), and the fact that . □
7. Localization of the Output
In this section, we want to show the potentiality of the obtained results for practical applications. Indeed, both Theorem 4.1(ii) and Theorem 6.1 allow one to explicitly estimate the probability that the output of a random Gaussian NN evaluated at the input belongs to a suitable set without resorting to computationally expensive Monte Carlo simulations. This is what we call output localization, which can suggest a suitable architecture design to estimate a target function f. Indeed, in statistical learning, given a training set
Hence, in the following we want to show how to use the upper bound on the total variation distance provided in Theorem 4.1(ii), for a shallow random Gaussian NN with univariate output, and the upper bound on the convex distance provided by Theorem 6.1, for a deep random Gaussian NN, to localize the output. Indeed, given a measurable convex set , by the definitions of both the total variation and the convex distances, we have
Let be a rectangle of . Because is an -dimensional centered Gaussian random vector with covariance matrix (12), we have
Now, we furnish numerical values of the constant Cbound in (56) and (57); in both cases, we take , that is, a ReLU activation function. Because the ReLU function is Lipschitz continuous with Lipschitz constant equal to 1, and , by Proposition 5.2(ii), we have
By the expression of the constants in Theorem 5.1, we have
By the definition of the quantities , (see (10)), we have
Table 2 gives the values of Cbound in (56) in the case of a shallow random Gaussian NN with architecture L = 1, , and for different values of . Table 3 gives the values of Cbound in (57) in the case of a deep random Gaussian NN with architecture L = 3, , and , for different values of . In both tables, we consider four different inputs,
|
Table 2. Values of Cbound in (56) for Different Inputs for a Shallow Random Gaussian NN with Architecture L = 1, , and ReLU Activation Function
| n | 1 | 10 | 102 | 103 | 104 | 105 | |||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | |
| Cb | |||||||||||||||||||
| 1 | 0.02 | 0.21 | 1.49 | 0.01 | 0.07 | 0.47 | 0.00 | 0.02 | 0.15 | 0.00 | 0.01 | 0.05 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | |
| 10 | 0.01 | 0.07 | 0.47 | 0.00 | 0.02 | 0.15 | 0.00 | 0.01 | 0.05 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | |
| n | 1 | 10 | 102 | 103 | 104 | 105 | |||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | |
| Cb | |||||||||||||||||||
| 1 | 0.02 | 0.21 | 1.49 | 0.01 | 0.07 | 0.47 | 0.00 | 0.02 | 0.15 | 0.00 | 0.01 | 0.05 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | |
| 10 | 0.01 | 0.07 | 0.47 | 0.00 | 0.02 | 0.15 | 0.00 | 0.01 | 0.05 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | |
| n | 1 | 10 | 102 | 103 | 104 | 105 | |||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | |
| Cb | |||||||||||||||||||
| 1 | 0.02 | 0.22 | 1.54 | 0.01 | 0.07 | 0.49 | 0.00 | 0.02 | 0.15 | 0.00 | 0.01 | 0.05 | 0.00 | 0.00 | 0.02 | 0.00 | 0.00 | 0.00 | |
| 10 | 0.01 | 0.07 | 0.47 | 0.00 | 0.02 | 0.15 | 0.00 | 0.01 | 0.05 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | |
| n | 1 | 10 | 102 | 103 | 104 | 105 | |||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | |
| Cb | |||||||||||||||||||
| 1 | 0.03 | 0.48 | 0.44 | 0.01 | 0.15 | 0.14 | 0.00 | 0.05 | 0.04 | 0.00 | 0.02 | 0.01 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | |
| 10 | 0.01 | 0.09 | 0.36 | 0.00 | 0.03 | 0.11 | 0.00 | 0.01 | 0.04 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | |
|
Table 3. Values of Cbound in (57) for Different Inputs for a Deep Random Gaussian NN with Architecture L = 3, , and ReLU Activation Function
| n | 104 | 105 | 106 | 107 | 108 | 109 | ||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 |
| Cb | ||||||||||||||||||
| 1 | 0.03 | 0.58 | 88.59 | 0.01 | 0.18 | 28.01 | 0.00 | 0.06 | 8.86 | 0.00 | 0.02 | 2.80 | 0.00 | 0.01 | 0.89 | 0.00 | 0.00 | 0.28 |
| 10 | 0.07 | 1.38 | 331.57 | 0.02 | 0.44 | 104.85 | 0.01 | 0.14 | 33.16 | 0.00 | 0.04 | 10.49 | 0.00 | 0.01 | 3.32 | 0.00 | 0.00 | 1.05 |
| n | 104 | 105 | 106 | 107 | 108 | 109 | ||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 |
| Cb | ||||||||||||||||||
| 1 | 0.03 | 0.58 | 89.30 | 0.01 | 0.18 | 28.24 | 0.00 | 0.06 | 8.93 | 0.00 | 0.02 | 2.82 | 0.00 | 0.01 | 0.89 | 0.00 | 0.00 | 0.28 |
| 10 | 0.07 | 1.38 | 331.85 | 0.02 | 0.44 | 104.94 | 0.01 | 0.14 | 33.18 | 0.00 | 0.04 | 10.49 | 0.00 | 0.01 | 3.32 | 0.00 | 0.00 | 1.05 |
| n | 104 | 105 | 106 | 107 | 108 | 109 | ||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 |
| Cb | ||||||||||||||||||
| 1 | 0.03 | 0.58 | 106.14 | 0.01 | 0.18 | 33.56 | 0.00 | 0.06 | 10.61 | 0.00 | 0.02 | 3.36 | 0.00 | 0.01 | 1.06 | 0.00 | 0.00 | 0.34 |
| 10 | 0.07 | 1.38 | 338.62 | 0.02 | 0.44 | 107.08 | 0.01 | 0.14 | 33.86 | 0.00 | 0.04 | 10.71 | 0.00 | 0.01 | 3.39 | 0.00 | 0.00 | 1.07 |
| n | 104 | 105 | 106 | 107 | 108 | 109 | ||||||||||||
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 |
| Cb | ||||||||||||||||||
| 1 | 0.03 | 1.90 | 2,998.50 | 0.01 | 0.60 | 948.21 | 0.00 | 0.19 | 299.85 | 0.00 | 0.06 | 94.82 | 0.00 | 0.02 | 29.99 | 0.00 | 0.01 | 9.48 |
| 10 | 0.07 | 1.69 | 3,027.47 | 0.02 | 0.54 | 957.37 | 0.01 | 0.17 | 302.75 | 0.00 | 0.05 | 95.74 | 0.00 | 0.02 | 30.27 | 0.00 | 0.01 | 9.57 |
whose Euclidean norm is, respectively, 0, strictly less than 1, 1 and strictly larger than 1, and
It is clear from Table 2 that for a shallow random Gaussian NN, the Gaussian approximation of the output is very good for any of the choices of the parameters n, Cb, and CW and of the input . It is also clear from Table 3 that for a deep random Gaussian NN, such as the considered NN with three hidden layers, the Gaussian approximation of the output is very good for some choices of the parameters Cb, CW, and n and of the input (also when the number n of neurons in the hidden layers is not excessively large), but it is very poor for other choices of these quantities. In order to analyze the influence of the parameters Cb and CW and of the input on the value of Cbound in (57), in Table 4 we reported the value of the constant C1 that appears in (57), and it is explicitly given in (39). As it can be seen from this table, the value of Cbound strongly depends on the value of C1, which, in turn, is closely related to the choice of the parameters Cb and CW and does not depend much on the norm of the input vector .
|
Table 4. Values of C1 in (39) for Different Inputs for a Deep Random Gaussian NN with Architecture L = 3, , and ReLU Activation Function
| CW | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 | 0.01 | 0.1 | 1 |
| Cb | ||||||||||||
| 1 | 1.55 | 14.80 | 85.04 | 1.55 | 14.80 | 85.00 | 1.55 | 14.80 | 83.86 | 1.55 | 14.78 | 33.09 |
| 10 | 0.36 | 3.52 | 31.83 | 0.36 | 3.52 | 31.83 | 0.36 | 3.52 | 31.82 | 0.36 | 3.52 | 30.28 |
References
- (2024) Gaussian random field approximation via Stein’s method with applications to wide random neural networks. Appl. Comput. Harmon. Anal. 72:1–38.Google Scholar
- (2024) Quantitative Gaussian approximation of randomly initialized deep neural networks. Machine Learning 113:6373–6393.Google Scholar
- (2003) On the dependence of the Berry-Esseen bound on dimension. J. Statist. Planning Inference 113:385–402.Google Scholar
- (2024) Non-asymptotic approximations of Gaussian neural networks via second-order Poincaré inequalities. Proc. 6th Sympos. Adv. Approximate Bayesian Inference, 45–78.Google Scholar
- (2024) A quantitative functional central limit theorem for shallow neural networks. Modern Stochastics Theory Appl. 11:1–24.Google Scholar
- (2011) Normal Approximation by Stein’s Method (Springer, Heidelberg, Germany).Google Scholar
- (2021) Non-asymptotic approximations of neural networks by Gaussian processes. Conf. Learn. Theory, 1754–1775.Google Scholar
- (2023) Quantitative CLTs in deep networks networks. Preprint, submitted July 12, https://arxiv.org/pdf/2307.06092.Google Scholar
- (2023) Random neural networks in the infinite width limit as Gaussian processes. Ann. Appl. Probab. 33:4798–4819.Google Scholar
- (2022) Rate of convergence of polynomial networks to Gaussian processes. Conf. Learn. Theory, 701–722.Google Scholar
- (2018) Deep neural networks as Gaussian processes. Internat. Conf. Learn. Representations (ICLR, Appleton, WI).Google Scholar
- (2024) Large and moderate deviations for Gaussian neural networks. Preprint, submitted June 25, https://arxiv.org/pdf/2401.01611.Google Scholar
- (2018) Gaussian process behaviour in wide deep neural networks. Internat. Conf. Learn. Representations (ICLR, Appleton, WI), 4:77–86.Google Scholar
- (2004)
On the maximal perimeter of a convex set in with respect to a Gaussian measure. Morel J-M, Teissier B, eds. Geometric Aspects of Functional Analysis, Lecture Notes in Mathematics (Springer, Berlin), 169–187.Google Scholar - (1996) Priors for infinite networks. Bickel P, Diggle P, Fienberg S, Krickeberg K, Olkin I, Wermuth N, Zeger S, eds. Bayesian Learning for Neural Networks, Lecture Notes in Statistics, vol. 118 (Springer, New York), 29–53.Google Scholar
- (2012) Normal Approximations with Malliavin Calculus (Cambridge University Press, Cambridge, UK).Google Scholar
- (2022) Multivariate normal approximations on the Wiener space: New bounds in the convex distance. J. Theoret. Probab. 35:2020–2037.Google Scholar
- (2022) The Principles of Deep Learning Theory. An Effective Theory Approach to Understanding Neural Networks (Cambridge University Press, Cambridge, MA).Google Scholar
- (2019) Multivariate second order Poincaré inequalities for Poisson functionals. Electron. J. Probab. 24:1–42.Google Scholar
- (1972) A bound for the error in the normal approximation to the distribution of a sum of dependent random variables. Berkeley Symp. Math. Statist. Prob. 6.2:583–602.Google Scholar
- (2023) Quantitative multidimensional central limit theorems for means of the Dirichlet-Ferguson measure. ALEA - Latin Amer. J. Prob. Math. Statist. 20:825–860.Google Scholar
- (2009) Optimal Transport: Old and New (Springer, New York).Google Scholar
- (2019) Non-Gaussian processes and neural networks at finite widths. Preprint, submitted September 30, https://arxiv.org/abs/1910.00019.Google Scholar
- (2019) Wide feedforward or recurrent neural networks of any architecture are Gaussian processes. Wallach H, Larochelle H, Beygelzimer A, d’Alché-Buc F, Fox E, Garnett R, eds. Adv. Neural Inform. Processing Systems, vol. 32 (Curran Associates, Red Hook, NY).Google Scholar

