Large-Time Behavior of Finite-State Mean-Field Systems With Multiclasses
Abstract
We study in this paper large-time asymptotics of the empirical vector associated with a family of finite-state mean-field systems with multiclasses. The empirical vector is composed of local empirical measures characterizing the different classes within the system. As the number of particles in the system goes to infinity, the empirical vector process converges toward the solution to a McKean-Vlasov system. First, we investigate the large deviations principles of the invariant distribution from the limiting McKean-Vlasov system. Then, we examine the metastable phenomena arising at a large scale and large time. Finally, we estimate the rate of convergence of the empirical vector process to its invariant measure. Given the local homogeneity in the system, our results are established in a product space.
Funding: This research was supported by Discovery Grant of the Natural Sciences and Engineering Research Council of Canada [NSERC 315660] and by Carleton University.
1. Introduction
Interacting particle systems with multiclasses, widely encountered in a variety of domains going from statistical physics, chemistry, communication networks, and biology to finance, have recently attracted the interest of many researchers, and several models have been proposed to understand their large-scale behavior. See, for example, Graham (2008), Collet (2014), Collet et al. (2016), Chong and Klüppelberg (2019), Knöpfel et al. (2020), Meylahn (2020), and Nguyen et al. (2020) and the references therein for an overview of recent advances on the subject. In multiclass systems, the particles come from different subpopulations, within which they are homogeneous, and thus, the entire system is heterogeneous but composed of homogeneous subpopulations. Thence, one can average over the local symmetries within the different classes to describe mean-field interactions through local empirical measures. Thus, the entire system is characterized by an empirical vector composed of the local empirical measures.
The focus of the current article is on a particular family of mean-field multiclass models describing the evolution of block-structured networks with dynamically changing multicolor nodes. The state-space here is a finite set of colors, and the particles are identified as the nodes of the network. This class of models was proposed in Dawson et al. (2020) to describe the dynamic of various physical phenomena, and the large-scale asymptotics were established. In particular, a multiclass propagation of chaos was proven to hold together with a law of large numbers, implying the convergence of the empirical vector toward the solution to a McKean-Vlasov system of equations as the total number of particles N in the system goes to infinity. Moreover, these authors studied the large deviations principle for the empirical vector process over finite time intervals.
We propose in this article to study the large-time behavior of the family of models introduced in Dawson et al. (2020). Our motivation comes from the interesting characteristics of these systems as well as the importance of their large-time behavior for many applications. Notice that numerous works on the large-time behavior of various systems of interacting particles exist in the literature. See, for example, Cox and Greven (1990), Dawson and Greven (1993; 1999), Dawson et al. (1995), Greven and den Hollander (2007), Döring and Mytnik (2013), and Kuehn (2015) and references therein. Therefore, our current contribution aims to be the continuation of the aforementioned references. The approach taken in this document is summarized as follows.
First of all, let us mention that the propagation of chaos and the law of large numbers established in Dawson et al. (2020) are valid over finite time intervals. Thence, on any finite time horizon (not too large), one could use the McKean-Vlasov limit system as an approximation for the large N-particle system because it is classical in the mean-field literature. However, when time tends to infinity, these results are no longer necessarily true, and care should be taken in using this approximation. Indeed, as we detail throughout this article, the validity of the approximation is intimately linked to the critical points of the McKean-Vlasov limit system and its stability. Intuitively, if the McKean-Vlasov limit system has several ω-limit sets, one can wonder which of these sets characterizes the large-time behavior of the large N-particle system.
As a starting point to tackle this problem, we study in Section 4 the asymptotics of the invariant measure of the empirical vector. In particular, we establish, under mild conditions, the large deviations principles for the invariant measure in two different cases, first when there exists a unique asymptotically stable equilibrium to the McKean-Vlasov system (Theorem 4) and then when the McKean-Vlasov system has multiple ω-limit sets (Theorem 5). More precisely, Theorem 4 states that when the McKean-Vlasov system has a unique globally asymptotically stable equilibrium ξ0, the value at a given state ξ of the rate function that governs the large deviations principle of the invariant measure is given by the minimum cost of transporting the system state from the globally asymptotically stable equilibrium ξ0 to ξ across all the possible paths in any time duration. Theorem 5 gives a generalization to the case where the McKean-Vlasov system has multiple ω-limit sets. To this end, one relies on the classical hypothesis of Friedlin and Wentzell (2012). Note that the large deviations principles of the invariant distribution have been established for interacting diffusions in Dawson and Gärtner (1989) and for finite-state mean-field systems on complete graphs in Borkar and Sundaresan (2012). We adopt here the control theory approach introduced in Biswas and Borkar (2011) for small noise diffusions and extended in Borkar and Sundaresan (2012) to interacting jump processes on complete graphs. Therefore, we further extend this approach to the heterogeneous case with block-structured interaction graphs.
Next, we investigate the metastable phenomena, in the sense of Friedlin and Wentzell’s (2012) small noise stochastic systems, emerging when the McKean-Vlasov limiting system has several ω-sets. Namely, in the case of multiple attractors, the empirical vector process, which is associated with a large but finite number of particles N, would transit between these attractors when the time is large. One then aims to estimate the most probable order in which these transitions occur and the average time spent in the neighborhood of each of the ω-limit sets. Understanding these phenomena is of great practical interest, and hence, our motivation. We describe in Section 5 the metastable phenomena in more detail and give important estimates. For this, we adopt the classical approaches of Freidlin and Wentzell (2012, chapter 6) and Hwang and Sheu (1990). The main ingredients are the large deviation properties of the empirical vector process over finite time periods established in Section 3, and the main tool is the Freidlin-Wentzell quasipotential derived from the rate function characterizing the large deviations principle. Subsequently, we develop in Theorem 7 an estimate of the time required for the empirical vector process to converge to its invariant measure. We find out that when the time is of the order , for all and being an appropriate constant, the empirical vector process is very close to its invariant measure. Interestingly, this coincides with the timescale found in Hwang and Sheu (1990) for diffusion processes and in Borkar and Sundaresan (2012) for finite-state mean-field systems on complete graphs. We underline that the metastable phenomena in the Section 5 are studied under Assumption 2. Note, however, that if these assumptions are not imposed, things are more complicated, and this is not pursued here. The reader may consult Zhou et al. (2012), Bouchet et al. (2016), and Tang et al. (2017) and the references therein for discussions and examples.
It is worthwhile to mention that, although the methodology used in the current work follows the one in Borkar and Sundaresan (2012) and Yasodharan and Sundaresan (2019) based on the small-noise large deviations of diffusion processes introduced in Hwang and Sheu (1990) and Freidlin and Wentzell (2012), we emphasize that our results are established for empirical vectors and thus on product spaces. In particular, the multiclass structure of the system breaks down the homogeneity and global interaction hypothesis made in Borkar and Sundaresan (2012) and Yasodharan and Sundaresan (2019). Thus, one cannot directly rely on the results obtained in Borkar and Sundaresan (2012) and Yasodharan and Sundaresan (2019) because the global empirical measures in our current model are not expected to satisfy the law of large numbers, nor is the large deviations principle established for the global empirical measure. Our strategy to overcome the difficulties arising from the heterogeneity is through a finer-grain analysis of each local empirical measure describing each subclass.
The rest of this paper is organized as follows. In Section 2, we revisit the family of models introduced in Dawson et al. (2020). One may also consult Dawson et al. (2020; section 2) for a full detailed description together with some examples. Section 3 is dedicated to the large deviations properties of the empirical vector process over finite time intervals. We first recall the principal results of Dawson et al. (2020) and then introduce some additional results used in the following sections. Then, we prove in Section 4 our first set of the main results. Namely, Theorem 4 gives the large deviations principle for the empirical measures when the McKean-Vlasov system has a unique globally asymptotically stable equilibrium, and Theorem 5 gives the large deviations principle of the invariant measure when there are multiple ω-limit sets. In addition, we recall some important concepts from the Friedlin-Wentzell large deviations theory. We then investigate in Section 5 the metastability of the N-particles system. Based on a set of results, we provide estimates of the metastable transitions. Finally, Theorem 7 gives the time required for the empirical vector process to converge toward its invariant measure.
Because the results of Sections 4 and 5 rely on the Freidlin-Wentzell program described in Freidlin and Wentzell (2012, chapter 6), we give in Appendix A the generalization of some of these results to our current setting. Finally, to facilitate the reading, we leave the lengthy proofs of some technical results in Appendix B.
2. The Setting
Consider a graph composed of r blocks of sizes , respectively, where is the set of the nodes and Ξ is the set of the edges. Denote by the total number of the nodes in the network. Moreover, suppose that each block Cj is a clique, that is, all the Nj nodes of the same block are connected. We divide the nodes of each block Cj into two sets:
The central nodes : connected to all the other nodes of the same block but not to any node from the other blocks. We set .
The peripheral nodes : connected to all the other nodes of the same block and all the peripheral nodes of the other blocks. We set .
Also, we denote by Nc (resp. Np) the total number of central (resp. peripheral) nodes in the graph. Notice that an example of such a graph is shown in Figure 1. Let be a finite set of K colors. Suppose that each node of the graph is colored by one of the K colors at each time. For each and (resp. ), denote by the stochastic jump process that describes the evolution of the color of the node n through time. Let be the directed graph where describes the set of admissible jumps. In addition, whenever , a node colored by z is allowed to jump from z to at a rate that depends on the current state of the node and the state of its neighbors (adjacent nodes). To characterize these neighborhoods, we introduce, for each block , the following local empirical measures describing, respectively, the state of the central and the peripheral nodes of the j-th block at time t,
The central nodes dynamic: Each central node jumps from color z to , with , at rate
(2.2)which depends on its current state and the states of its neighbors through the empirical measures and .The peripheral nodes dynamic: Each peripheral node jumps from color z to , with , at rate
(2.3)which depends on its state and the states of its neighbors through the local empirical measures .
Notice that the rate functions and depend on the number of nodes and within each category, but we omit this dependency to not overload the notations, because they do not play a direct role in the study conducted in this paper. In particular, the dependency is not only through the empirical measures and but also through the proportions of nodes of the different categories. For example, one possible dependency is described in Dawson et al. (2020) where we suppose the existence of some measurable functions and such that
For any probability measures and any real numbers with ,
For any and any real numbers such that ,
In such a case, the rate functions write as
We make the following assumptions throughout the paper.

The directed graph is irreducible.
The rate functions and are Lipschitz and uniformly bounded away from zero, that is, there exists c > 0 such that for all and .
For each block , there exist such that, as ,
(2.4)
Given the compactness of the state-space , the space of probability measures over is also compact by Prokhorov’s theorem. Therefore, the rate functions and are continuous because Lipschitz, and uniformly bounded from above, that is, there exists a constant such that for all , we have and .
3. Large Deviations of the Empirical Measures
This section gathers large deviations results of the empirical vector process over finite time intervals. We start by introducing some additional notations that will be used in the sequel. Denote by the Skorokhod space of càdlàg functions defined on with values in and by the set of the probability measures over it. Denoting by the full description of the N particles over the finite time interval , we let be the vector of “historical” empirical measures defined by
Thus, . Denote by the law of XN with initial condition . Note that the distribution of MN depends on the initial condition only through its empirical vector defined by
Define by the distribution of MN, which is the pushforward of under the mapping GN.
Recalling (2.1), consider the -valued empirical vector process defined as
Notice that , and is the projection , at time t, of MN, that is,
Denote by the distribution of μN. The flow μN takes values in the product space . Also, denote by the n-th particle’s law with initial condition zn in the case of noninteraction, that is, when all of the particles are independent of each other and the color of each node changes with a constant rate equal to 1 for all allowed transitions , and all other transition rates are zero. Thus, the law of the entire noninteracting system is given by . Moreover, denote by the distribution of the corresponding historical empirical vector MN, where νN is the initial empirical vector. Therefore, the Radon-Nikodym derivative at any is given by (see Dawson et al. (2020), equation (4.10)),
Equip the Skorokhod space with the metric
For any , define the rate matrices
Denote by τ the log-Laplace transform of the centered Poisson distribution with parameter 1 given by , and let be its Legendre transform defined by
Define, for any ,
Finally, we introduce some spaces and topologies of interest related to the large deviations principle of the sequence of probability measures associated with the sequence of vectors of “historical” empirical measures . Although this result is not presented in the current paper (cf. Dawson et al. 2020, theorem 4.2), these spaces are used in the sequel and thus are introduced here to facilitate the reading. First, consider the Polish space , where
We endow the set with the weak topology , that is, the weakest topology under which as if and only if
3.1. Large Deviations Over Finite Time Intervals
We recall here the large deviations principle for the sequence of probability measures over finite time intervals.
Suppose that weakly. The sequence of probability measures obeys a large deviations principle in the space , with speed N, and rate function given by (3.12).
Moreover, if a path satisfies , then, for and are absolutely continuous, and there exist rate matrices and , with , such that
In such a case, there exist unique rate matrices and (up to almost everywhere equality with respect to ) such that
Conversely, if a path satisfies the following: μ is absolutely continuous, , and there exist time-varying rate matrices and such that μ satisfies (3.17), (3.18), and (3.19), then the good rate function evaluated at μ is given by the left-hand side of (3.18).
See Dawson et al. (2020, theorem 4.2) for the proof of the first statement. The converse statement follows by a simple generalization of Léonard (1995b, theorem 7.1). □
The action functional S characterizes the difficulty of the passage of μN near μ in the time interval . Indeed, according to Theorem 1, the probability of such a passage behaves like as .
Observe from (3.18) that if the rate function , then μ must be the solution to the McKean-Vlasov system (3.10) with initial condition .
The following result states the uniform large deviations principle for with respect to the initial condition ν over compact sets.
For any compact set , any closed set , and any open set , we have
See Dembo and Zeitouni (2010, corollary 5.6.15). □
3.2. Large Deviations at Initial and Terminal Times
Fix T > 0. Recall that is the distribution of the empirical process , with initial conditions given by the empirical vector . Let be the empirical measure at time T, and let be its distribution. The next result states the large deviations principle for the sequence .
Suppose that weakly. The sequence of probability measures obeys a large deviations principle in the space with speed N and good rate function
Moreover, is bounded for all , and its infimum is achieved.
Recall that the space is equipped with the product topology induced by the product metric , where, for all ,
Thus, it is easy to see that the application is continuous. Therefore, a simple application of the contraction principle (cf. Dembo and Zeitouni 2010, theorem 4.2.1) proves the validity of the large deviations principle. In addition, because is a good rate function, its infimum is achieved over closed sets. Finally, the boundedness follows from Lemma 1 below. □
We next state some useful technical results.
The following statements hold:
There exists a constant such that, for any , there is a piecewise linear and continuous path μ, with and , having constant velocity in each linear segment and satisfying .
For any , we have .
There exists a constant C3 such that, for each , there exists a such that implies that .
See B.1. □
Let, for and be rate matrices such that the solution to the system
See B.2. □
Let, for and be rate matrices such that the solution to the system
See B.3. □
The mapping is uniformly continuous.
This is a mild generalization of Borkar and Sundaresan (2012, lemma 3.3). Fix T > 0, , and let . From Lemma 1 there exists a constant C3 such that, for any with , we have . Let be two points in the space equipped with the metric defined by
Traverse the path from ν1 to ν2 in time with cost at most .
Consider the path given by with , where μ is the optimal -path μ from ν2 to ξ2. Therefore, travel from ν2 to ξ2 in the duration along the path .
Traverse the path from ξ2 to ξ1 in time with cost at most .
Thus the minimum cost for traversal from ν1 to ξ1 is at most the sum of the previous paths. Hence, by Lemmas 2 and 3, one obtains
Note that for u > 0. Thus, using again Lemma 1 leads to
Now, obbserve that
Denote by the law of the initial empirical measure vector , and let denote the joint law of . We next establish the large deviations principle for the sequence .
Suppose that the sequence satisfies the large deviations principle with speed N and good rate function . Then, the sequence of joint laws satisfies the large deviation principle with speed N and good rate function
Here, is the joint distribution of is the distribution of , and is the conditional distribution of given . Thus one can write
Because both the sequences and obey the large deviations principle, one can apply Feng and Kurtz (2006, proposition 3.25) in order to derive the rate function corresponding to the large deviations principle of the sequence in the product space. To this end, we first verify that the conditions of application of Feng and Kurtz (2006, proposition 3.25) are satisfied. Let weakly. By Theorem 2, the sequence of laws of the terminal measure satisfies the large deviations principle with speed N and good rate function . Moreover, recall that . From Dawson et al. (2020, lemma 4.6) we have that for any ,
Moreover, using Varadhan’s lemma (see Léonard 1995a, proposition 2.5) we obtain, for every ,
Furthermore, by defining
One can observe that, for any , the mapping is continuous. Indeed, for any , define the mapping
Because f is continuous, and is also continuous by Lemma 4, the mapping is jointly continuous. Let weakly, and for each N, let ξN denote a point where the supremum in (3.27), corresponding to νN, is attained. This is consistent because the space is compact. Thus for each N. Moreover, by the same compactness argument, the sequence has a convergent subsequence that converges to for some . Let denote this subsequence. Therefore, as . By the continuity of η we have, as ,
From the definition of ξN, one can observe that, for any . Hence, by the continuity of η
Thus , from which we deduce that is continuous because is jointly continuous. Thus, the mapping is indeed continuous.
Now, notice that when there are N particles in the system, the initial empirical measure takes values in the product space , where
We are now in a position to apply Feng and Kurtz (2006, proposition 3.25). Indeed, we have shown that the sequence is exponentially tight. Moreover, the convergence in convergence in (3.26) is uniform for ν in compact subsets of . Furthermore, the function is continuous in ν. Hence, because satisfies the large deviations principle with good rate function , then satisfies the large deviations principle with good rate function given by (3.24). This concludes the proof. □
4. Large Deviations of the Invariant Measure
By Assumption 1, the graph of the allowed transitions is irreducible. Moreover, the state space is finite; therefore, for each fixed total number of particles N, there exists a unique invariant measure for the Markov process . Hence, there is a unique invariant measure, denoted by , for the -valued Markov process μN. The goal of this section is to investigate the large deviations properties of the sequence under two separate scenarios. First, we consider the case where the limiting McKean-Vlasov system (3.10) has a unique globally asymptotically stable equilibrium. Then, we treat the general case with multiple ω-limit sets.
4.1. Unique Globally Asymptotically Stable Equilibrium
We first establish the large deviations principle for the invariant measure in the case where the limiting McKean-Vlasov system has a unique globally asymptotically stable equilibrium ξ0.
Suppose that Assumption 1 holds true. Moreover, suppose that the McKean-Vlasov system (3.10) has a unique globally asymptotically stable equilibrium ξ0. Then, the sequence satisfies a large deviations principle with speed N, and a good rate function s given by
The rest of this section is dedicated to the proof of Theorem 4. The proof is based on the control-theoretic approach introduced in Biswas and Borkar (2011) for small-noise diffusions and in Borkar and Sundaresan (2012) for finite-state mean-field systems on complete graphs. We proceed through several lemmas. We start by establishing a subsequential large deviations principle.
Let be a sequence of natural numbers going to . Then, there exists a subsequence such that satisfies a large deviations principle with speed Nk and a good rate function s verifying
Moreover, , and there exists some such that .
The space is compact since it is finite, and then is also compact. Hence, the product space is also compact. Moreover, the space endowed with the product metric is a metric space, and thus it has a countable basis. Therefore, there exists a sequence such that satisfies the weak large deviations principle in (Dembo and Zeitouni 2010, lemma 4.1.23). Moreover, because is compact, satisfies the strong large deviations principle with a good rate function and speed Nk (Dembo and Zeitouni 2010, lemma 1.2.18).
Fix an arbitrary T > 0. Recall that is the probability measure of the initial empirical measure vector . Set . Then, by Theorem 3, the sequence of joint laws satisfies the large deviations principle along the subsequence with speed Nk and good rate function . Using the continuity of the projection and the contraction principle (Dembo and Zeitouni 2010, theorem 4.2.1), the sequence of the terminal probability distributions satisfies a large deviations principle with the good rate function
Because is the invariant measure and , one has . Thus, by the uniqueness of the rate function, we deduce that
Finally, because every rate function is nonnegative, we have . In addition, with s being a good rate function, it attains its minimum, and thus, there exists some such that . This concludes the proof. □
Notice that Equation (4.3) has multiple solutions. The trivial is one of them. To see this, take the McKean-Vlasov path μ of duration T given by (3.10) starting at some initial condition ν and ending at ξ; then, , and the right-hand side of (4.3) vanishes when . Therefore, one shall identify more conditions satisfied by the rate function s.
Observe from (3.22) that (4.3) is equivalent to
Hence, one can see (4.3) as a Bellman equation associated with an optimal control problem, with s being the corresponding value function of a minimization problem over path space, with paths defined on , for , and ending at ξ. Therefore, one must determine the optimal control problem for which Equation (4.3) is the dynamic programming equation and ξ is the terminal condition. Because the terminal condition is fixed, we shall define translated and reversed-time paths that start from ξ at time 0. In particular, for and a path satisfying
For each , there exists a path and families of rate matrices and defined on such that, for ,
Let be an infinite duration path, and let ψm be its restriction to the finite interval given by
Equip the space of infinite duration paths with the following metric:
One can easily observe that the restriction ψm is continuous for all m. Moreover, define the reversed restriction by
For any , consider the sets
Thus is the set of infinite duration paths , with the corresponding restrictions on time intervals having the costs bounded by B, for any . On the other hand, the set comports the paths of duration , with corresponding reversed path having cost bounded by B. We next prove that the set is compact for any .
First, observe that is a subset of a metric space; thus it is enough to prove that it is sequentially compact. Take an infinite sequence and the corresponding restriction sequence . Note that the sets , for , are compact. Indeed, with being a good rate function, the corresponding level set is compact, and then, because for each , we have that is compact because it is a continuous image of a compact. Therefore, one can find an infinite subset such that converges. Moreover, one can take a further subsequence represented by the infinite subset such that converges. We can continue this procedure for all and eventually take the subsequence along the diagonal. This subsequence converges for every interval . Take , its pointwise limit for every t. The restriction and the corresponding satisfy , and thus, . This guarantees that is sequentially compact and thus compact.
Fix . From Lemma 5, there exists such that , and then, by Lemma 1,
Take in the sequel . Starting from any location , the minimum cost of transporting the system from ν to ξ in mT units of time is also bounded by . Indeed, consider the path consisting of traversing the McKean-Vlasov path with initial condition ν in units of time; the corresponding cost is zero. Then, go to ξ in T units of time. The cost of the last traversal is bounded by , as stated by Lemma 1. Now, if we consider the translated and reversed time paths that start from ξ at time 0, they have a cost at most B and stay within for the duration . Let
Next, we prove that is compact. To this end, and because is a subset of a compact set, it suffices to show that it is closed. Let be a point of the closure of , and then one can find a sequence such that . By definition of , we have . In addition, set for some . Consider the corresponding translated and reversed paths and μ. Therefore, using the lower semicontinuity property of the good rate function , we obtain
We are now in the position to conclude the proof. Given that is nonempty and closed, by the continuity of the restriction ψm, the image set is nonempty and closed. Moreover, because is compact, the intersection is also compact. Furthermore, note that . Indeed, take an optimal path that starts from ξ and passes by ν at time , and by at time mT then, its restriction to the interval is necessarily optimal (if not, would not be optimal on ). Therefore, , and thus is a nested decreasing sequence of subsets. Hence, is in turn a nested decreasing sequence of nonempty, compact, and closed subsets, and thus, by Cantor’s intersection theorem, its intersection is not empty. Take and let , for and , be its reversal restriction on the interval , and let be the reversal restriction of on . Thus for all . Thence, from Theorem 1, there exist families of rate matrices and such that satisfies (3.17) on with initial condition . Define for all and for . Thus, the rate matrices and are defined on such that satisfies (4.10) on with initial condition . The equality in (4.11) follows because the path satisfies (4.8) on any interval by the principle of optimality. This concludes the proof of the lemma. □
Next, we show that the optimal path given in Lemma 6 must end up in an invariant set of the dynamics given by the time-reversed McKean-Vlasov system.
Let
The path given in Lemma 6 remains in and satisfies the dynamics (4.10) with initial condition for some . Moreover, notice that the integral term in (4.11) is increasing with m because the integrand is nonnegative. Hence, because the equality is valid for all , the term must decrease as m increases. But we know from Lemma 1 and (4.12) that , and then there exists such that as .
Take a subsequence of and let be its limit. Moreover, denote by ν the limit of the subsequence , and consider the path of duration T given by
Thus the following convergence,
By the nonegativity of the rate function ST, one further finds
In addition, Lemma 4 tells us that the rate function ST is uniformly continuous in both its arguments. Thence, we deduce that . This means that the path, say μ, that goes from ν to in T units of time with no cost , is necessarily the McKean-Vlasov path that satisfies (3.10) on , with initial condition and terminal condition (see Remark 3). Hence, the reversed-time path satisfies (4.13) for with initial condition . This proves that Ω, the ω-limit set of , is contained in an ω-limit set of the reversed-time McKean-Vlasov dynamics (4.13). The lemma is proven. □
The following result proves that the rate function s vanishes at the equilibrium point ξ0.
If the McKean-Vlasov system (3.10) has a unique globally asymptotically stable equilibrium ξ0, then .
Consider the Mckean-Vlasov dynamics (3.10) with initial condition satisfying (see Lemma 5). Because ξ0 is the unique asymptotically stable equilibrium, . Moreover, the McKean-Vlasov path μ has no cost, that is, for each T > 0. Therefore, from Equation (4.3), one obtains
Taking the limit as , and using the lower semicontinuity of s, one gets
Finally, the next result gives the unique form of the rate function.
If the McKean-Vlasov system (3.10) has a unique globally asymptotically stable equilibrium ξ0, then the solution to (4.3) and (4.11) is unique and is given by (4.1).
From Lemma 7, the ω-limit set Ω of the path given in Lemma 6 is contained in an ω-limit set of the reversed-time McKean-Vlasov dynamics (4.13), which is both positively and negatively invariant. Moreover, the path stays within (see the proof of Lemma 6), and thus, . Because is compact and Ω is closed being an ω-limit set, then it is nonempty, compact, and connected. But by the assumption that the forward McKean-Vlasov system (3.10) possesses a unique globally asymptotically stable equilibrium ξ0, ξ0 is thus the unique nonempty, compact, connected set in that is both positively and negatively invariant for the dynamics in (4.13). Thus we necessarily have . Therefore, letting in (4.11), we obtain, by lower semicontinuity of the rate function s,
But from Lemma 8, we have , and then
Finally, applying (4.3) shows that is upper bounded by the right-hand side of (4.1), and thus the equality holds. The uniqueness follows because the rate function is unique (Dembo and Zeitouni 2010, section 4.1.1). □
We are now ready to conclude the proof of Theorem 4. Take an arbitrary sequence of positive numbers going to . From Lemma 5, there exists a subsequence such that satisfies the large deviations principle with speed Nk and a good rate function s that satisfies equation (4.3). Moreover, from Lemma 9, the rate function s is uniquely determined by Equation (4.1). Therefore, for every sequence, there exists a further subsequence such that satisfies the large deviations principle with speed Nk and the same good rate function s given by (4.1). Hence, because the rate function s corresponding to the subsequence satisfying a large deviations principle does not depend on this subsequence, one can conclude that the sequence of probability measures satisfies the large deviations principle with speed N and good rate function s defined by (4.1). □
4.2. Multiple ω-Limit Sets
We consider now the general case where the McKean-Vlasov system (3.10) admits multiple ω-limit sets. To this end, we rely on the classical Freidlin-Wentzell program (cf. Freidlin and Wentzell 2012). The same strategy was adopted in Borkar and Sundaresan (2012) for finite-state mean-field systems on complete graphs. We begin by introducing the central concepts.
First, recall the important notion of quasipotential defined, for any , by
Roughly speaking, measures the difficulty for the empirical vector process to move from ν to ξ in a finite time interval.
The mapping is uniformly continuous.
Replacing the metric by the product metric , the proof of (Borkar and Sundaresan 2012, lemma 3.4) holds verbatim. □
Using the Freidlin-Wentzell quasipotential, we define the following equivalence relation on :
We make throughout the section the following Freidlin-Wentzell assumptions (Freidlin and Wentzell 2012, chapter 6, section 2).
In , there exist a finite number of compact sets such that:
For any two points ν1 and ν2 belonging to the same compact, we have .
For each and imply .
Every ω-limit set of the McKean-Vlasov system (3.10) lies completely in one of the compact sets Ki.
For any compacts Ki and , let us introduce
Let L be a finite set, and let be a subset of L. An oriented graph consisting of edges (m, n) is called a W-graph if it satisfies the following conditions:
Every point is the initial point of exactly one edge;
There are no closed cycles in the graph.
The second condition above can be replaced by the following condition. For any point , there exists a sequence of edges leading from it to some point . Denote by G(W) the set of W-graphs. Take the indices corresponding to the compact sets given in Assumption 2, and we define the following quantity,
Suppose that both Assumptions 1 and 2 hold true. Then the sequence satisfies the large deviations principle with speed N and a good rate function s given by
Again, the proof is split into several lemmas.
Assume that Assumption 2 holds true. The rate function s satisfies the following assertions:
There exists ξ0 in some compact , with , that satisfies ;
The rate function s is constant on each of the compacts , that is, for all , for some nonnegative real numbers .
Let μ be the McKean-Vlasov path given by (3.10) and starting at some with . Note that the existence of such a point is guaranteed by Lemma 5. Take such that . By Assumption 2, for some . Using the same arguments as in the proof of Lemma 8, the first statement follows.
Fix and let . Thus . Hence, by (4.16) and (4.17), for any , there exists T > 0 and a path of length T starting at ν and ending at ξ such that . Consequently, using (4.3), one obtains
Reversing the roles of ν and ξ, one gets
Combining the two last inequalities leads to . Letting , we obtain , which proves the second statement. □
Assume that Assumption 2 holds true. Then, the rate function s, the solution to equations (4.3) and (4.11), is uniquely given by (4.21).
Using the same arguments as in the proof of Lemma 9 together with Assumption 2, one can conclude that the ω-limit set Ω of the path given in Lemma 6 is contained in some compact , with . Therefore, letting in Equation (4.11), for some , which gives that is lower bounded by the right-hand side of (4.21). The upper bound is obtained by Equation (4.3), and thus the proof follows. □
Using Lemma 11, Theorem 8, and the definition of the large deviations principle, it is easy to see that the values of the rate function s at the compact sets Ki are given by , with defined by (4.20).
Take an arbitrary sequence of positive numbers going to . From Lemma 5, there exists a subsequence such that satisfies the large deviations principle with speed Nk and a good rate function s that satisfies Equation (4.3). Moreover, from Lemma 12, the rate function s is uniquely determined by Equation (4.21), where for . Therefore, for every sequence, there exists a further subsequence such that satisfies the large deviations principle with speed Nk, and the same good rate function s given by (4.21). Hence, because the rate function s of the subsequence satisfying a large deviations principle does not depend on this subsequence, we conclude that the sequence of probability measures satisfies the large deviations principle with speed N, and good rate function s defined by (4.21). □
5. Metastability and Convergence to the Invariant Measure
We study in this section the metastable phenomena that occur when the total number N of particles in the system as well as the time t are large. First, let us briefly summarize the main results of the previous sections and their consequences.
From the law of large numbers (cf. Dawson et al. 2020, corollary 3.1), as and for converging initial conditions , the sequence converges weakly and uniformly over any finite time interval toward the deterministic solution μ of the McKean-Vlasov system in (3.10) with initial condition ν. This suggests that, when and over any finite time interval , one can approximate the trajectories of the empirical vector process μN by the solution to the McKean-Vlasov system in (3.10) with initial condition ν.
Moreover, given that the graph of allowed transitions is irreducible, there is a unique invariant measure for μN. Therefore, for each N, the distribution of μN converges toward as . Thence, the large N behavior is described by the large deviations properties of the invariant measure established in Theorems 4 and 5. Thence, the natural question one might ask is whether we can interchange the and limits, namely
This classical question is related to the limiting behavior of the McKean-Vlasov system (3.10). In particular, two distinct cases must be considered. The most simple situation is when the McKean-Vlasov system (3.10) has a unique globally asymptotically stable equilibrium . In this case, one can prove that the unique invariant measure of the empirical process vector μN converges toward the point mass as . Moreover, because is unique, it is independent of the initial condition. This leads to a justification of interchange of the limits and . A detailed discussion of this scenario is given in Benaïm and Le Boudec (2008).
The second and more complicated case, which we are interested in here, is when the McKean-Vlasov system (3.10) has multiple ω-limit sets, depending on the initial condition. In this case, starting at a given and letting , the solution to the McKean-Vlasov system in (3.10) goes to an ω-limit set corresponding to this initial condition. Moreover, recalling that the Birkhoff center of the solution μ to the McKean-Vlasov system (3.10) is the closure of the set of recurrent points, that is, the set of points such that , it is well known that the support of any limit point of , as , is a compact subset of the Birkhoff center of μ. See, for example, Benaïm and Le Boudec (2008, Theroem 3). Therefore, most of the time, for large but finite N, the empirical vector process remains close to the Birkhoff center of μ. However, the difficulty here is that the Birkhoff center contains multiple ω-limit sets, stable equilibrium, and/or limit cycles, depending on the initial condition. Thence, metastable phenomena are likely to arise. Here is an example.
Let the initial condition converge weakly to a given as . Then, on any finite time horizon, for a large but finite N, the empirical vector process μN would track the solution of the McKean-Vlasov Equation (3.10) starting at ν with a high probability. Therefore, as t becomes large, μN would enter a neighborhood of the ω-limit set of (3.10) corresponding to the initial condition ν. However, because N is finite, the process can exit the basin of attraction of this ω-limit set and likely remains in a neighborhood of another ω-limit set for a large amount of time before transiting again to the next one, and so on. This is an example of metastability. The goal of this section is thus to study such phenomena. The main ingredient is the large deviations properties developed in Sections 3 and 4.
The metastable phenomena were first studied for diffusion processes with a small-noise parameter. The two main references are Freidlin and Wentzell (2012, chapter 6), in which the authors studied these phenomena under the hypothesis summarized in Assumption 2, and Hwang and Sheu (1990), where slightly more general assumptions have been made. More recently, an extension to finite-state mean-field models on complete graphs was established in Yasodharan and Sundaresan (2019). We propose in this section an extension of the aforementioned results to mean-field models with jumps on block-structured graphs detailed in Section 2. The main idea is to consider the empirical vector process as a small-noise perturbation of the deterministic solution μ to the McKean-Vlasov system (3.10). Here, plays the role of the small-noise parameter ε considered in Hwang and Sheu (1990) and Freidlin and Wentzell (2012) in the sense that, as , we recover the “nonperturbed” McKean-Vlasov equation (3.10). Thence, under Assumption 2, one considers an embedded Markov chain Zn, for which the state space is the union of small neighborhoods of the compact sets and whose transitions probabilities allow us to estimate the exit and entering times of the empirical vector process μN in the neighborhood of the compact sets Ki, thus describing the metastability of the finite N-particles system.
5.1. Metastable Phenomena Estimates
Let us introduce some additional notations. For a set , let denote the open δ-neighborhood of A and denote its closure. Moreover, let the stopping time denote the first exit time from A. Recall Assumption 2 and let r0 and r1 be two positive numbers such that and , where we recall that is the product metric that equips the product space . Denote by the set from which we delete the r0-neighborhoods of , and let be the closure of . Furthermore, denote by the r1-neighborhood of Ki, and . Consider the following stopping times:
We first give an estimate of the stopping time τ1 of the first reentrance into the r1-neighborhood of one of the compact Ki. Note that similar results were established in Hwang and Sheu (1990, lemma 1.3) for small-noise diffusion processes, and in Yasodharan and Sundaresan (2019, lemma 3.6) for complete interaction mean-field systems with jumps.
Given and r0 small enough, there exists such that for any we have
First, notice that . By Lemma 20, there exists and such that for all and , we have
Moreover, , where . Because the compact set F does not contain any ω-limit set, Corollary 3 allows us to deduce that there exists a constant such that , and thus . Taking N large enough gives (13). □
Let be the indices corresponding to the compact sets given in Assumption 2. For any , introduce the following stopping times:
For a W-graph g, set . For and , let denote the set of W-graphs in which there is a sequence of arrows leading from i to j. For any subset , define
The next lemma gives upper and lower bound estimates for the mean entrance time into small neighborhoods of a set of compacts indexed by , starting from the neighborhood of a given compact Ki, with . Similar estimates have been obtained in Hwang and Sheu (1990, part I, lemma 1.6) and Yasodharan and Sundaresan (2019, lemma 3.10) for small-noise diffusion processes and finite-state mean-field systems on complete graphs, respectively.
Let , and let . Given , there exists and such that for any and , we have
By the strong Markov property, one obtains
Notice that the sum term in the last inequality corresponds to the expectation of the number of steps from ν until the first entrance in . Thus using the upper bound estimate given in Freidlin and Wentzell (2012, chapter 6, lemma 3.4) we find
Furthermore, Lemma 21 allows us to deduce that, for all sufficiently small r0 and sufficiently large N,
Define
The following result gives the estimate of the probability that the first entry of μN into a neighborhood of a set takes place via a given compact set Kj, with , starting from a neighborhood of Ki, with .
Let and . For any , there exists such that for any , sufficiently large N, and all we have
This follows by applying (Freidlin and Wentzell (2012, chapter 6, lemma 3.3) and making use of the estimate in (A.3). □
We now introduce the important notion of cycle, which, roughly speaking, describes how the process runs through its lifetime. Indeed, as explained above, for large but finite N, and over large time intervals, there are passages of the process μN between the neighborhoods of the compact sets Ki. The cycles then describe the most probable order in which the trajectories of μN traverse these neighborhoods and the time required to go from one compact to another. Notice nevertheless that the notion of cycle used here was introduced in Hwang and Sheu (1990), which is slightly different from the classical Freidlin and Wentzell notion of cycle (Freidlin and Wentzell (2012)).
Recall the definition of in (4.19), and set . From the estimate in (A.3) of the one-step transition probabilities of the Markov chain Zn, one can notice that, starting from a neighborhood of a given compact Ki, the most likely set that will be visited by the process μN, for large enough N, is the one that reaches the minimum . Denote by if . This gives us an oriented graph structure where the nodes are given by the set , and the edges are given by the relation , for . We refer to this graph by L. Moreover, for , we say that if there exists a sequence of arrows leading from i to j, that is, if there exists in L such that .
A cycle π in L is a subgraph of L satisfying the following:
and imply ,
For any in π, and .
For proof of the existence of cycles in L, one can consult Hwang and Sheu (1990, lemma A1). Next, we describe the decomposition of the set L into a hierarchy of cycles. Setting , the cycles of rank 1, or the 1-cycles, are defined as
We use the superscript 1 to refer to the 1-cycles, namely . For two 1-cycles, with , let and
We say that if , and if there is a sequence of arrows leading from to . This gives a cycle of second order, or a 2-cycle, and we use the notation to denote them. In this way, represents the exit rate from a cycle , as specified by Lemma 16 below.
By recurrence, assuming that we have defined up to m-cycles, we set
For , let and
We say that if , which gives us the -cycle. We keep going until we reach a given m for which the set of m-cycle is a singleton. From now on, we will use, with a slight abuse of notation, the notation πk to refer to both the k-cycle πk and the set of elements of L constituting it. Recall that for . Before stating some important results about cycles, we first introduce an example to help the reader have a clear picture of this notion.
Let the set of indices be , and consider the matrix corresponding to the values of , for
Set . Using the definition above, we find three 1-cycles, , characterized by the edges , and , characterized by the edges , and , and finally, , characterized by the edges and . Thus .
Now, in order to find the 2-cycles, we use again the construction above. Straightforward computations give . Thus, we deduce that . Moreover, we obtain , and thus, . Finally, we find , from which we deduce that . Hence, we have , which represents the collection of 2-cycles. Denote by the two elements of L2. Now, in order to identify the set of 3-cycles, we find again by simple calculations of the following quantities , and . Moreover, and . Therefore, and , which gives us the unique 3-cycle , the only element of L3. Now, we stop, because L3 is a singleton. See Figure 2 for an illustration.
The next result gives an estimate of the mean exit time from a cycle.

Let πk be a k-cycle, and let . Moreover, denote by the set of compacts not contained in πk. Then, given , there exist and such that for any and , we have
We have, from Hwang and Sheu (1990, lemma A3) and Hwang and Sheu (1990, corollary A4), that . Using this together with Lemma 14 leads to the result. □
The next result gives an estimate of the probability of going from one k-cycle to another without passing by any other element of Lk.
Let be two distinct k-cycles, and let . Moreover, denote . Then, given , there exist and such that, for all , and , we have
From Lemma 15 we have, for each and large enough N,
Therefore, summing over the disjoint compacts Ki we obtain
From the sums above, we select the term decreasing more slowly than the remaining ones, which gives us
Finally, from Hwang and Sheu (1990, lemma A5), we have that , which leads to the stated result. □
5.2. Convergence to the Invariant Measure
As aforementioned, there exists, under Assumption 1, a unique invariant measure for the empirical vector process μN. Thus, the distribution of μN converges, as time , toward . However, one might investigate the corresponding rate of convergence. Therefore, using the results from the previous section, we show that, when the time is of order , with and Λ a suitable constant detailed above, the empirical measure is very close to its invariant measure . Interestingly, despite the heterogeneity introduced by the block structure, a similar constant appears in the case of small-noise diffusion processes (Hwang and Sheu 1990, lemma 1.3) and in the case of homogeneous mean-field systems with jumps in Yasodharan and Sundaresan (2019, lemma 3.6). Before stating our main result, let us introduce further notations and intermediate results.
Let be such that . Moreover, define
Let denote the transition probability kernel associated with the empirical process μN. The next result gives a lower bound for the transition probability of reaching a small neighborhood of when T is of order for some .
Given , there exist , r > 0, and such that, for all , , and , we have
Replacing the space by the product space together with the corresponding metrics, the proof given in Yasodharan and Sundaresan (2019, theorem 3.21) adapted from Hwang and Sheu (1990, theorem 2.3, part I) holds verbatim. □
Under the conditions of Theorem 6, for all , and N sufficiently large, we have
The proof follows verbatim the proof given in Yasodharan and Sundaresan (2019, corollary 3.22). □
We state now the main result of this section, which gives the time scale at which the empirical vector converges toward its invariant measure.
There exists a constant such that, for any , there exist and such that, for all and ,
Let , and let T0, δ0, r, r1, and be as in the statement of Theorem 6. By Corollary 2 we have that, for any and ,
Moreover, by the Markov property and using Corollaries 2 and 1, we have that, for any , and some fixed t,
Define such that for , and
Thence we find,
Hence, by repeating the previous steps k times using the Chapman-Kolmogorov property given that μN is Markov, we find
Choose . Thus, . Moreover, using the property as , one obtains for large k
Choosing ε small enough such that , one finally obtains
6. Conclusion and Discussion
We have addressed in this paper the large-time behavior of the empirical measure vector associated with a family of finite-state mean-field models with multiclasses. In particular, we have established the large deviations principle for the invariant measure when the McKean-Vlasov limiting system has a unique asymptotically stable equilibrium and when there are multiple ω-limits sets. Also, we established various metastability phenomena estimates together with the speed of convergence of the empirical measure vector to its invariant measure. These results are of interest for various real-world applications ranging from engineering systems to the spread of infectious diseases. In particular, the multispecies setting represented by the block graph structure can be used as a model for various phenomena.
Many interesting questions remain nevertheless open and are worthwhile to explore. For instance, for the family of block graphs analyzed in the current paper, the degree of each node is O(N). An interesting extension is to consider other scaling regimes for the degree of the nodes. For example, the number of neighbors of the peripheral nodes could grow as . Even though it is sublinear in N, there could be a strong interaction. It would be interesting to study the large-time behavior of the vector empirical measure process in such cases.
The asymptotic analysis conducted in the current paper is on the vector of empirical measures. However, one might be interested to study the large deviation of the global empirical measure. Given the heterogeneity of the particles composing the global empirical measure, the latter is not a Markov process. Yet from the large deviations of the vector of empirical measures, one can use the contraction principle to deduce the large deviations principle for the sequence of global empirical measures. The question then is to understand the rate function. Another interesting question is to find criteria for the uniqueness of stationary distribution of the global empirical measure.
A natural extension of the current work is to consider the case of countable state spaces. Notice that if one replaces the finite space by a countable state space , the corresponding set of probability measures over E is no longer compact or finite-dimensional. Therefore, one has to impose stronger conditions on the rate functions to study the asymptotics of the system. For instance, the large deviations principle of the empirical measures has been established in Feng (1994) and Léonard (1995a) under different assumptions on the Levy kernels governing the jumps. However, the large deviations principle for the family of invariant measures is not available in the general countable state-space context and thus remains an open problem. In a recent contribution, Yasodharan and Sundaresan (2021), sufficient conditions have been established for a class of interacting particle systems on countable state space under which the sequence of invariant measures satisfies the large derivation principle with a rate function governed by the Freidlin-Wentzell quasipotential. Nonetheless, the proofs in Yasodharan and Sundaresan (2021) hold only under specific conditions on the transition graph and the forward transition rates. Also, the results in Yasodharan and Sundaresan (2021) are restricted to the case where the limiting McKean-Vlasov system has a unique globally asymptotically stable equilibrium and do not hold when there are multiple ω-limit sets. Thus, the latter case remains open to exploration. In particular, the Freidlin-Wentzell approach for studying large deviations of the invariant measures used in the current paper is not directly applicable in the case of countable state space because it requires the uniform large deviations principle over open subsets of , which is not available. Consequently, one cannot use this approach to study the related exit problems and metastability phenomena. It is then a challenging open question. Of course, these challenges are heightened in the multiclasses context considered in the current paper due to the technicalities arising from the product space setting.
The authors thank the anonymous referees for having read the paper with great care and having made several very useful suggestions that improved the exposition.
Appendix A: Freidlin-Wentzell program
We give here generalizations to our setting of a series of lemmas introduced in Freidlin and Wentzell (2012) in the case of diffusion processes and generalizations in Borkar and Sundaresan (2012) in the case of jump processes in one homogeneous population. These results play an important role in the study of the large-time behavior of the system. Because this is a slight generalization, we will give full proofs only when needed and refer to the previous references otherwise.
(Freidlin and Wentzell 2012, chapter 6, lemma 1.2). For any and any compact set , there exists a T0 such that for any there exists a function , with , satisfying with .
The proof of Borkar and Sundaresan (2012, lemma A.1) holds verbatim by replacing the metric by the product metric . □
(Freidlin and Wentzell 2012, chapter 6, lemma 1.6). Let all points of a compact set be equivalent to each other, but not to any other point in . Then, for any , ν, , there exist a T > 0 and a function defined on , with for all , and .
Using Lemma 1, the proof follows verbatim the proof of Borkar and Sundaresan (2012, lemma A.2). □
Let , and define the stopping time
(Freidlin and Wentzell 2012, chapter 6, lemma 1.7). Let all points of a compact set be equivalent to each other, and let . For any , there exists a such that, for all sufficiently large N and all , we have
Using Corollary 1, the proof of Borkar and Sundaresan (2012, lemma A.3) holds verbatim by replacing the metric by the product metric . □
(Freidlin and Wentzell 2012, chapter 6, lemma 1.8). Let K be an arbitrary compact subset of , and let G be a neighborhood of K. For any , there exists a such that, for all sufficiently large N and all ν belonging to , we have
The proof of Borkar and Sundaresan (2012, lemma A.4) holds verbatim by replacing the metric by the product metric . □
(Freidlin and Wentzell 2012, chapter 6, lemma 1.9). Let K be a compact subset of not containing any ω-limit set entirely. Then, there exist positive constants c and T0 such that, for all sufficiently large N, any , and any , we have
The proof of Borkar and Sundaresan (2012, lemma A.5) holds verbatim by replacing the metric by the product metric . □
Let K be a compact set not containing any ω-limit set entirely. There exists a positive integer N0 and a positive constant c such that for and any , we have
See Freidlin and Wentzell (2012, chapter 6, p. 149). □
Recall that r0 and r1 are positive numbers such that and . We denote by and let be the closure of for . Moreover, we denote by the r1-neighborhood of Ki and . Moreover, recall the following stopping times and , and consider the embedded Markov chain of states at hitting times of neighborhood of the stable limit sets . Let
(Freidlin and Wentzell 2012, chapter 6, lemma 2.1). For any , there is a small enough such that for any r2 satisfying , there is an r1 satisfying such that for all sufficiently large N, for all , the one-step transition probabilities of Zn satisfy
Replacing the space by the product space endowed with the product norm , the proof of Borkar and Sundaresan (2012, lemma A.6) holds verbatim. □
Finally, the next result gives the limiting behavior of the empirical measure vector μN when we first let t go to infinity and then let N go to infinity.
(Freidlin and Wentzell 2012, chapter 6, theorem 4.1). Assume that Assumption 2 holds true. Then, for any , there exists , which can be chosen arbitrarily small, such that the -measure of the r1-neighborhood γi of the compact Ki satisfies
Choose small positive such that the inequalities in (A.3), (A.1), and (A.2) are satisfied for sufficiently large N with replacing ε. One can notice that the Markov chain Zn is irreducible and thus has an invariant measure. By Freidlin and Wentzell (2012, chapter 6, lemma 3.2), and by selecting the exponential terms that decrease more slowly, when N is large enough, the values of the normalized invariant measure of the chain Zn lie in the interval
Recall the following formula, which expresses, up to a factor, the invariant measure of the process μN in terms of the invariant measure of the chain Zn on (see Freidlin and Wentzell 2012, chapter 6, equation (4.1)):
Using the estimates in (A.2) and (A.1), we find that
By summing over , we get
Using again the formula (A.4) together with the definition of the stopping times σn and τn, we find
By inequality (A.1), we find that , and by Corollary 3 we have that is bounded above by some constant. Using this together with the lower bound in (A.6), we obtain, by dividing in (A.5) by , the lower and upper bounds for the normalized measure of the r1-neighborhood γi. The theorem is proven. □
Appendix B: Technical Proofs
We give here the proof of some technical results.
B.1. Proof of Lemma 1
This is a mild generalization of Borkar and Sundaresan (2012, lemma 3.2) to the multi-populations case. We prove each of the three assertions, respectively.
Let us ignore for the moment the constraint by supposing that all possible transitions are allowed. Let with and , and consider the constant velocity path given by
(B.1)from which we can see that(B.2)a constant velocity. Note that (B.1) can be rewritten as follows. For all and t in(B.3)(B.4)Now we construct rate matrices and that ensure the traversal of this constant velocity path. Note that there is conservation of mass between the initial and terminal states, and thus for all and . Therefore, on can construct mass transport parameters and such that for any two states with (resp. ) and with (resp. ), the quantity (resp. ) is the fraction of the excess of mass (resp. ) that goes from z to . In particular, the coefficients for satisfy, for all , the following conditions:
(B.5)(B.6)(B.7)and(B.8)The first condition follows from the definition of the fractions . The second attests that no mass is transferred from a state with no mass excess, and no mass is received by a state with mass excess. The third point tells us that there is no mass destruction, and the last point stipulates that no mass is created. By using these mass transfer parameters, we construct the rate matrices and as follows. The diagonal elements are given, for , by
(B.9)for all that satisfy and otherwise. The off-diagonal elements are given by(B.10)if and , and if , for any . Using these definitions of the rate matrices and , and the properties of the mass transport coefficients and , it is easy to prove that for and (see Borkar and Sundaresan 2012, p. 360),Let us now evaluate the difficulty of the passage at this constant velocity path μ. Theorem 1 tells us that if , then is given by (3.18), and this is particularly true if the 2r integral terms in (3.18) are finite. Notice that if, for a given j, and for , then is not bounded. Therefore, we need to provide a bound for (3.18) in the case of the constant velocity path. We start by upper bounding the sums in (3.18) for which the rates are strictly positive on . To this end we introduce the set
Thus, from the definition of the Legendre transform , the right-hand side of (3.18) can be written as
(B.11)From (B.9) and (B.10), we find that . Moreover, from Assumption 1, we obtain . Therefore, (B.11) is upper bounded by
(B.12)which in turn is bounded by(B.13)with referring to the total variation distance. Observe that for , the function is upper bounded, say by . Moreover, for a fixed z, . Therefore, the terms are bounded by(B.14)On the other hand, using the change of variable and then (B.2), one obtains
(B.15)Therefore, (B.12) is upper bounded by
(B.16)where and are two constants that do not depend on and for any and .Let us now consider the case where the pairs , that is, for which . Using , the corresponding integral terms in (3.18) become
(B.17)Combining this with (B.16) gives for some positive constants and that do not depend on and .
Finally, let us consider the case where only transitions in are allowed. Because the directed graph is irreducible, the Markov chain is also irreducible, and thus there exists a finite sequence of intermediate points through which one can move from ν to ξ in steps
Therefore, one can construct a piecewise linear path μ with constant velocities on each of the m segments such that each segment is covered in time duration T/m. Indeed, define for each the constant velocity path
and take for . Hence,which completes the proof of the first assertion.This follows immediately from the previous result and (3.22).
Fix and take in (B.16). We can always find a such that, if , then (B.16) is bounded by
(B.18)
Combining this with (B.17) gives that, for any such that ,
The result then follows from (3.22). □
B.2. Proof of Lemma 2
We generalize Borkar and Sundaresan (2012, lemma 7.1). From Theorem 1, because , the rate function is given by (3.18). Moreover, we can verify that for all . Using this together with (3.18) gives to us
B.3. Proof of Lemma 3
We generalize Borkar and Sundaresan (2012, lemma 7.2). By construction, we have that and . Fix . Using we find that, for all ,
From (3.11), and using the properties of the function, we can easily verify that
Therefore, we find
Introducing the change of variables , we obtain
Finally, using Assumption 1, we find (3.23). □
References
- (2008) A class of mean field interaction models for computer and communication systems. Perform. Eval. 65(11–12):823–838.Google Scholar
- (1999) Convergence of Probability Measures (John Wiley & Sons, New York).Google Scholar
- (2011) Small noise asymptotics for invariant densities for a class of diffusions: A control theoretic view (with erratum). Preprint, submitted, https://arxiv.org/abs/1107.2277.Google Scholar
- (2012) Asymptotics of the invariant measure in mean field models with jumps. Stoch. Syst. 2(2):322–380.Link, Google Scholar
- (2016) Perturbative calculation of quasi-potential in non-equilibrium diffusions: A mean-field example. J. Stat. Phys. 163:1157–1210.Google Scholar
- (2019) Partial mean field limits in heterogeneous networks. Stochastic Process. Appl. 129:4998–5036.Google Scholar
- (2014) Macroscopic Limit of a bipartite Curie-Weiss model: A dynamical approach. J. Stat. Phys. 157:1309–1319.Google Scholar
- (2016) Rhythmic behavior in a two-population mean-field Ising model. Phys. Rev. E. 94:042139.Google Scholar
- (1990) On the long term behavior of some finite particle systems. Probab. Theory Related Fields 85:195–237.Google Scholar
- (1987) Large deviations from the McKean-Vlasov limit for weakly interacting diffusions. Stochastics 20(4):247–308.Google Scholar
- (1989) Large deviations, free energy functional and quasipotential for a mean field model of interacting diffusions. Mem. Amer. Math. Soc. 78(398).Google Scholar
- (1993) Hierarchical models of interacting diffusions: Multiple time scale phenomena, phase transition, and pattern of cluster-formation. Probab. Theory Related Fields 96:435–473.Google Scholar
- (1999) Hierarchically interacting Fleming-Viot processes with selection and mutation: Multiple space time scale analysis and quasi-equilibria. Electron. J. Probab. 4:1–81.Google Scholar
- (1995) Equilibria and quasi-equilibria for infinite collections of interacting Fleming-Viot processes. Trans. Amer. Math. Soc. 347(7):2277–2360.Google Scholar
- (2020) Propagation of chaos and large deviations in mean-field models with jumps on block-structured networks. Preprint, submitted December 4, https://arxiv.org/abs/2012.02870.Google Scholar
- (2010) Large Deviations Techniques and Applications, 2nd ed. (Springer, New York).Google Scholar
- (2013) Longtime Behavior for Mutually Catalytic Branching with Negative Correlations, vol. 38 (Springer Proceedings in Mathematics & Statistics, Boston).Google Scholar
- (2006) Large Deviations for Stochastic Processes (Mathematical Surveys and Monographs, American Mathematical Society).Google Scholar
- (1994) Large deviations for empirical process of mean-field interacting particle system with unbounded jumps. Ann. Probab. 22(4):2122–2151.Google Scholar
- (2012) Random Perturbations of Dynamical Systems, 3rd ed. (Springer, Berlin-Heidelberg).Google Scholar
- (2008) Chaoticity for multiclass systems and exchangeability within classes. J. Appl. Probab. 45:1196–1203.Google Scholar
- (2007) Phase transitions for the long-time behavior of interacting diffusions. Ann. Probab. 35(4):1250–1306.Google Scholar
- (1990) Large-time behavior of perturbed diffusion Markov processes with applications to the second eigenvalue problem for Fokker-Planck operators and simulated annealing. Acta Appl. Math. 19:253–295.Google Scholar
- (2020) Fluctuation results for general block spin Ising models. J. Stat. Phys. 178(1):1175–1200.Google Scholar
- (2015) Multiple Time Scale Dynamics (Springer International Publishing, Switzerland).Google Scholar
- (1995a) Large deviations for long range interacting particle systems with jumps. Annales de l’I.H.P. Probabilités et Statistiques, Tome 31 31(2):289–323.Google Scholar
- (1995b) On large deviations for particle systems associated with spatially homogeneous Boltzmann type equations. Probab. Theory Related Fields 101:1–44.Google Scholar
- (2020) Two-community noisy Kuramoto model. Nonlinearity 33(4):1847–1880.Google Scholar
- (2020) On mean field systems with multi-classes. Discrete Contin. Dyn. Syst. 40(2):683–707.Google Scholar
- (2017) Potential landscape of high dimensional nonlinear stochastic dynamics with large noise. Sci. Rep. 7(15762):1157–1210.Google Scholar
- (2019) Large time behaviour and the second eigenvalue problem for finite state mean-field interacting particle systems. Preprint, submitted September 9, https://arxiv.org/abs/1909.03805.Google Scholar
- (2021) A sufficient condition for the quasipotential to be the rate function of the invariant measure of countable-state mean-field interacting particle systems. Preprint, submitted October 25, https://arxiv.org/abs/2110.12640.Google Scholar
- (2012) Quasi-potential landscape in complex multi-stable systems. R. Soc. Interface 9:3539–3553.Google Scholar

