Here is a proof that convergence in proba implies convergence weakly, without having to deal with «-distance sets» manually. I don’t claim novelty; this is about writing stuff down as a way to understand it, and figure out the relation between the pieces.
Let’s first factor out a few too many simple lemmas:
If is a sequence of random variables (almost surely) in , then the following are equivalent:
The proof uses the Markov inequality for one direction, and splitting the expectation for the other (OK, that’s where the «-distance sets» are hiding):
Assume first that and fix . We have by assumption, hence .
Conversely, assume that for all , . Then, for any by assumption that . Since this holds for any , we conclude
If is a sequence of random variables (almost surely) , then, for any the following are equivalent:
Clearly (1)(2). Conversely, assuming (2) holds, for , while for , with chosen arbitrarily.
In other words, the convergence
is monotonous in
:
if it holds for
,
it holds for any
.
This implies we only need to check it for
,
for which it doesn’t matter that we truncate with
.
If are random variables (in some metric space), then the following are equivalent (for any ):
This is part of (Kallenberg 2021, lemma 5.2 p102).
We have the following sequence of equivalent statements:
A simple consequence is
If for all , then iff .
(omitted)
We can now prove that convergence in proba implies convergence in distribution.
By the Portmanteau theorem, recall that iff for any countinuous bounded function , we have . So, fix any such . By the continuous mapping theorem, since , we also have ; furthermore are bounded, say by . Then as needed.
Here is the morale I get out of this story: For bounded random variables (in ), convergence in probability is convergence to zero of their expected distance (lemma 4). Convergence weakly is convergence of the expected values, under any bounded continuous function. Since those preserve convergence in probability, and force boundedness, we can use the aforementioned characterization of convergence in proba (and the triangle inequality).
Admittedly, with all the extracted lemmas above taken together, this makes for a long proof of , but lemma 3 can be considered known, and that’s really all that’s needed for the proof.
Actually, after reading this math.se question, I think lemma 1 can be adapted to cover s such that with (the proof goes through by simply replacing by ). This means that for non-negative dominated sequences , convergence is equivalent to convergence in proba, in both cases to . Then, given , with dominated by with , we can apply the adapted lemma 1 to , so that This is pretty neat.
denotes the minimum↩︎