2.3 Large Sample Theory

2.3.1 Delta Method

Proof. By the definition of the derivative, we have that \[ \phi (t) = \phi (\theta ) + \phi '(\theta )(t - \theta ) + o(\|t - \theta \|), \] i.e. \begin{equation} \phi (t) = \phi (\theta ) + \phi '(\theta )(t - \theta ) + R(\|t - \theta \|) \end{equation} where \(\lim _{h \to 0} \frac {R(h)}{h} = 0\). Since \(r_n(T_n - \theta )\) converges in distribution, we know that \(r_n(T_n - \theta ) = O_p(1)\), which implies that \(r_n \|T_n - \theta \| = O_p(1)\). We also have that \(\|T_n - \theta \| = o_p(1)\), which implies \(R(\|T_n - \theta \|) = o_p(\|T_n - \theta \|)\). Thus \[ r_n R(\|T_n - \theta \|) = r_n o_p(\|T_n - \theta \|) = o_p(r_n \|T_n - \theta \|) = o_p(O_p(1)) = o_p(1). \] Using this along with (6), we have the second part of the theorem. Noting that \(r_n \phi '(\theta )(T_n - \theta ) \xrightarrow {d} \phi '(\theta ) T\), and applying Slutsky’s theorem, we get the first part as well. □

Proof. By definition, \[ \phi (t) = \phi (\theta ) + \boldsymbol {\nabla } \phi (\theta )^\top (t - \theta ) + \frac {1}{2}(t - \theta )^\top \boldsymbol {\nabla }^2 \phi (\theta )(t - \theta ) + R(\|t - \theta \|), \] where \(R(h) = o(\|h\|^2)\). Since \(\boldsymbol {\nabla } \phi (\theta ) = 0\), we actually have \begin{equation} \phi (t) = \phi (\theta ) + \frac {1}{2}(t - \theta )^\top \boldsymbol {\nabla }^2 \phi (\theta )(t - \theta ) + R(\|t - \theta \|). \end{equation} Note \(r_n^2 R(\|T_n - \theta \|) = r_n^2 o_p(\|T_n - \theta \|^2) = o_p(\|r_n(T_n - \theta )\|^2)\). Since \(r_n(T_n - \theta )\) converges in distribution, so does \(\|r_n(T_n - \theta )\|^2\), and so \(\|r_n(T_n - \theta )\|^2 = O_p(1)\). Thus \begin{equation} r_n^2 R(\|T_n - \theta \|) = o_p(O_p(1)) = o_p(1). \end{equation} Now by the continuous mapping theorem, we have that \begin{equation} \frac {1}{2}(r_n(T_n - \theta ))^\top \boldsymbol {\nabla }^2 \phi (\theta )(r_n(T_n - \theta )) \xrightarrow {d} \frac {1}{2} T^\top \boldsymbol {\nabla }^2 \phi (\theta ) T. \end{equation} So combining (7), (8), (9) and using Slutsky’s lemma, we get the desired convergence in distribution. □

2.3.2 Fisher Information

Before we do anything, we have to make several assumptions.

1.
We have a “nice, smooth” model, i.e. the Hessian is Lipschitz-continuous. To be rigorous, the following must hold: \[ \big \|\nabla ^2 \ell _{\theta _1}(x) - \nabla ^2 \ell _{\theta _2}(x)\big \|_{\mathrm {op}} \le M(x)\,\|\theta _1 - \theta _2\|_2 \qquad \mathbb {E}_\theta [M^2(x)] < \infty \]
2.
The MLE, \(\hat {\theta }_n \in \arg \max _{\theta \in \Theta } P_n \ell _\theta (x)\), is consistent, i.e. \(\hat {\theta }_n \xrightarrow {p} \theta _0\) under \(P_{\theta _0}\).
3.
\(\Theta \) is a convex set.

Proof. Set \(\Psi (x) = \nabla _\theta \ell _\theta (x)\), we have that \(\mathbb {E}_\theta [\Psi ] = 0\), and that \begin{align*} \mathbb {E}[(\delta - g(\theta ))\Psi ] &= \mathbb {E}[\delta \Psi ] \\ &= \mathbb {E}[\delta \nabla \ell _\theta ] \\ &= \mathbb {E}\!\left [\delta \frac {\nabla p_\theta }{p_\theta }\right ] \\ &= \int \delta \nabla p_\theta \, d\mu (x). \end{align*}

Under good regularity conditions, we have that \[ \mathbb {E}[(\delta - g(\theta ))\Psi ] = \nabla \int \delta (x) p_\theta (x)\, d\mu (x) = \nabla g(\theta ). \]

We take \[ \gamma = \nabla g(\theta ), \quad C = I_\theta \] to get the desired result.

Proof. Take \begin{align*} g(\theta ) &= v^\top \theta \\ \delta &= v^\top \hat {\theta }(X). \end{align*}

Applying the Cramer-Rao theorem, \[ \mathbb {E}\!\left [(v^\top (\hat {\theta } - \theta ))^2\right ] \ge v^\top I_\theta ^{-1} v \] and \[ \mathbb {E}\!\left [(v^\top (\hat {\theta } - \theta ))^2\right ] = \mathbb {E}\!\left [\mathrm {tr}\!\left ((\hat {\theta } - \theta )(\hat {\theta } - \theta )^\top v v^\top \right )\right ] = v^\top \mathrm {Cov}(\hat {\theta })\, v. \]

Search definitions, theorems, and topics across the notes.