unlike the method of moments, which matches only the first
k sample moments (biased in small samples, possibly inefficient), MLE is asymptotically unbiased, consistent and efficient under regularity conditions, achieving the information-theoretic lower bound: