E[g(X)] = sum of g(x) P(X = x), or the integral of g(x) f(x) dx. You do not need the distribution of g(X) itself.
The trap it prevents, and the one it causes. E[g(X)] is generally not g(E[X]). E[X^2] exceeds (E[X])^2 by exactly the variance, which is the most-used instance.
The nickname comes from students applying it without realising it needs proof - but the substantive point for an interview is the inequality direction, which follows from Jensen and explains why convexity has value.