AI-generated tests assert what the code does, not what it should do — so they pass on bugs and fail on fixes. Why coverage broke, and what to measure instead.