A school system that wants to know whether its students are learning must measure something, and what it can measure most cheaply is performance on a standardised test. The appeal is obvious. A common instrument allows comparison across schools that differ in everything else, and it replaces the judgment of individual teachers, which is variable and sometimes prejudiced, with a procedure that treats every candidate identically.
The difficulty begins when the measure is also used to allocate rewards. Once a school's funding, or a teacher's promotion, depends on test scores, effort shifts towards whatever raises scores. Some of that effort raises learning too, and some of it does not: narrowing the syllabus to the tested subjects, coaching in question formats, and in the worst cases, discouraging weak students from sitting the examination at all. A score that rises for these reasons no longer tells you what it told you when nothing depended on it.
This is sometimes stated as a general law: a measure that becomes a target ceases to be a good measure. But the law, stated that broadly, would counsel abandoning measurement altogether, which cannot be right either. A system that measures nothing does not thereby become fair; it simply allocates rewards on the basis of reputation, inspection visits and whatever the inspector happened to see.
The useful question is narrower. It is whether a particular measure can be gamed more easily than the underlying quality can be improved. Where gaming is cheap and improvement is expensive, the measure will degrade quickly. Where the only practical way to raise the number is to teach better, it will not. That is a question about the design of the instrument, and it has different answers for different instruments.