How training can produce increasingly impressive performance while producing progressively less of the learning it was designed to create. (Written for Smartermarx Magazine Sept. 2026)
A Sad Tale of Tails
In 1902, French colonial authorities in Hanoi (capital city of Vietnam) faced a substantial rat problem. The issue became particularly concerning amid Hanoi’s modernization efforts, as a newly developed sewer system (intended as a conspicuous achievement of French colonial engineering) also created an extraordinarily accommodating habitat and transportation network for the city’s rodent population. With concerns about bubonic plague mounting, colonial authorities began an aggressive extermination campaign, and thousands of rats were killed. On some days, the numbers reportedly climbed well into the tens of thousands. Unfortunately for all involved, killing a remarkable number of rats and solving a rat problem turned out to be two very different things.
In an effort to expand the campaign, authorities eventually offered civilians a bounty for each rat killed. Understandably, no one was particularly interested in having thousands of decomposing rat carcasses delivered to municipal offices, so officials settled on a (ugh…) more convenient bit of evidence: the tail. Turn in a rat tail, collect the bounty, and everyone could reasonably assume that somewhere there was now one less rat. At first glance, it was a pretty sensible arrangement. A rat tail was easy to identify and count, seemingly difficult for a rat to surrender voluntarily, and presumably a fairly reliable indication that its former owner had met an unfortunate end. Unfortunately, the people of Hanoi soon discovered there was a serious flaw in the arrangement.

Sadly, over time, live rats began appearing without tails. Rather than kill the animals, some enterprising participants simply removed the part with monetary value and released the rats, leaving them capable of producing future generations of similarly profitable tails. Historical accounts even describe reports of rats being imported into the city and breeding operations appearing outside Hanoi. The authorities had wanted fewer rats. What they had actually incentivized was producing rat tails. To be fair, the rat tail had not been a foolish thing to measure. Under ordinary circumstances, the number of tails collected would likely correlate well with the number of rats killed. The problem arose when success became increasingly defined by the measure itself. Once the proxy—the tail—became the target, people found increasingly effective ways to produce the desired number without necessarily producing the condition that number was intended to represent.
This “rat-tail problem” turns up in many other pursuits, including–you guessed it–art education. For example, an accurately replicated shape, an immaculate drawing, a beautifully controlled gradation, the number of completed studies an aspiring artist creates, or even the speed with which an exercise can be successfully completed may all provide useful information about a learner’s developing specific desired abilities. In the right context, each can serve as a meaningful indicator of progress in the appropriate context. However, we can sometimes easily forget that none of these things is learning itself.

That distinction can become especially important within a skills-based curriculum. If an exercise is designed to cultivate a particular perceptual, cognitive, or motor capability, then the visible product of the exercise is often one of our most accessible indicators that the desired development is occurring. The danger emerges when the teacher or learner treats that evidence as the developmental goal itself. At that point, it becomes possible to get considerably better at producing the appearance of successful training without becoming correspondingly better at the thing the training was designed to develop. And, as those unfortunate rats in Hanoi might attest, people can become remarkably good at giving us exactly what we ask for.
From Monetary Policy to a General Problem of Measurement
Before looking more closely at this issue in art education, let’s take a moment to better understand its history. More than seventy years after the Hanoi rat debacle, this general type of problem would take on a more formal identity in a different context: British monetary policy.
In 1975, economist Charles Goodhart (then an adviser at the Bank of England) was considering a problem that can arise when governments rely on statistical indicators to guide economic policy. He noticed that a relationship could look fairly reliable as long as people were simply observing it, but could begin to break down once policymakers started using it to influence or control behavior. His original formulation was considerably less catchy than the version most often repeated today: “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.”
Put more simply, the relationship between a measure and what it represents can change once people know that something important depends on it. If rewards, penalties, funding, promotions, or other consequences are tied to a particular number, people naturally begin adjusting their behavior to improve that number. Sometimes those adjustments also improve the thing the number was meant to represent. Sometimes they do not. And when people become increasingly successful at improving the measure without improving the underlying condition, the measure becomes less trustworthy as an indicator of success.
At almost exactly the same time, American social scientist Donald T. Campbell was developing a remarkably similar warning in the context of social programs and their evaluation. Campbell developed his work independently (an early version was presented in 1974, and Assessing the Impact of Planned Social Change appeared as an occasional paper in 1976) so it would be misleading to characterize his contribution as an extension of Goodhart’s. Campbell’s formulation was particularly relevant to education. He argued that quantitative social indicators become increasingly vulnerable to distortion as greater consequences are attached to them. He then offered achievement testing as an example. Test scores, Campbell reasoned, may be perfectly useful indicators of broader educational competence when instruction is actually directed toward developing that competence. However, when producing the test score itself becomes the goal, teaching can begin changing in ways that improve the score without producing a corresponding improvement in the broader capability the score was intended to indicate.
If that sounds suspiciously like our tale of rat tails, it should. A couple of decades later, anthropologist Marilyn Strathern gave Goodhart’s observation the beautifully compact formulation by which it is probably best known today. Writing in 1997 about assessment, accountability, and what she described as an expanding “audit culture” within British universities, Strathern wrote:
“When a measure becomes a target, it ceases to be a good measure.”
That sentence traveled considerably farther than the wordier rule of monetary policy. Today, Goodhart’s Law, as it has come to be known, is invoked in discussions of education, healthcare, business, policing, scientific publishing, artificial intelligence, organizational management, and just about any other arena in which something complicated must be represented by something comparatively easy to count.
However, there is an important qualification. Goodhart’s Law is not an argument against measurement. I cannot stress that enough. In fact, I would argue that abandoning measurement in art education would be an incredibly poor response to this issue. If we want to make meaningful claims about learning, we need some way to distinguish actual changes in performance from intuition, preference, or the vague impression that someone seems to be “getting better.” Clear learning outcomes, observable behaviors, comparative assessments, and performance standards can provide essential information about whether training is doing what we believe it is doing. Within our own curriculum, measurable outcomes are deliberately tied to individual exercises precisely so that instructors and learners can identify progress, diagnose gaps, and make informed decisions about what should happen next.
The problem, then, is not that we measure. The problem begins when we forget what the measurement is for. An exercise performance, test result, completion time, or finished drawing can be valuable evidence of a developing capability. But the evidence and the capability are not interchangeable. Once we begin optimizing the former without continually checking its relationship to the latter, an apparently improving measure can begin concealing stagnant—or even diminished—learning. And this leaves us with a deceptively simple question for the studio:
When a student becomes better at the thing we can easily measure, are they actually becoming better at the thing we intended to teach?
In Regard to Skill Training
This distinction becomes particularly important when we turn our attention to skill training, where it is easy to conflate three related but different things: the exercise, performance on the exercise, and learning.
An exercise is simply the training environment we create. It presents a learner with a particular problem, set of constraints, or collection of experiences intended to produce useful development.
Performance is what the learner does within that environment. It is the observable result: how accurately a shape is replicated, how smoothly a gradation is produced, how consistently a line is executed, how efficiently a task is completed, or how successfully some other criterion is satisfied.
Learning, however, is the change we hope those experiences produce in the learner—the acquisition, modification, or refinement of knowledge, skills, behaviors, or useful perceptual and motor associations through experience and practice. Within our curriculum, learning is understood as something that changes how a task can subsequently be understood, approached, and performed.
These three aspects of development obviously overlap. If an exercise is well designed, thoughtfully repeated engagement with it should produce better performance, and that improved performance should provide evidence that useful learning is taking place. However, as easy as it might be to think so, these things are not synonymous. A student can improve performance on an exercise without necessarily developing all the capabilities the exercise was intended to cultivate. They may discover a highly specific strategy that works extremely well under the task’s particular conditions. They may learn a model’s quirks, memorize a procedure, become dependent on a scaffold, or simply grow very efficient at meeting the criteria by which success is judged. Their measurable performance improves—and perhaps dramatically so—but what exactly has been learned?
One useful way to answer that question is to see what happens when the conditions change. This is one reason we incorporate elements of what is called interleaved learning into our own curriculum. Interleaved learning is a practice strategy in which related tasks, problem types, or conditions are alternated rather than practiced exclusively in uninterrupted blocks. This is not an argument for interleaving over blocked practice; we use both deliberately. Blocked practice has real advantages—particularly when establishing early fluency through repetition—while interleaving periodically requires the learner to retrieve, discriminate among, and adapt developing strategies as conditions change. That added difficulty can temporarily suppress practice performance, yet in many tasks appropriately designed interleaving improves later discrimination, retention, and transfer. Within our curriculum, blocked practice helps establish stable perceptual-motor associations, while strategically introduced variation helps test whether those associations remain useful when the problem no longer looks exactly like the one on which they were learned.
Can the learner apply the same underlying capability to a new subject, a different orientation, an unfamiliar configuration, or a more complex problem? Can they continue to perform when a particular guide or scaffold is removed? Does the skill remain useful when it must be combined with other developing abilities?
This is the problem of transfer. Transfer refers to the ability to apply something learned in one context to a new or different context. People often assume transfer will occur as a natural consequence of practice, but that is not necessarily the case. Within our curriculum, early exercises are therefore not intended as terminal achievements. Shape Replication, for example, builds perceptual-motor skills that can later be reapplied to establishing proportions, spatial relationships, form construction, and increasingly complex representational problems. The curriculum deliberately revisits and recombines earlier skills because successful performance in one tightly constrained setting does not guarantee useful performance elsewhere.
The Shape Replication exercise illustrates the distinction particularly well, and the next section will explain it in more detail if you are not familiar with it. Its purpose is not merely to produce shapes that closely match a model. Through repeated practice, learners are meant to build useful associations between perceived angles, proportions, curvatures, and the motor behaviors that can reproduce them. The finished shape gives us useful evidence that this process may be occurring, but it is not the process itself.
This is why a student who produces a slightly less accurate shape while developing increasingly adaptable perceptual-motor relationships may, in an important sense, be learning more than a student who produces a nearly perfect shape through excessive reliance on strengtheners or other task-specific strategies. This difference may not be obvious when you place the products of both approaches side by side. The difference often becomes more apparent later, when the learner encounters a problem for which those supports or highly specific strategies are no longer available. The ultimate test of a training exercise is not how good the student becomes at the exercise. It is what remains when the exercise is gone. That is where performance begins to tell us something meaningful about learning—and where Goodhart’s Law becomes particularly useful for understanding how a seemingly successful curriculum can quietly drift away from the abilities it was designed to cultivate.
A Goodhart Trap in Shape Replication
One of the clearest examples of how Goodhart’s Law can emerge within our own curriculum appears fairly early in the exercise we call Shape Replication. For those unfamiliar with the exercise, its basic structure is quite simple. A learner is presented with a transparent model sheet containing a series of two-dimensional shapes, usually configurations of straight and curved lines contained within small rectangular boundary boxes. Next to the model, the learner establishes corresponding blank boxes (with measurement tools) and tries to reproduce each shape inside as faithfully as possible. Early efforts rely heavily on visual estimation and comparison, although a variety of measurement strategies can be introduced when appropriate.
The boundary box itself (like an envelope) is important. It gives the learner a stable frame of reference against which to judge the shape’s location, angle, proportion, and direction. For example, a line that appears to enter the left side of a box about one-third of the way down can be compared against that boundary rather than evaluated in isolation. The box therefore acts as both a spatial constraint and a perceptual anchor, helping the learner make relational judgments about where things are located and how they relate to one another.

After a group of replications is completed, the learner can check the results using the transparent model sheet as an overlay—essentially a clear sheet carrying the correct target configuration that can be placed directly over the drawing. Any significant deviation between the attempted shape and the model becomes immediately visible. In the default exercise, we generally delay these checks until a learner has completed a column of four shapes so the overlay functions as feedback rather than a mechanism for constant correction.
At first glance, it’s easy to assume the purpose of all this is simply to teach the student to copy shapes more accurately. Accuracy certainly matters, and the transparency gives us a convenient and comparatively objective way to assess it. But accuracy is not the primary developmental target. The more important goal is what we describe as perceptual-motor mapping.
Human vision does not provide us with a perfectly objective measurement of the world. Learning to draw does not somehow remove the contextual influences, biases, and constructive processes that make perception fundamentally non-veridical. What training can do, however, is help us establish increasingly useful associations between what we perceive and how we physically respond to it. With repeated Shape Replication practice, a learner begins to associate particular perceived angles, proportions, curvatures, distances, and directional relationships with increasingly familiar patterns of motor behavior. Put simply: when I experience something that looks like this angle, this series of hand and tool movements tends to produce a mark that looks like that angle. With enough varied, feedback-rich experience, those relationships can become increasingly stable, efficient, and adaptable. The curriculum therefore treats Shape Replication less as a simple copying exercise (although it can be colloquially labeled as such) than as a platform for calibrating perception and action.
This distinction gives us the ingredients for a Goodhart problem. The developmental target is a useful and increasingly adaptable relationship between perceptual experience and motor response. One convenient indicator of that development is replication accuracy. However, the journey toward more stable calibrations of perception and action also often brings an array of perfectly legitimate strengtheners. If a learner consistently struggles to locate where particular lines should begin and end, an instructor may suggest adding origin-and-destination points along the perimeter of the boundary box. These points can be established with a ruler, divider, or caliper and give the learner additional spatial anchors. Additional air tracing, further subdivision of the reference, and more frequent transparency checks can likewise be introduced when a learner needs greater support.
There is nothing inherently problematic about any of this. The additional points can reduce the task’s difficulty enough to keep a struggling learner productively engaged within what we call an appropriate Zone of Proximal Learning (ZPL)—our operational adaptation of Vygotsky’s Zone of Proximal Development. They can help focus attention on particular relationships, provide clearer feedback, and allow the learner to accumulate successful perceptual-motor experiences that might otherwise remain out of reach. But now imagine what happens if accuracy quietly becomes the dominant goal—even though accuracy is, quite legitimately, one of the things the learner is trying to pursue. Two additional points improve the match, so perhaps four would improve it further. Four become eight. Eight become twelve. Before long, a complex shape that originally required the learner to judge larger angles, proportions, distances, and relationships has effectively become a connect-the-dots exercise. And the transparency overlay may reward the strategy beautifully. The replicated shape becomes more accurate. The errors become smaller. The measurable performance improves. Yet the learner may now be spending progressively less time confronting the very perceptual-motor problem the exercise was designed to provide.
This is where the distinction between a strengthener and a dependency becomes important. The problem is not that the additional dots make the exercise easier. Well-designed scaffolding should make an otherwise unmanageable challenge more accessible. The problem emerges when the support begins replacing rather than supporting the targeted experience—when increasingly successful performance depends upon a strategy that will not accompany the learner into the broader situations for which the training is intended. In fact, our teaching manual explicitly warns instructors to watch for an overemphasis on accuracy at the expense of other goals. Accuracy matters, but it sits within a hierarchy of priorities rather than above them. Learners can begin sacrificing light, deliberate mark-making, and consistent line quality in pursuit of a “pixel-perfect” match. They can also memorize particular motor patterns if shapes are repeated without sufficient variation. For this reason, Shape Replication exercises incorporate changes in orientation and other variations intended to keep the learner responding to the perceptual problem rather than simply memorizing a solution.
This is a small but useful example of why Goodhart effects do not require dishonesty, laziness, or any deliberate attempt to “game the system.” A student adding more and more measurement points may be doing exactly what we normally encourage learners to do: trying very hard to improve. They see a measurable disparity, discover a strategy that reduces it, and understandably conclude that more of the successful strategy should produce even better results. And according to the most obvious measure, it does. The problem is that the measure and the intended learning have begun to separate. The student has become better at producing the evidence we associate with perceptual-motor development while potentially spending less time engaging the experience intended to produce that development. In our Hanoi analogy, the tail count is going up beautifully. We just need to make sure we are still getting fewer rats.
The Scaffold That Eats the Lesson
The Shape Replication example points toward a broader instructional problem. A support (strengthener or scaffold) can improve immediate performance while simultaneously reducing engagement with the very difficulty responsible for the intended learning. This does not make scaffolding undesirable. In fact, it is quite the opposite. Appropriate scaffolding is fundamental to how we structure learning within our curriculum. A task that is far beyond a learner’s present abilities may produce little more than confusion or frustration. A thoughtfully chosen support element can reduce selected aspects of that difficulty enough to keep the learner operating within an appropriate Zone of Proximal Learning—challenged beyond what they can comfortably accomplish alone, but still able to make meaningful progress with guidance.
In our curriculum, we often refer to these supports as strengtheners. Depending on the exercise and the learner, a strengthener might involve additional measurement points, air tracing, a transparency overlay, a grid or axis guide, physical masking, a printed reference model, more frequent feedback, or any number of other temporary aids. These tools are not concessions or lesser forms of learning. They adjust a problem so the learner can engage meaningfully with the aspect we are trying to develop. The important question is what, exactly, is the support supporting?
I have written elsewhere about the somewhat careless way artists often use the word crutch. A ruler, grid, photograph, projector, tracing, or any other tool has no intrinsic property that makes it a “crutch.” The designation only becomes meaningful relative to a particular goal. A useful operational definition is this: a crutch is a tool, strategy, or process that compensates for a skill deficit by bypassing essential perceptual, cognitive, or motor work. In evaluative or representational contexts, a crutch is any aid that compensates for a skill that is implied by the final product.
Consider something as familiar as a grid. If the immediate training objective is to develop more useful, larger-scale freehand judgments of proportion and spatial relationships, a sufficiently dense grid may allow the learner to avoid a meaningful amount of precisely that work. But if another exercise aims to isolate value organization, edge control, or material handling, establishing the underlying structure with a grid may be a sensible way to keep proportion problems from consuming attention that should be invested elsewhere. The nature of the grid has not changed; the learning objective has.
The same is true of transparency overlays, measurement devices, tracing, masking, instructor demonstrations, and just about every other form of support we might introduce. We cannot determine their educational value simply by asking whether they make the task easier. We must ask whether they make the task easier in the right way for the right reasons. This is also why dependency, by itself, is not necessarily the problem. Artists depend on brushes, easels, reference materials, lighting systems, measuring devices, and countless other technologies. Dependence on a tool tells us very little unless we first identify the capability being developed or assessed. The relevant concern in a training environment is whether continued dependence on a particular support is beginning to replace the developmental work that support was initially introduced to make accessible.
That distinction can be surprisingly hard to detect because the finished product may keep improving. A student using increasingly dense measurement scaffolds may produce a more accurate drawing. A learner checking a transparency after every mark may produce fewer visible errors. An instructor who intervenes at every difficult moment may help generate beautifully polished exercises. In each case, the immediate performance can look better while the learner is being asked to solve progressively less of the problem independently. This is precisely why effective scaffolding generally includes some expectation that support will be adjusted or faded as performance stabilizes. Our teaching manual repeatedly applies this principle: printed models, overlays, isolators, and other aids may be introduced when useful, but reliance decreases as the targeted capability strengthens. The larger instructional aim is adaptive performance, not permanent proficiency under one carefully supported set of conditions.
Fading support does something else that matters in the context of Goodhart’s Law: it tests what the learner actually owns. Remove a few perimeter points. Delay the transparency check. Rotate the model. Change the subject. Ask the learner to diagnose the problem before the instructor supplies the answer. If performance immediately collapses, that does not necessarily mean the scaffold was a mistake. It may simply reveal that the support still does more of the work than the learner can do alone. That is useful information. From this perspective, a good strengthener does not eliminate the important problem. It presents the problem at a resolution the learner can currently manage. As ability develops, the support can be reduced, and the learner can assume more of the burden. A problematic support does something different. It allows the learner to become increasingly successful while continually avoiding the capability the exercise was intended to develop. And once again, the finished products may not tell us which of those two things has happened. That is why the instructor cannot simply ask, “Did this tool improve the drawing?” The more useful question is: “Did this tool improve the learner?”
Why Struggle Can Be Better Evidence Than Success
Something many of my students have heard me say over the years is that, while I certainly encourage everyone to strive for pristine success in any exercise, I am quite pleased by a completed challenge that looks more like a metaphorical battlefield—full of corrections, abandoned approaches, misjudgments, retries, and other visible evidence that the learner had to fight for every bit of progress. The reason is fairly simple: a clean exercise tells me that a student performed well. One with signs of a struggle may tell me that a great deal of learning took place.
Of course, this is not an argument that failure is somehow preferable to success, nor that frustration should be manufactured for its own sake. Struggle is useful only when it is directed toward a meaningful problem and remains within a range where the learner can still generate, test, evaluate, and refine solutions. Random confusion, overwhelming difficulty, or the repeated reinforcement of ineffective behaviors is not productive simply because it feels hard. However, learning science gives us good reason to be suspicious of the assumption that the smoothest immediate performance necessarily reflects the best learning. One relevant idea is what Robert and Elizabeth Bjork have described as desirable difficulties—conditions that can make practice more effortful or temporarily depress performance while improving long-term retention, discrimination, and transfer. Spacing, interleaving, retrieval, problem variation, and other forms of carefully calibrated difficulty can make a learner appear less successful during training even while producing more durable learning. Our teaching manual makes this distinction explicit: learning that feels easy can be shallow, while appropriately structured difficulties can improve adaptability and long-term retention.
A related concept is productive failure. In this instructional approach, learners first attempt a complex problem before receiving explicit instruction that makes a clean solution more likely, after which their attempts are consolidated through structured guidance. The immediate attempt may be unsuccessful, but the effort can expose gaps in understanding, activate prior knowledge, generate useful questions, and create a richer context for the instruction that follows. In our curriculum, the Form Box exercise uses this dynamic by asking students to confront a complex observational problem before they complete the isolated form studies that will later give them better tools to solve it.

I often use a simple metaphor to explain the idea. Imagine two people standing at the entrance to a room filled with furniture, but the room is completely dark. Both are asked to cross the room and reach the opposite side. The first person has a terrible time of it. They bump into a table, find a chair with their shin, discover a shelf the hard way, adjust course, encounter another obstacle, and eventually stumble their way to the exit. The second person somehow makes it across almost effortlessly, perhaps because their particular path happened to contain few obstacles; they simply glide through the room and reach the other side.
If our only measure is successful completion of the task, the second person clearly performed better. But who learned more about the room? The first person now carries a surprisingly detailed map of the environment. Each collision provided more information, each incorrect assumption required an adjustment, and each failed route narrowed future possibilities. The second person certainly learned something as well: from that particular starting point, a relatively direct path could carry them safely to the exit. But they encountered fewer obstacles, explored fewer alternatives, and therefore acquired less information about the room as a whole. This is why the visible disorder of a difficult training exercise does not necessarily trouble me. When I see evidence that a learner has attempted something, recognized a disparity, tried another approach, reassessed, corrected, and tried again, I am often looking at the physical record of a substantial feedback loop. The paper may not be pristine, but the learner has spent a great deal of time doing what deliberate practice requires: identifying problems and modifying behavior in response. And this is where Goodhart’s Law begins to implicate the teacher just as much as the student. If we make immediate success our metric for effective teaching, we may be tempted to remove the very difficulties that drive durable learning.
Yes, Teachers Can “Goodhart” the Curriculum Too
Up to this point, it might be tempting to frame Goodhart’s Law primarily as a problem of learner behavior: students discover increasingly efficient ways to satisfy the visible demands of an exercise while inadvertently drifting away from the learning those demands were intended to promote. That would be far too easy on those of us who are instructors—but, unfortunately, teachers can Goodhart a curriculum too. In fact, instructors may be especially vulnerable because we constantly judge whether our teaching is effective. We look at student performance, completion rates, progression through a program, the quality of finished exercises, the amount of correction required, and any number of other observable outcomes. These things can all provide valuable information. The problem begins, once again, when one or more of them quietly becomes the definition of instructional success.
Imagine, for example, that we begin judging teaching quality primarily by how quickly students progress through a curriculum. An instructor who wants students to succeed may naturally intervene earlier, simplify difficult tasks, increase the use of strengtheners, or advance learners as soon as the visible requirements appear to have been satisfied. Progression speeds up.
Or perhaps the valued outcome is the appearance of the student work. The cleaner and more accurate the exercises, the more successful the instruction appears. Again, perfectly reasonable. So the instructor demonstrates more, corrects more frequently, supplies more measurements, catches errors earlier, and prevents students from wandering too far into approaches that may produce an unattractive result. The work appears to improve.
Or perhaps we value consistency across a program. Students should meet a recognizable “standard,” so instructors increasingly steer them toward the procedures most likely to produce that standard quickly and reliably. The classroom begins to look very successful—and it may be. But each of these situations contains the same question that has followed us from Hanoi: what exactly have we optimized?
If the instructor takes on too much of the problem-solving, the learner may produce increasingly impressive work while developing less independence. If supports remain in place because they reliably produce cleaner work, the student may have fewer opportunities to discover what happens when those supports disappear. If difficult exercises are softened until nearly everyone succeeds immediately, we may have improved our completion rate by systematically removing experiences that would have produced greater adaptability later. The danger is particularly difficult to see because none of this requires a bad teacher. Quite the opposite, as any of this can emerge from instructors who are attentive, generous, hardworking, and deeply invested in their students’ success. And that is precisely why Goodhart’s Law is useful. It reminds us that good intentions do not protect a measure from becoming a bad target. Within our own curriculum, the teacher’s role is therefore defined less as someone who produces successful student exercises and more as someone who manages the conditions under which useful development can occur. Our teaching manual describes the instructor as a diagnostician, identifying the learner’s present needs; an architect, calibrating challenges and supports; and a coach, providing useful feedback while gradually fading assistance as independence develops.
That final part matters. A strengthener that reliably improves performance can tempt us to leave it in place, but its purpose is temporary: to modify the challenge while the learner develops a capability that eventually makes some portion of the support unnecessary. Our manual therefore emphasizes reducing scaffolds as performance stabilizes. The intended destination is not permanent success under ideal instructional conditions, but increasingly adaptable performance that survives changes in context, support, and difficulty.
Sometimes immediate correction is appropriate. At other times, a demonstration can save hours of unproductive confusion. Sometimes adding structure is exactly what keeps a learner within an appropriate challenge zone. But there are also moments when the most useful thing an instructor can do is wait.
- Allow the student to notice the disparity.
- Allow them to propose an explanation.
- Allow them to try an imperfect solution.
- Allow them to discover why it failed.
This has implications well beyond our particular curriculum. Any skill-based educator can fall into the same trap. A music teacher can make every performance sound better by correcting every phrase before the student learns to diagnose it. A coach can continually position an athlete so precisely that performance collapses when the cue disappears. A writing instructor can edit every weak sentence into clarity while leaving the writer little better equipped to recognize why the sentence was weak in the first place. The teacher improves the product while the learner does not necessarily improve by the same amount.
This is one reason that judging educational quality from visible student output alone can be so misleading. Beautiful work, rapid progression, high completion rates, or consistent results may reflect excellent teaching. But they may also reflect an environment in which instructors have become exceptionally good at helping learners satisfy the indicators by which the program itself is being judged. The distinction can be hard to see from the outside, and sometimes even from the inside. So the instructor has to keep returning to a more demanding question than “Are my students succeeding?” They must ask “What are my students becoming increasingly able to do without me?” That may be one of the most useful defenses against Goodhart’s Law that a teacher has.
So What Should We Measure Instead?
At this point, it would be easy to draw the wrong conclusion from Goodhart’s Law: perhaps the safest response is simply to stop measuring things. But I think you’d agree that wouldn’t help very much. In a skills-based curriculum, measurement is enormously useful. If we cannot identify what an exercise is intended to develop, observe changes in performance, compare those changes over time, or determine whether developing abilities survive increasingly difficult conditions, then we are left largely with intuition and impression. Our own curriculum places considerable emphasis on explicit learning objectives and observable learning outcomes for precisely this reason. The exercise, the intended capability, and the evidence used to assess development should be made as clear as possible.
The lesson of Goodhart’s Law is therefore not “don’t measure.” It is “don’t ask one convenient measure to tell you more than it can.” One useful response is what we might broadly call triangulation: looking for convergence among several different indicators rather than letting a single proxy stand in for the whole of learning. If several forms of evidence point in the same direction, we can be much more confident that the capability we care about is actually developing.
If we look back at our Shape Replication exercise through this lens, it would look like this:
Accuracy matters. If a student’s replications become progressively less faithful to the models, it would be rather difficult to argue that the exercise is proceeding successfully. The transparency overlay therefore gives us useful information. But accuracy is only one piece of the picture. We might also ask whether the learner is maintaining light, deliberate, and consistent line quality. We might look at how much measurement or other scaffolding is required to achieve their result. We might rotate or alter the model to see whether performance holds up under changes in orientation. We might introduce unfamiliar configurations rather than endlessly repeating the same shape configurations. We might delay feedback, or revisit the capability later to see whether it was retained. Most importantly, we might look for evidence that the developing perceptual-motor relationships reappear in later proportion, form, spatial, or representational challenges. In other words:
Accuracy + line quality + reduced scaffold dependence + variation + retention + transfer tells us considerably more than just accuracy alone. No single one of these needs to become the number.
This is also why the broader structure of a curriculum matters so much. If students simply repeated the same Shape Replication exercise until they could produce nearly flawless transparency matches, we might become extraordinarily good at measuring one increasingly specialized form of performance. Instead, earlier capabilities are repeatedly revisited under changing conditions, combined with other skills, challenged by greater complexity, and eventually carried into problems that look very different from the exercises in which those capabilities first developed. And that changing context is highly informative. The manual explicitly treats transfer as something that cannot simply be assumed. Early exercises are intended to function as stepping stones into broader capabilities, and learners are repeatedly asked to reapply earlier skills under new conditions. Shape Replication experience and development should contribute to later judgments of proportion and spatial relationships; value exercises should reappear in form studies; isolated competencies should eventually operate together within complex representational work.
The same logic underlies the deliberate combination of blocked and interleaved practice within the curriculum. Repetition can help stabilize an emerging skill, while strategic variation asks whether that skill remains available when the learner can no longer rely on precisely the same problem, sequence, or context.
Scaffold fading offers another source of evidence. If a learner initially requires six additional measurement points, then four, then two, and eventually none, that trajectory tells us something important that the final drawing alone cannot. The same is true when transparency checks become less frequent, instructor intervention decreases, or a learner begins independently selecting appropriate strategies rather than waiting to be told which one to use. In each case, we are asking not merely whether performance improved, but what the learner increasingly contributed to that performance.
This changes the role of assessment. A transparency overlay is not a scoreboard declaring that one student earned a 96% while another earned an 89%. It is a diagnostic signal. It tells us something about the relationship between an attempted replication and a target at a particular moment, under a particular set of conditions. Its real instructional value comes from what we do with that information next.
Did the error reveal a weak perceptual judgment?
A motor-control problem?
Overdependence on one strategy?
A failure to maintain previously developed line quality?
Does the student recognize the disparity without being shown?
Can they propose an effective correction?
Does the same problem recur when the model is rotated?
Does it disappear only when you add more measurement points?
The solution, then, is not to merely replace one target with several. A learner may still need clear, immediate performance targets within an exercise, but the instructor must evaluate those targets as part of a broader synthesis of information relative to the actual learning objective. Accuracy, line quality, scaffold dependence, adaptability, retention, and transfer may each provide useful evidence, but none should stand in for the capability the exercise was designed to develop. This clarification matters because almost any assessment system can become vulnerable to Goodhart’s Law. Even a sophisticated rubric containing seven different categories can simply produce seven new targets if everyone begins optimizing the rubric rather than the capabilities it was designed to reveal. Adding more measures does not solve the problem if we turn them into another collection of boxes to check. The defense is not measurement complexity for its own sake. It is maintaining a clear hierarchy:
Learning objective first. Evidence second. The objective tells us what we are trying to develop. The measures help us determine whether that development appears to be occurring. And whenever a measure begins exerting enough influence to change how learners or instructors approach the task, we should ask whether it still tells us what we think it does. That may be the most useful way to live with Goodhart’s Law rather than attempting to escape it.
Measure carefully, draw on multiple sources of evidence, change the conditions, fade the supports, and look for retention and transfer. Never allow the thing that is easiest to count to quietly become the thing that matters most. A good measure should help us understand learning; it should never be allowed to define it.
Remember: The Measure Is Not the Skill
Goodhart’s Law is useful here not because it offers some sweeping argument against measurement, but because it exposes a fairly ordinary mistake in how we interpret performance. The rat tail was not the rat problem. A test score is not education. And a flawlessly replicated shape is not perceptual-motor fluency. However, each of these things may be strongly related to the condition we actually care about. In fact, that relationship is precisely what makes the measure useful in the first place. We counted rat tails because we assumed they should tell us something about the number of rats being killed. We examine test scores because we believe they should tell us something about what students know or can do. We compare a replicated shape with its model because we believe that the resulting disparities can tell us something about the perceptual and motor relationships a learner is developing. The problem begins when that relationship is quietly reversed—when producing the indicator becomes more important than developing the capability the indicator was intended to reveal.
This is especially important in a curriculum like ours because we deliberately place great value on observable performance. The Waichulis Curriculum is skills-first and performance-focused, with exercises designed around identifiable perceptual, cognitive, and motor objectives rather than vague impressions of artistic development. Clear learning objectives and measurable outcomes matter precisely because they let us examine whether the experiences we provide produce the changes we intend. Goodhart’s Law does not weaken that commitment. It places an important condition on it: measurement must remain subordinate to the thing being measured.
That requires a certain humility from both learner and instructor. A measure that served us well yesterday may become less informative once behavior reorganizes around it; a scaffold that revealed progress at one stage may begin concealing an undesired dependency at another; and a familiar exercise may eventually become too specialized to tell us much about broader capability. The relevant test is whether improvement survives changed conditions, reduced support, unfamiliar problems, and ultimately transfer.
Many of the things we most want from education—adaptability, fluency, discrimination, judgment, independence, and transfer—are difficult to observe directly, so we still need proxies, exercises, rubrics, comparisons, overlays, standards, and other forms of evidence. But the easier a proxy is to count, reward, display, or optimize, the easier it becomes to confuse that proxy with the goal itself. Perhaps this is the most useful contribution Goodhart’s Law can make to art education: successful performance should remain evidence in an ongoing investigation, not the final verdict on learning.
Ultimately, the goal is not to produce better numbers, nor even to produce students who become exceptionally good at our exercises. It is to develop learners who can carry increasingly useful capabilities into situations we did not specifically prepare for, solve problems without constant intervention, and continue refining those abilities long after formal instruction has ended. In that sense, an exercise’s disappearance from the learner’s life should not mean the loss of anything important. If the training has done its job, what mattered most was never contained in the exercise to begin with. A good curriculum does not produce students who become just extraordinarily good at its exercises. It makes the exercises progressively unnecessary.