Over the past few days, I have received messages from school leaders and teachers across the country. Many have shared stories of unexpected marking decisions, disputed reviews and outcomes that simply do not seem to align with the evidence in front of them.
Individually, these stories could be dismissed as isolated incidents. Taken together, they raise a more important question:
Just how reliable is the SATs marking system?
At my school, our concerns began when we reviewed the scanned transcripts of two pupils’ KS2 GPS papers. Both pupils achieved a scaled score of 99. Both were one point below the expected standard. When we examined the transcripts made available to the school, we discovered that four pages, containing eight questions, were missing from each child’s script.
The review process did not resolve the issue. Those questions simply remained marked as zero. I then had to make a more formal complaint. No one thought to take the initiative and say – this is odd. At this point, the issue is no longer about whether an answer deserves a mark. It is about whether the answers were available to be marked at all.
As school leaders, we are repeatedly told that SATs outcomes are robust, reliable and subject to rigorous quality assurance. Yet if a school can identify substantial sections of a script that appear not to have been considered, it is reasonable to ask what confidence we should have in the resulting score.
Unfortunately, this was not the only concern.
One pupil answered the following GPS question:
Explain how the change of modal verb affects the meaning of the second sentence.
- We will walk home from the station.
- We might walk home from the station.
The pupil wrote: “One is saying it will definitely happen. The other one is saying it could happen.”
Many teachers would recognise this as a clear explanation of the difference between certainty and possibility. The answer was initially awarded a mark. During the review process, despite this question not being challenged, the mark was removed. Perhaps there is a technical reason why the response no longer met the mark scheme. If so, it deserves explanation. Because from the perspective of a Year 6 pupil, a parent or a teacher, it appears to demonstrate exactly the understanding the question was designed to assess.
We found a similar issue in the reading paper. The question asked what Amelia Earhart changed about her plans after her accident in Hawaii.
One child answered: “She had to make unexpected major repairs to her plane.”
The response received no credit. Of course, the mark scheme may have been seeking a more specific answer regarding the route she ultimately chose (Landing). However, the child’s response demonstrates inference. They recognised that an accident occurred, understood that significant repairs were required and linked those repairs to a change in circumstances. In other words, they demonstrated understanding. Yet understanding alone was not enough.
For me, this is where the debate becomes larger than individual questions. Children are taught that reading requires inference. They are encouraged to explain meaning in their own words. They are praised for demonstrating understanding rather than simply copying text. Yet assessment often appears to reward only the answer that most closely mirrors the mark scheme writer’s wording.
That may create consistency. But does it always capture understanding?
What concerns me most is that these examples sit alongside concerns about missing script pages, review decisions and outcomes that ultimately contribute to national statistics, inspection discussions and judgements about school effectiveness.
When a child is on a 80 or 90 statndadised score , one disputed answer may not matter. When a child is on 99 or 109, it matters enormously. One mark often changes the outcome. One mark changes how success is reported. One mark changes the narrative around a child’s achievement.
I am not arguing for lower standards. I am not suggesting that every appeal should succeed. What I am asking is whether we are sufficiently curious when reasonable questions are raised about the accuracy and consistency of outcomes. Lets face it, the people marking do not care as much about this as school leaders – that’s for sure.
If pages can apparently disappear from scripts, if marks can be removed from answers many teachers would deem correct, and if pupils demonstrating clear understanding can receive no credit, then surely it is right for school leaders to ask difficult questions.
Because the real issue is not whether a child scored 99 or 100. The real issue is trust in a system that is overburdened with accountability.
When the Stakes Are Even Higher
The concern becomes even greater when the pupils affected belong to groups that schools are working hardest to support.
Two of the scripts we challenged this year related to pupils with identified SEND (They had slightly larger buffed scripts which we ordered from the STA). Had those reviews not taken place, those children would have remained below critical thresholds despite evidence that marks (whole pages) had gone missing during the scanning process. The impact was not insignificant. The outcome of those reviews contributed directly to our reported results and would have moved our SEND attainment from broadly in line with national figures to above national averages.
That matters because attainment data is used to evaluate the effectiveness of schools, the impact of interventions and, ultimately, whether disadvantaged and vulnerable groups are receiving the support they need. We know that the IDSR is something Ofsted seem to feel is the bible of school performance and they most certaily aren’t intrested in your internal data. If marking inconsistencies disproportionately affect pupils who already face additional barriers, then this is not simply a question of assessment accuracy. It becomes a question of equity.
Assessment systems will never be perfect. However, when a single mark can alter a child’s outcome, a school’s reported performance and the perceived success of provision for vulnerable learners, we should expect the highest possible levels of consistency, transparency and accountability.
The Scripts We Never Check
Perhaps the biggest concern is that these examples only come from scripts that were scrutinised.
As headteachers, we do not routinely review every paper. The process is time-consuming, complex and often falls during one of the busiest periods of the school year. In reality, most schools look closely at a small number of scripts, usually those sitting on key boundaries where a single mark could make a difference.
This year, for the first time in my career, that scrutiny did lead to success. Two pupils had their results upgraded to Greater Depth following reviews. That should be reassuring. Yet it also leaves me asking a far more uncomfortable question.
What about the schools whose outcomes broadly align with expectations, who quite reasonably conclude that the results are “about right” and therefore don’t invest hours examining individual papers? What about the answers that were marked inconsistently but sit unnoticed because nobody has the capacity to challenge them?
The examples highlighted in this blog only came to light because we looked (It took my very busy assistant head a day to do this). They were not obvious administrative errors. They were questions of administration, interpretation, judgement and marking accuracy that had a direct impact on children’s outcomes and on our published results. If significant issues can be found in the relatively small number of scripts that receive detailed scrutiny, it is entirely reasonable to ask how many similar cases remain undiscovered across the thousands of papers that are never revisited. We looked at 15 papers and sent back 13.
The question is not whether mistakes happen. Every assessment system will contain some level of human error. The question is whether we have enough confidence in the consistency of marking to accept that the results children receive on results day are truly the results they have earned. When one mark can be the difference between Expected or Greater Depth that is not a question we should dismiss lightly.
A national assessment system only works if pupils, parents and schools can have confidence that every child’s work has been considered, every mark has been awarded consistently and every result accurately reflects what that child knows and can do. When so many colleagues are sharing similar experiences, perhaps the question is no longer whether mistakes happen.
The question is whether we are willing to have an honest conversation about how often they happen, and what that means for children sitting just one mark away from success.
Leave a comment