Irina Slutskaya and Olympics 2002 | Page 12 | Golden Skate

Irina Slutskaya and Olympics 2002

The point of trimmed mean ...

Forgive my puzzlement, but I still can't tell whether you are arguing in favor of the trimmed mean or against it.

If we consider the case of nine judges, which method would you like the best?

(a) Use all nine scores.

(b) Drop the highest and lowest and average the middle 7?

(c) Drop the two highest and the two lowest and average the remaining 5?

(d) Use the median.

(At one time the ISU had 14 judges and they dropped the lowest three and the highest two. At another time they had 14 judges but eliminated some of the marks by random draw before averaging. They also replaced half the panel between the SP and the LP, presumably to achieve some degree of consistency between ther two segments while at the same time breaking up voting blocs.)
 
I think any statistical analysis of 6.0 scores is meaningless, because the judge typically would decide the placement and manipulate the score to achieve the desired ranking.

I felt exactly the opposite. There is a huge volume of fascinating statistical literature about how to decide the rightful winner from a list of ranked ballots. (In the last election the U.S. State of Maine experimented with a ranked ballot in some political elections, with inconclusive results.)

Anyway, the idea of choosing the desired placement, then scribbling down some meaningless 5.8s and 5.9s after the fact to make it come out that way, is exactly what we want to evaluate. In this theory the "rightful winner" is the skater who the majority of the judging panel feels (for whatever reason, good or bad) deserves to win. (Definimng "majority of the panel" is the bugaboo, of course, in the case of differences in the rankings of the individual judges.

The IJS, on the other hand -- add up a column of numbers to get a total score -- what's the fun of that? ;)
 
I was wondering why I'd never dealt with an inference type analysis on trimmed mean.

If you have studied the distribution of sample medians you may have encountered a version of this topic. The median is just the most extreme example of the trimmed mean. (You eliminate the bottom half of the scores and the top half and take the one middle score that is left.) This belongs to the class of "non-parametric" statistics -- that is, what can we infer from a sample in the complete absence of assumptions about the underlying population of data? (Answer: not much. ;) )

However, as applied to figure skating a quick glance at the protocols will quickly convince you that the median and the mean are almost always pretty close. So the problem does not really arise, and the discussion is something of a red herring that detracts from more significant issues involving human judges.

In particular, the case of the Chinese judge that Baron V raises has nothing to do with statistics and everything to do with national prejudice and questionable integrity.
 
I felt exactly the opposite. There is a huge volume of fascinating statistical literature about how to decide the rightful winner from a list of ranked ballots. (In the last election the U.S. State of Maine experimented with a ranked ballot in some political elections, with inconclusive results.)

Anyway, the idea of choosing the desired placement, then scribbling down some meaningless 5.8s and 5.9s after the fact to make it come out that way, is exactly what we want to evaluate. In this theory the "rightful winner" is the skater who the majority of the judging panel feels (for whatever reason, good or bad) deserves to win. (Definimng "majority of the panel" is the bugaboo, of course, in the case of differences in the rankings of the individual judges.

The IJS, on the other hand -- add up a column of numbers to get a total score -- what's the fun of that? ;)

Are you talking about the scores or the ordinals. I think the ordinals can provide interesting information; the scores are merely the vehicle to rank the skaters. Irina was upset with the 5.6 she received in SLC, but that score was merely used to place her behind Sasha and Michelle when he thought Irina was the best of the three technically but deserved to be ranked behind them.
 
The scores were supposed to give a general sense of how well someone skated, not just be used to rank the skaters. When I was only 10 years old and watching skating on TV, I came up with my own system for what jump layouts should be worth within the 6.0 framework, and that's what the judges were supposed to be doing too.
 
I'd have to agree 6.0 is more interesting statistically. I just keep forgetting what I'm supposed to do in it :drama:
 
Are you talking about the scores or the ordinals. I think the ordinals can provide interesting information; the scores are merely the vehicle to rank the skaters.

What was interesting to me personally was neither the scores nor the ordinals. It was the final step: given these 9 lists of ordinal rankings, who deserves the medal? All of the ways of deciding -- majority of ordinals, OBO, etc. -- have their advantages and disadvantages.
 
The scores were supposed to give a general sense of how well someone skated, not just be used to rank the skaters.

True, it was kind of a delicious hibrid in that respect.

Still, there were many contests where the judge had already given one skater 5.8, 5.8 and another 5.8, 5.9. Then the last skaters goes and the judge thinks that she did a tiny bit better than the first but not quite as well as the second. What scores will he give the last skater?
 
Still, there were many contests where the judge had already given one skater 5.8, 5.8 and another 5.8, 5.9. Then the last skaters goes and the judge thinks that she did a tiny bit better than the first but not quite as well as the second. What scores will he give the last skater?

5.7/5.9, 5.9/5.8, 6.0/5.7, and 5.6/6.0 all create the desired result. Those latter two would likely not be an accurate reflection of the technical/artistic qualities of that skater, however, so they wouldn't be used. :)
 
Of course. She did do a one foot step sequence in the SP - but how it applies to her LP, I still don't know.

As for Kwan's one foot spiral sequence, she really was the best in the field in terms of skating skills.

Musicality, carriage and artistic impression sure. But I wouldn’t say she had incredible skating skills. Her edge change spiral is iconic because of the way she *performs* it which is magical - but the actual blade technique involved in an inside to outside edge change isn’t difficult compared to a spiral sequence with a turn or something like Sasha’s skid spiral. It shows good edge control but the edge itself isn’t deep nor is there any risk involved. There’s a reason change of edge isn't considered part of footwork complexity. Slutskaya’s one foot footwork however is a good demonstration of skating skills because it showed her ability to maintain speed while doing turns - and turns kill momentum. Of course it’s flaw is the same thing that makes it remarkable- that it’s only done on one foot so it isn’t the best indication of complete skating skills without showing anything on the other foot.

Michelle’s 2002 SP wasn’t her best display of her skating skills or choreo at their best anyways. Look at the half-done spread eagles in her footwork or the barely held y-spiral — there so many unfinished movements and she’s capable of better.
 
The scores were supposed to give a general sense of how well someone skated, not just be used to rank the skaters. When I was only 10 years old and watching skating on TV, I came up with my own system for what jump layouts should be worth within the 6.0 framework, and that's what the judges were supposed to be doing too.

Awww, did you keep it? I would love to see what 10 y/o BoP thought! :biggrin:
 
Musicality, carriage and artistic impression sure. But I wouldn’t say she had incredible skating skills. Her edge change spiral is iconic because of the way she *performs* it which is magical - but the actual blade technique involved in an inside to outside edge change isn’t difficult compared to a spiral sequence with a turn or something like Sasha’s skid spiral. It shows good edge control but the edge itself isn’t deep nor is there any risk involved. There’s a reason change of edge isn't considered part of footwork complexity. Slutskaya’s one foot footwork however is a good demonstration of skating skills because it showed her ability to maintain speed while doing turns - and turns kill momentum. Of course it’s flaw is the same thing that makes it remarkable- that it’s only done on one foot so it isn’t the best indication of complete skating skills without showing anything on the other foot.

Michelle’s 2002 SP wasn’t her best display of her skating skills or choreo at their best anyways. Look at the half-done spread eagles in her footwork or the barely held y-spiral — there so many unfinished movements and she’s capable of better.

:confused: Whom would you call the best in the field when it comes to skating skills, then? Kwan competed between 1994-2006 - we can say the best in 1994 was Yuka Sato, and best in 1995 was probably Lu Chen.
 
Why do you keep saying there is no risk involved in doing a change edge spiral? You can easily lose your balance doing that; many people in fact did have that happen to them when it became essentially mandatory in CoP, including all of Slutskaya, Yu-Na, Cohen, and Ando off the top of my head! It's easier to do a 3-turn in basic upright position and follow it up with just a brief average quality spiral. It's also not accurate to say Kwan's Y-positions were "unfinished". She did finish them off with clear movements, she just didn't hold the positions out, and that's fine. Not everything needs to be be held, movement is allowed to be fluid, and the ability/choreography she showed there is better than people who hold a weak Y position for 2 seconds (which is what so many people started doing on their spirals when CoP rolled around).

It's very galling that you try to say her change edge spiral wasn't on deep edges. Those were certainly deep edges, better than pretty much anyone else ever on a spiral with that feature and that amount of extension and speed. If she is the only person who was capable of utilizing the edges like that, then how is that not great skating skill? In any case, she did not need to show her absolute best in order to deservedly beat what other people put out there. Competitions are all relative and her base quality was still better than the others. I will agree with you however that the later part of her footwork sequence in the SP was underwhelming; she could have finished off one of the steps better and she completely left out a 3-turn + bracket turn that had originally been there. Her LP footwork sequence was also underwhelming, but then almost everyone's were. I think Butyrskaya actually had the best footwork sequence in the LP, followed by Hughes, and then Suguri just for the speed.

Awww, did you keep it? I would love to see what 10 y/o BoP thought! :biggrin:

It didn't change much, tbh. Mainly some aspects about assessing the quality, after I was learning the jumps for myself. I loved numbers as a kid though; I was put in special math sessions during elementary school and graded the regular math quizzes for the teacher, or went around the room answering questions during class, while the teacher was busy working 1-on-1 with kids to teach them their multiplication and division. Maybe I'm slightly on the autism scale. :think:
 
In particular, the case of the Chinese judge that Baron V raises has nothing to do with statistics and everything to do with national prejudice and questionable integrity.

It was absolutely connected with statistics, i don't understand how is not. If only average scores were implemented, without taking out maximum and minimum, Chinese team would win the Olympics. That wouldn't be the most fair result, because all the other judges gave a little bit higher overall scores to German team compare to Chinese. That was an example why trimmed mean is better to use than the simple average score. The question of national bias was resolved only after it, by taking into account results from other competitions of the same judge and how her overall judging of different Chinese teams deviated from the scores of the other judges. So, method (b) from your examples is best to be used - (b) Drop the highest and lowest and average the middle 7, because that is the method that in the best way equally involved opinion of a higher number of the judges in the final scores, which is the point of a judging in the context we are talking about. If method (a) is used then opinion of some judges may affect in an unfair way the opinions of others, like in the example above. If method (c) is used then not enough judges opinions would be involved in the final results. Ordinal judging in figure skating practice was 'behaving' as method (a), median as method (c). Now, using higher number of judges would be even better solution, but method (b) is good enough statistical measure to deal with the rankings in a fine way after those opinions are given. I'm not sure where I made you puzzled when my point of saying that BOP 'judgement' is biased was by using the philosophy of trimmed mean from current figure skating envirovement. I mean, why would I use something as an argument if i'm not for that something!
 
5.7/5.9, 5.9/5.8, 6.0/5.7, and 5.6/6.0 all create the desired result. Those latter two would likely not be an accurate reflection of the technical/artistic qualities of that skater, however, so they wouldn't be used. :)
While true that the numbers ought to have some explanatory utility for the viewer and the skater... what exactly would it mean to be an "accurate" reflection of the technical/artistic qualities of the skaters? All the viewer knows from the 5.6/6.0 is that they were the best in artistry in that competition, and that there were possibly 4 who were better on the technical side. What meaning do I assign these numbers, and how did we decide it?

It's also why I find CoP meaningless (though a bit more utility with the way it calls difficulty vs quality, so we have some sort of breakdown). What does a "9.75" mean? What are we comparing that against? In this competition? Across two judges? Across two competitions?

It's the ranking and the differential between competitors, respectively, that have meaning, IMO. The fixation behind the meaning of these numbers is creating a lot of, well, mania, I think.

Maybe it's just a bit too difficult to understand for me :laugh:
 
It's also why I find CoP meaningless (though a bit more utility with the way it calls difficulty vs quality, so we have some sort of breakdown). What does a "9.75" mean? What are we comparing that against? In this competition? Across two judges? Across two competitions?

It's the ranking and the differential between competitors, respectively, that have meaning, IMO.

Maybe it's just a bit too difficult to understand for me :laugh:

The biggest problems with ordinals were:
1) susceptibility(?) to manipulate results in an unfair way (In 6.0 system, random individuals were the one ranking the skaters. In COP, those individuals are giving the assessment of figure skating elements and skaters abilities, and the ISU system is the one who is ranking the skaters)
2) lack of meaning of the given numbers, which include a feedback to the skaters. So when 8.75 is given, that number is telling the skater how good is their skating, and what they need to do to be a better skaters. For example, in a marking guide of components for Ice Dance, you can see what is the 'meaning' of different category of numbers (0.25-0.75, 1.00-1.75, 2-2.75 going to 9.00-9.75 and 10) Here it is, on page 20 https://www.isu.org/figure-skating/...ok-for-referees-and-judges-2019-20-final/file So, in COP system, ideally you are comparing skaters to 0-10 scale. The logic is similar with judging the students in the school/college. Rankings are not giving some special information to the skaters, except to know what is their placement (if they were the best in some limited field of some competition).
 
I don't think we had anyone winning worlds with a clean 3+3 between 2002-2006 (maybe we should let Irina have it for attempting such a difficult one at 2005 though). We'd have to wait till Miki Ando (then we had a barrage of them until Asada doing "only" a 3A+2T lol, and then Ando winning with a 3Lz+2Lo the next season, and then Kostner's 3T+3T lol). Before that, Kwan actually did do a clean 3+3 in the LP in 2001 IIRC.

The second 3T in Kwan's combo at 2001 Worlds was underrotated so it was not clean. She would've got < under CoP. I wrote about it on the forum years ago and I got strongly condemned but that's patent truth, just look at the slow-motion: https://youtu.be/lsQbrwqNE0A
 
Kwan's 3+3 at 2001 Worlds was fine. Not perfect, but fine. Unlike so many skaters today she doesn't drastically swing around on her toepick during the takeoff; she vaults straight back and up into the air. The jump is airier and neater than what we tend to see these days. Her second lutz in the program was the one that had more than a 1/4 missing, despite being smooth. Her first lutz and the flip in that program were tight on the rotation too. Thankfully there is much more to skating than just nitpicking jump rotation, which is similarly why I think she deserved to beat Slutskaya in that 2002 SP despite the wonky flip. There are actually issues with Slutskaya's own flip in comparison - she has what could be considered too long of a break in the required transition into the jump and she also totally frontloads her jumps in the SP.
 
The second 3T in Kwan's combo at 2001 Worlds was underrotated so it was not clean. She would've got < under CoP. I wrote about it on the forum years ago and I got strongly condemned but that's patent truth, just look at the slow-motion: https://youtu.be/lsQbrwqNE0A

That didn't really matter in 2001 because 3-3s were relatively uncommon until the Yuna/Mao rivalry blossomed. Sasha Cohen had an impressive medal haul despite landing just one her entire career. Irina was credited with the first 3Lz-3Lo at the 2000 GPF that was wayyyyy short, and she still got technical marks reflecting successful completion of the element. I think the judges were more impressed by the effort and risk more than anything else. Once most top skaters were doing them, it became more important to distinguish good ones from bad ones.
 
Back
Top