The choices for qualis seem to be:
a. Have separate panels of judges for both groups (as is), with both groups skating on the same day.
b. Have one round mid-morning and another round in the evening (instead of the afternoon), and have one panel of judges score both sets of skaters after an aftenoon nap.
c. Split qualis into two days, both groups starting at the same time. Same set of judges.
Questions:
1. Would the same judges apply the same standards to both groups, if they were on separate days or at different times? I'm not so sure. Judges tend to start slowly and warm up to a pattern of scoring. If they came back for the second panel that same day, then they might start the second panel in a groove. The same judges could score differently, based on having already judged 20+ skaters earlier in the day. If this happened, it would cause a more predictable double advantage to the skaters who perform in the second group.
2. Is it necessarily a disadvantage to have two groups of judges? It depends. In general, it's considered an advantage to skate in the later group, unless the skate is late in the first group. If the morning group of judges is more generous with scores, then this could neutralize the disadvantage of starting early. If the later group of judges is more generous, then it is a double whammy to the early group.
3. Is it an advantage or disadvantage to have an extra rest day, if the groups were split into two mornings or two afternoons/evenings? I think that depends on the skater. Some skaters would get more nervous with a longer wait after qualis before the short program, and others would prefer the extra rest/practice day.
4. If they're going to make skaters perform on the same day with different sets of judges, and they were willing to have an evening group in Calgary, why do they insist on a morning group, instead of an afternoon group and an evening group, particularly when a great number of skaters hate skating in the morning group? At least more people can show up for the evening group, once school and work is out, and these are usually the most affordable tickets.
The mitigations for having one tough set of judges and one easier set is not basing the short program cut-off based on the top 30 scores, but instead by the top 15 scores in each group, and factoring the scores to 25%. I, frankly, think they should jettison the scores after the quali rounds, except for recording personal bests, and at most, use them to weight the SP groups along the lines they use for Ice Dancing. I don't think it's a bad thing to have 18-20 judges working on a major event.
Random selection is not the same as choosing 9 judges up front and having 3 treated as phantoms. Random selection is done before each phase of the competition, and the chances of not being selected for any panel are greatest if the quali round is the only round judged, but slim after the SP/FS panel is chosen. The chances of having the scores count in at least one phase are solid. So it's unlikely a judge is sitting there for nothing, unless s/he's neither selected in the quali round nor selected for the SP/FS panel.
Random selection allows a greater variety of judgements over the course of a competition, because while there is overlap, the chances are not great that the same 9 judges' scores will count for all phases. As Mathman has pointed out, this can be for better or worse, depending on who is selected and whether they've been colluding. While this may not be better than taking all 12 scores, the theory being the more opinions, the more accurate, it still is different than having a single panel of 9 score everything.
It is fairly easy to tell which judges' scores are not counted for each stage, particularly if one is a programmer. Whether it is easy to know which judge gave which set of scores is dependent on whether they tell other people truthfully which set is theirs. Certainly there's a lot of speculation about which judges were which. On the other hand, if a country has three-four judges and one-two contenders, there is still plenty of room for complicated dealing, and the scores we assume are from the judge from that skater's country (or that country's arch-rival) might very well be from a judge who had made a deal to hold up or down that skater. Again, as Mathman has pointed out, it is possible for a "bloc" to be diluted in the random selection by being tossed out in whole or part or to be fortified by random selection, by making up a large portion of the counted judges. (The bigger the bloc, the more likely it is to have its members selected.)
It is very true that the person with the most points wins. Just as the cross country skier who has the fastest time by .1 over 100k wins the race. S/he gets the medal, the trophy, the bonus, the diamond earings, the car, the ISU money, etc. But what the figure skating scores also show is whether the scores are meaningful in an abstract sense, i.e., whether they are statistically significant, or could any X skaters have reached Y place. This is mathematically possible to determine regardless of whether there is random selection or purely trimmed mean. Ordinals flattened the differences between skaters: there could be a .1 point difference or a 10 point difference, and it all looked the same: 1st and 2nd. Absolute scores show scale, and statistically significance across multiple scores tells the story of how likely that outcome is.