All Hokie, All the Time. Period. Presented by First Bank & Trust Company

H
HokieEngineer22 OP
May 11, 2026 at 11:09 AM ET
Professional statistician and occasional poster here.
At risk of throwing fuel on the fire, I just can’t see a strong, data-driven argument for VT not being one of the top 16 teams. RPI is a joke and not at all a legitimate metric. Even so, the committee seemingly ignored it in their seeding in many of teams anyway (Arkansas was #1 in RPI and got the #4 seed). Massey has the best well-known public ratings and has VT at #14. I also developed my own rating system earlier this year where VT finished at #12, ironically right above LSU. Not here to whine, but in my view claiming VT was not among the best 16 teams is not supported by the facts. At risk of being conspiratorial, I can’t help but suspect things might have turned out differently if we had SEC patches on the jersey in both the seeding and the hotel situations.

17 Replies

N
nkhooch
2mo
As a metric person, I would probably just say the RPI has shortcomings
and as opposed to calling it a joke (not a statistical term :-)). Massey is a good one and I am sure yours is too. As a statistician, you know that there is no perfect metric. Each metric has to decide what to include and how much to weight it. There is some subjectivity there. No way around it. You also know that there is a margin of error around any result. What's lousy is that there is a precise cutoff at 16. And our team keeps hovering around that. From what I have seen here, it may not have been the metrics/formulas that got us. People listed categories - wins against top 50, top 10, etc. as other things that were considered. The weighting of those is also subjective. The strength of the SEC conference clearly helps there and hard for us to compete with that in a slate of ACC opponents. That's likely the place where the SEC gets an advantage over every other conference. But, it's not that it's all unearned. They had the best OOC W/L pct of any conference. Pete said that he believed our schedule incorporated a number of strong OOC opponents. And that some of them did not turn out to be as strong as we hoped. But, please keep posting. There's always room for more data-driven commentary.
H
HokieEngineer22 OP
2mo
The statistical term re: the RPI would be “extremely underparameterized”
It is absolutely not grounded in any real statistics, it just averages a few things together that sound like they might be useful (record, opponents record, opponents opponents record). If you tried to present that as a legitimate model you would be laughed out of the room at every job I’ve ever had. Even a simple linear regression based would be more defensible. On your point about the strong conference schedule being what helps the SEC in the RPI rankings, you are exactly right and that’s why a more sophisticated model that can go beyond simple record comparison is needed. Any regression based analysis could better control for unequal scheduling. The NET in basketball is a good example of the NCAA realizing this and switching to a more robust, albeit still imperfect, metric. There is no excuse for this kind of antiquated “metric” having so much control over college softball/baseball. A volunteer graduate student could cook up something better in 30 minutes.
N
nkhooch
2mo
Come on now. The win/loss record "sounds" useful?
In what universe is the win/loss record of a team not the most fundamental statistic? And where averages are not considered a statistic? Does your metric exclude those 3 RPI elements? I think the word "joke" might apply there. They may not be sufficient, but they are significant. You start with W/L record. In a closed system (where everyone plays everyone) that is sufficient and how post season play is determined. In an open system, where there are way more teams that can be played you have to figure out a way to give the teams a relative rank. You quickly see W/L record is not sufficient. You have to include the strength of the teams that the W/L record was achieved under. But, then you realize that you need the secondary impact of strength of schedule or someone playing a team like Grand Canyon will boost your strength of schedule without considering who Grand Canyon has played. Do you recall UNC starting off the year 17 - 0? What did everybody know? That it was against a weak schedule. And sure enough, they finished the season 15-19 when they started to play better competition. If there is a subjective component in the RPI it would be the weighting of the 3 components. You could stress test that by changing the weightings and seeing how the teams shift. But, for sure, all 3 components of the RPI are important. After that, you can add in scoring margins. You can see the importance of that. And it goes from there. Can you share one of the linear regression models? Just dependent and independent variables, not coefficients. Yes, a grad student could come up with models. But, you have to be able to communicate them. Coaches need to be able to have a clear understanding of what drives the metric so they can consider that in scheduling. And fans and analysts need to be able to discuss them. One positive of the RPI is that it is easy to understand. Yes, the NET is a more robust model. But, you know, that metric is NOT the sole determinant of NCAA tournament seedings and selections. It's an input.
H
HokieEngineer22 OP
2mo
That’s not quite how building robust rating systems work.
I’m somewhat reluctant to even engage further, but I can’t really help myself. You do not build robust statistical models by cherry picking a few things like record and subjectively weigh them together as has been done with the RPI. You pick a loss/evaluation function to evaluate performance, then iterate on your chosen model to try to improve performance by seeing how it performs on unseen data. In essence, you want the model to learn what should be weighted from the data rather than just randomly guessing. In this case, the mean squared error (MSE) of a particular rating system’s predicted scoring differential between two teams would be the obvious evaluation function. A simple linear regression model with an indicator variable signifying if the game was played at home/away/neutral, an indicator variable for the opponent, and an indicator variable for the team themselves for each game could be the predictor variables that could be selected. The target variable would be the scoring differential between the two teams. This allows the model to learn weights to put each team’s performance in each game in context of the opponent they played. Simple linear regression utilizes a closed form solution to minimize MSE, exactly our goal here. The coefficients for each team would then give you an estimate of team strength properly controlled for location and schedule. You would then evaluate how well your model does on unseen data by holding out the 20% or so of a given season’s games and evaluating the MSE of your models’ predictions. This is the sort of methodology things like the NET, KenPom, Massey, etc. take when creating and evaluating predictive models. It sounds simple, but this approach will give a far more accurate representation of team strength when it comes to predicting which team is better than what the RPI does. ChatGPT could probably do a good job explaining this further if you were interested in learning more. Of course you are correct that selection is not being done based purely on metrics. Instead, we are allowing people to make arbitrary decisions that give them the flexility to advance their own agendas. Just take a look at the college football playoff committee. I would claim that VT (and essentially all non-SEC teams) would be better off in a system where decisions were made based of a fair, objective evaluation criterion like a rating system with publicly audible code than some committee.
H
Hokie_Forester
2mo
Best post I have seen on the subject. There is absolutely no logical
argument for VT not being a top 16 seed.
T
TomA
2mo
Well, how wast Angela Tincher not on the USA Olympic team, yet…
she no-hits them… Clearly a major SEC and West Coast bias in the sport.
MidloHokie MidloHokie
2mo
Can you name the three pitchers that were selected?
Osterman, Finch, Abbott. Most would include those three on a list of the best five ever (I would also include AT). At that point I would have taken AT over Finch who had graduated in '02 but Mike Candrea was the coach so . . . . It was 2008. At that point the debate was really just getting started about west coast versus everywhere else superiority. Sure the bias was there but the game had evolved in the west. The first East Coast team win the ASA national championship did not occur until 2005 when the NOVA based Shamrocks won it all, ironically, the year after AT played for them.
V
Va AandM
2mo
As Ronald Coase said “If You Torture the Data Long Enough, It Will Confess”
And it seems the NCAA is the modern Inquisition
B
BEST2VT
2mo
Yep, pick the teams you want and then select and weight the stats to
support it. No matter if those stats and weightings change every year because ‘we know more now.’ ** Edited by BEST2VT at 5/11/2026, 12:59:02 PM
V
Va AandM
2mo
Ah, Finagle’s Rule: draw curve then plot your points. Works every time
B
baumerj
2mo
It feels the same every time this happens. The committees
whether we're talking basketball, softball, whatever seem to adjust their story to justify the results they came up with. All we ask for is to come up with the rules, then use them. Everyone knows that the RPI is extremely flawed, so why are we still using it? I say come up with a BCS type system and take the committee out of it. In the BCS, polls were factored in anyway. At this point though, all there is to do get a big chip on your shoulder and go beat the snite out of everyone you play. This team has the tools to do it.
S
skautdog
2mo
Only comeback is to win. Go Hokies.
H
hokietony
2mo
You're "RPI is a joke" line, although probably true, was the death ....
of your analysis comparison. Committee apparently still loves it and harps on it as well as Top 25 wins and SOS. Until that love of RPI fades, the use of other data and DSR/KPI just won't happen.
H
HokieEngineer22 OP
2mo
Then they have contradicted themselves because they didn’t follow RPI…
…all that closely in the seeding. If so, Arkansas would have been the #1 overall team. I do agree that it’s unlikely that the reliance on RPI will change because it gives SEC teams an inherit advantage because of the Midmajor teams they are able to schedule and their strong conference schedule. So there is a strong incentive for them to support the current system.
H
HokieNL
2mo
Committee contradicts self. Story at 11! 😉
H
hokietony
2mo
In general they did, only 1 non-top 16 RPI team is a host.
H
HokieEngineer22 OP
2mo
That exception is sort of what I am trying to point out.
They seemingly feel free to ignore it when the committee feels like it doesn’t actually reflect the best 16 teams. In the case of Virginia Tech, any legitimate metric you look at will show them among the best 16 teams. If there was a rigorous and data-driven process for softball like basketball gets, I feel confident we would have been in the 12-16 range because that’s what the facts reflect.