When models disagree, measuring ensemble variance helps spot risky inputs that...
https://wiki-room.win/index.php/What_Metrics_Should_I_Track_Besides_Accuracy_for_Risky_AI%3F
When models disagree, measuring ensemble variance helps spot risky inputs that need extra attention. For example, cases in the top 1-2% of variance can be routed to human review to prevent errors