17 points Betelbuddy 1 hour ago 6 comments
asamoahf 1 hour ago | parent
So this fixes dependence between judges and leaves dependence between all the judges and the truth untouched, which is the failure people are actually worried about when they say eight models agreed. You still want a small human-labelled anchor set to break it. The number I'd find interesting is how much smaller that anchor set gets once you model the dependence, since that's the real saving.
Same shape as offline policy evaluation. Correlated logging errors survive any amount of re-weighting, and one real experiment would be probably what pins them.
Forgeties79 14 minutes ago | parent
Tsarp 59 minutes ago | parent
troupo 6 minutes ago | parent
It shouldn't even be a debatable question.
bryzaguy 3 minutes ago | parent