33 points theanonymousone 1 hour ago 22 comments
dgellow 44 minutes ago | parent
andriy_koval 28 minutes ago | parent
tetec1 21 minutes ago | parent
jcmontx 42 minutes ago | parent
traceroute66 31 minutes ago | parent
"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".
So basically pinky-promise benchmarking ?
I'm not sure I follow the value here ?
kadoban 22 minutes ago | parent
traceroute66 17 minutes ago | parent
But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?
demibabs 18 minutes ago | parent
bix6 29 minutes ago | parent
demibabs 17 minutes ago | parent
How does that work?
traceroute66 14 minutes ago | parent
My gut feeling is that any serious real-world company with a proprietary codebase worth looking at would not be handing out the crown jewels to a third party. License or not.
I don't doubt somebody licensed their codebase to them, I just have my doubts about who the "who" could be.
InsideOutSanta 1 minute ago | parent
At any rate, I'm not sure it matters whose codebase it is. I'd even say that a shitty codebase might make for a better test.
IshKebab 17 minutes ago | parent
There's only two or three sane options here - you can easily try them all and pick yourself.
rovr138 2 minutes ago | parent
bdlowery 9 minutes ago | parent
Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change.
thereitgoes456 5 minutes ago | parent
siddbudd 2 minutes ago | parent
tucnak 1 minute ago | parent
lmeyerov 8 minutes ago | parent
One lesson of running botsbench.com, in a slightly different domain, is to measure for model contamination every time.
visiondude 4 minutes ago | parent