63 points plurby 51 minutes ago 49 comments
mohamedkoubaa 25 minutes ago | parent
SkyeCA 7 minutes ago | parent
decodingchris 21 minutes ago | parent
amluto 21 minutes ago | parent
Mooty 19 minutes ago | parent
WarmWash 16 minutes ago | parent
The course looks like it is something that a human could do in 15 seconds, while Astra took 5 minutes.
pixl97 13 minutes ago | parent
pixl97 14 minutes ago | parent
A different way to think of this is, consciousness is just a near real time video game with causal influence.
zezcko 18 minutes ago | parent
Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game.
Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?
vablings 16 minutes ago | parent
pcstl 15 minutes ago | parent
micromacrofoot 14 minutes ago | parent
another trick is to have it build something in a sandbox and have it add a human-editable setting to point it to places outside of the sandbox
seems like they're somewhat more willing to build a metaphorical gun as long as they're not pulling the trigger
syntaxing 18 minutes ago | parent
valine 17 minutes ago | parent
It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.
robots0only 15 minutes ago | parent
jvanderbot 12 minutes ago | parent
There's also a very tangible limitation of the bitter lesson.
If, over time, compute climbs, and so compute-bound data-driven general architectures beat bespoke architectures (this is the bitter lesson), then it is not necessarily true that the most general architecture now beats all available bespoke architectures now (or even in the near/mid future - the crossover point is "eventually").
Bitter lesson is most tangible for long-running research directions. Sometimes you need something working as best as possible now.
bethekidyouwant 11 minutes ago | parent
publicmail 9 minutes ago | parent
valine 5 minutes ago | parent
VBprogrammer 8 minutes ago | parent
I wouldn't let him loose on the road though.
I think, at the very least, the guardrails would have to deterministic, ideally with super human senses, for people to accept self driving cars on the road.
moffkalast 8 minutes ago | parent
tintor 6 minutes ago | parent
It is easy to make car driving demos.
prometheus1992 15 minutes ago | parent
N_A_T_E 13 minutes ago | parent
jrflo 8 minutes ago | parent
blorenz 15 minutes ago | parent
comboy 14 minutes ago | parent
WarmWash 13 minutes ago | parent
onlyrealcuzzo 11 minutes ago | parent
But I imagine this is orders of magnitude more expensive / less efficient than whatever Waymo is already doing, right?
The cool thing is that 1) it's theoretically more generalizable, 2) if we wait 18 months, it'll be 100x cheaper, and another 100x cheaper likely in 18 more months - at that point - something like a Mac Studio inside a humanoid could have these generalized capabilities, and a lot of Robotics problems start to look more feasible - especially when you consider how much better the models could be if highly specialized.
famouswaffles 7 minutes ago | parent
famouswaffles 9 minutes ago | parent
SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516
ZeroBench - https://zerobench.github.io/
Robot Arms - https://openai.robocurve.org/gpt-6-astra/
dyauspitr 6 minutes ago | parent