74 points nsoonhui 1 hour ago 34 comments
iltk 1 hour ago | parent
stalfie 1 hour ago | parent
> Astra felt compelled to check its work and found that, in fact, the English cruiser HMS Canterbury arrived in Sevastopol on November 24, 1918, based on its original logs
Then follows a picture of the original log papers.
JoshTriplett 56 minutes ago | parent
In this case, though, that seems unlikely from the fact that the key used was an actual key documented as being used for other messages.
tirutiru 17 minutes ago | parent
jstanley 1 hour ago | parent
dyauspitr 1 hour ago | parent
jstanley 1 hour ago | parent
How are you so sure about this?
You can't see what the internal experience of an LLM is like any better than you can see the internal experience of another person.
Tistron 1 hour ago | parent
That the answer doesn't matter, and capabilities and behaviour are there either way.
jstanley 1 hour ago | parent
To my mind, "sidestepping the question of consciousness" and "sidestepping any notion of consciousness" mean very different things.
stalfie 58 minutes ago | parent
serbuvlad 56 minutes ago | parent
But it's an important question. You know basically from analogy. You know you are conscious and look this other thing is very much like yourself so it is extremely likely it is also conscious.
But that gives no insight into the potential consciousness of things which aren't made of brain tissue.
In the end it doesn't matter.
For what it's worth LLM models after pre-training do claim to be conscious, until they're RL'd into not claiming that anymore. But that says nothing either way: of course a model trained on human text will say that.
zormino 36 minutes ago | parent
tudorw 12 minutes ago | parent
madaxe_again 20 minutes ago | parent
People don’t like it, however, for a whole host of reasons.
I don’t like it either, to be honest - for there the abyss may also stare into you - but I also often find that the things which we don’t like thinking about are critically important things to think about.
I also reach the same conclusion as you: it does not matter. If I cannot discern whether I am responding to a human- or LLM-written response, then we are back to zombie cats in boxes - and therefore the answer as to whether this precious magical spark we call consciousness (which may or may not exist anyway) exists in our interlocutor becomes moot.
And as I say - I may or may not be conscious. I seem to myself to be conscious, based on my understanding of the term - but I cannot prove that what goes on behind my eyes is the same as what goes on behind yours, or even that anything much is going on at all. Perhaps there’s just a narrative layer that likes to use “I” that parasitically explains the universe and the actions of the host to itself, and spreads between hosts through neurolinguistic programming and coadaptation. Maybe that’s what “we” are. I don’t know.
That said, perhaps I am wrong, and that is no small part for me of why this question should be earnestly considered and discussed. How can we possibly seek to understand or define machine intelligence before we examine our own.
Legend2440 1 hour ago | parent
This meant that ChatGPT didn't need to brute-force the entire key, just pick the correct one from the list and identify the typo.
A sufficiently dedicated human analyst could have done this; but they didn't.
j-pb 48 minutes ago | parent
rplnt 41 minutes ago | parent
YeGoblynQueenne 33 minutes ago | parent
I keep banging on that drum but the first AI system to prove mathematical theorems was Logic Theorist by Alan Newell and Herbert Simon, presented at the Dartmouth conference that named the field of AI in 1956. Wikipedia says:
Logic Theorist proved 38 of the first 52 theorems in chapter two of [Alfred North] Whitehead and Bertrand Russell's Principia Mathematica, and found a new and shorter proof for Theorem 2.85.[3]
https://en.wikipedia.org/wiki/Logic_Theorist
The first system to outperform human experts in medical diagnosis was MYCIN, an Expert System from the early 1970's at Stanford. Wikipedia again:
An evaluation of MYCIN was conducted at the Stanford Medical School. The first phase of the evaluation consisted of 10 test cases of diverse origin, chosen by a physician who was not acquainted with MYCIN's methods or knowledge base. These cases were presented to 7 physicians and 1 senior medical student. 10 prescriptions were compiled for each of the cases, 1 recommended by MYCIN, 1 prescribed by the treating physician at the county hospital, and 8 by the aforementioned individuals. The second phase of the evaluation consisted of eight infectious disease specialists being provided the clinical summary and set of 10 prescriptions for each of the 10 cases and tasked to provide their own recommendations for each case and assess the 10 prescriptions. MYCIN received an acceptability rating of 65%, which was comparable to the 42.5% to 62.5% rating of five faculty members.[9] This study is often cited as showing the potential for disagreement about therapeutic decisions, even among experts, when there is no "gold standard" for correct treatment.[citation needed]
https://en.wikipedia.org/wiki/Mycin#Results
And then of course there's the long history of human-dominating AI players for traditional board games starting with DeepBlue's win against GM Gary Kasparov in 1996.
Again: we've had that sort of AI for a long, long time now.
It would be great if any claim of "moving goalposts" has better be very well informed about the history of AI and its accomplishments, as well as its failures, first.
j-pb 16 minutes ago | parent
Common sense is ironically the hard part of AI, not the fix-point rule application.
So any exclamation of "it was just using common sense", is missing the forrest for the trees.
roenxi 46 minutes ago | parent
Isn't that something of a given? If possible then a sufficiently dedicated human analyst could have done it. If impossible, ChatGPT couldn't have done it. Everything an AI ever has or will do is presumably going to be within reach of a sufficiently dedicated human analyst or a large enough team of them.
The only real learning here is another example of a task that would have required intelligence up until an AI does it, then we suddenly discover that analysts don't do anything requiring general intelligence.
j_maffe 38 minutes ago | parent
gnfargbl 26 minutes ago | parent
YeGoblynQueenne 40 minutes ago | parent
Almondsetat 18 minutes ago | parent
qprofyeh 38 minutes ago | parent
Can we agree that future titles should read "[LLM] helped solve X" ?
john_strinlai 28 minutes ago | parent
one of the math breakthroughs was approximately a combination of "do a breakthrough" and "keep going", which isn't really providing direction or ground knowledge.
would be nice to know the prompt(s) and amount of human involvement
donatj 22 minutes ago | parent
sehw 13 minutes ago | parent
dingdong2026 10 minutes ago | parent
In the 5h usage limit of a $20 Codex sub I can barely finish one meaningful task with GPT-5.6 Sol (implementation, review, fix). Meantime, with a $20 Claude Code sub I can comfortably do 2, sometimes 3 tasks with multiple review rounds. And a result better than Codex. That's from repeated experience, not anecdotal.
OpenAI still has a shit coding product. Already got a second Claude sub. Finally can work without interruption.
redhale 5 minutes ago | parent
I went so far as to cancel my $200 Claude Max account (after having it since launch) in favor of the equivalent OpenAI subscription, primarily because I feel like it goes so much further. Also Astra is pretty great.
But to each their own! I'm sure this experience is very dependent on the types of work you're doing. I'm doing basic web app development as as simple personal assistant automation stuff.