68 points tomjakubowski 13 hours ago 27 comments
bloppe 13 hours ago | parent
TZubiri 13 hours ago | parent
Strands of evidence? My best guess would be that:
0- this is ai generated slop
1- it's using that watermarking technique
2- it's obviously detectable and degrades quality
3- it's amplified when inferencing on its own content and generates slop
akk0 12 hours ago | parent
a_t48 6 hours ago | parent
harimau777 4 hours ago | parent
tomjakubowski 7 hours ago | parent
EagnaIonat 3 hours ago | parent
As I understand the paper they are saying the reasoning/thinking you see is actually a translation of what is actually going on, and stuff can be lost in the translation. Similar to what was observed in j-space.
WithinReason 1 hour ago | parent
fellowniusmonk 13 hours ago | parent
ck2 12 hours ago | parent
then we'll have to "flip" other models to be snitches on the other agents
then they'll make double-agents
the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves
yeah this won't end well, at all
jplusequalt 12 hours ago | parent
They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.
cousinbryce 12 hours ago | parent
ck2 12 hours ago | parent
"Is he still in the grandmother's house?"
"We would like to speak to him."
(btw Google's "AI" explains the meaning of that moment/sentence perfectly as if it gets it, creepy)
pixl97 9 hours ago | parent
pixl97 9 hours ago | parent
And yes, agents are already being used in things like cyber warfare in which other AIs attempt to poison them while they are working.
We don't have the hardware for sovereign AI quite yet, but at the current rate of growth it's not that many years out.
EagnaIonat 3 hours ago | parent
That was the media hype about it. There were two incidents.
1. Using a RL to train a model, it found that it got rewarded for certain garbage phases, so continued to talk that way.
2. Certain Latin words for fish/birds were used instead of "fish" or "bird". Just a token issue.
joegibbs 7 hours ago | parent
bcorigliano 11 hours ago | parent
And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.
pixl97 9 hours ago | parent
mnkv 11 hours ago | parent
I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.
trhway 10 hours ago | parent
chubot 4 hours ago | parent
applicative 3 hours ago | parent
shawntan 3 hours ago | parent
But arguably, a larger model will not need the chain of thought a smaller model does, which means simply by scaling we're already reducing CoT.
If the people who were relying on CoT are panicking now, they should've been panicking when perceptrons became multi-layer perceptrons.