638 points km144 2 hours ago 558 comments
throwaway2027 2 hours ago | parent
handfuloflight 2 hours ago | parent
cronin101 2 hours ago | parent
staticman2 2 hours ago | parent
aoeusnth1 2 hours ago | parent
ThouYS 2 hours ago | parent
danw1979 2 hours ago | parent
hmokiguess 2 hours ago | parent
RGS1811 2 hours ago | parent
carlos-menezes 2 hours ago | parent
sailfast 2 hours ago | parent
lgessler 1 hour ago | parent
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
loopmonster 1 hour ago | parent
bibimsz 1 hour ago | parent
variety8675 2 hours ago | parent
emadabdulrahim 2 hours ago | parent
akhilome 2 hours ago | parent
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
m4tthumphrey 2 hours ago | parent
gruez 2 hours ago | parent
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
thejazzman 2 hours ago | parent
EricBurnett 2 hours ago | parent
KyleTheDev 2 hours ago | parent
I agree that it's sort of stupid, not a fan.
ealready_value 1 hour ago | parent
giancarlostoro 2 hours ago | parent
iAMkenough 2 hours ago | parent
Everyone that doesn't gets served some animated bullshit.
gruez 1 hour ago | parent
Yep, you're right. I tried on my phone and got the scroll through image.
mbreese 2 hours ago | parent
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
dionian 2 hours ago | parent
iAMkenough 2 hours ago | parent
thebitguru 2 hours ago | parent
halyconWays 2 hours ago | parent
josefresco 2 hours ago | parent
halyconWays 1 hour ago | parent
oefrha 1 hour ago | parent
swader999 2 hours ago | parent
amluto 2 hours ago | parent
bibimsz 1 hour ago | parent
serchinastico 1 hour ago | parent
dbbk 2 hours ago | parent
petesergeant 2 hours ago | parent
Catloafdev 2 hours ago | parent
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
gekoxyz 2 hours ago | parent
ithkuil 2 hours ago | parent
mavamaarten 2 hours ago | parent
brandon272 2 hours ago | parent
gwking 2 hours ago | parent
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
brandon272 2 hours ago | parent
fragmede 27 minutes ago | parent
drnick1 2 hours ago | parent
username_my1 2 hours ago | parent
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
ygouzerh 2 hours ago | parent
meric_ 2 hours ago | parent
https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
Eliezer 1 hour ago | parent
j_heffe 1 hour ago | parent
adastra22 4 minutes ago | parent
aray07 2 hours ago | parent
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
dgroshev 2 hours ago | parent
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
sashank_1509 2 hours ago | parent
dgroshev 2 hours ago | parent
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
wren6991 1 hour ago | parent
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
dgroshev 1 hour ago | parent
Claude is just comically bad nowadays.
redox99 1 hour ago | parent
Gander5739 2 hours ago | parent
tomhow 2 hours ago | parent
km144 2 hours ago | parent
tomhow 2 hours ago | parent
mupuff1234 2 hours ago | parent
Lord_Zero 2 hours ago | parent
setsewerd 2 hours ago | parent
WarmWash 2 hours ago | parent
petesergeant 2 hours ago | parent
mupuff1234 2 hours ago | parent
Less companies involved means less pressure to go fast.
nozzlegear 2 hours ago | parent
roughly 2 hours ago | parent
icrbow 2 hours ago | parent
Lord_Zero 2 hours ago | parent
Gander5739 2 hours ago | parent
alpineman 2 hours ago | parent
skunkworker 2 hours ago | parent
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
ekckekcjekfj 2 hours ago | parent
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
meerita 2 hours ago | parent
gopalv 2 hours ago | parent
nickandbro 2 hours ago | parent
keeganpoppen 2 hours ago | parent
cogythea 2 hours ago | parent
jdmoreira 2 hours ago | parent
buntp 2 hours ago | parent
wren6991 2 hours ago | parent
FergusArgyll 1 hour ago | parent
frshgts 1 hour ago | parent
glub 2 hours ago | parent
calibas 2 hours ago | parent
We can't test it properly because it knows it's being tested.
johntb86 2 hours ago | parent
pookieinc 2 hours ago | parent
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
jbellis 2 hours ago | parent
meric_ 2 hours ago | parent
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
randomblock1 2 hours ago | parent
viccis 2 hours ago | parent
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
scrollop 2 hours ago | parent
thibran 2 hours ago | parent
joshstrange 2 hours ago | parent
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
bayesianbot 2 hours ago | parent
bleonard 1 hour ago | parent
So longer threads get cheaper and one-shots stay the same price.
sznio 2 hours ago | parent
system2 2 hours ago | parent
lanyard-textile 2 hours ago | parent
ygouzerh 2 hours ago | parent
mavamaarten 2 hours ago | parent
Zambyte 1 hour ago | parent
adastra22 8 minutes ago | parent
booty 2 hours ago | parent
copperx 1 hour ago | parent
NorwegianDude 31 minutes ago | parent
The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.
enraged_camel 1 hour ago | parent
anthonypasq 1 hour ago | parent
enraged_camel 1 hour ago | parent
Game_Ender 49 minutes ago | parent
copperx 1 hour ago | parent
mgw 2 hours ago | parent
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
alvis 2 hours ago | parent
sharkjacobs 2 hours ago | parent
God I hope so
kantahayashi 2 hours ago | parent
mavamaarten 2 hours ago | parent
lgessler 2 hours ago | parent
boc 2 hours ago | parent
fastball 2 hours ago | parent
mikeocool 1 hour ago | parent
neilellis 1 hour ago | parent
drbscl 1 hour ago | parent
Trasmatta 1 hour ago | parent
unddoch 1 hour ago | parent
abtinf 2 hours ago | parent
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
felixgallo 2 hours ago | parent
ryanscio 2 hours ago | parent
felixgallo 2 hours ago | parent
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
esafak 1 hour ago | parent
abtinf 2 hours ago | parent
onlyrealcuzzo 2 hours ago | parent
This is news to me. Excited to try it out! Thanks.
nchmy 2 hours ago | parent
KeplerBoy 1 hour ago | parent
polalavik 2 hours ago | parent
roughly 2 hours ago | parent
Can you give more details here? This sounds intriguing.
sidrag22 1 hour ago | parent
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
cbg0 2 hours ago | parent
qlte 1 hour ago | parent
https://artificialanalysis.ai/models/releases/claude-opus-5-...
Opus 5.5 Medium = $1.34
GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
margorczynski 53 minutes ago | parent
notatoad 50 minutes ago | parent
mlcruz 1 hour ago | parent
aennassiri 2 hours ago | parent
tag2103 2 hours ago | parent
jacobgold 2 hours ago | parent
km144 2 hours ago | parent
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
CPLX 2 hours ago | parent
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
port3000 2 hours ago | parent
booty 2 hours ago | parent
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
Syntaf 2 hours ago | parent
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
cbg0 2 hours ago | parent
booty 2 hours ago | parent
"benchmark margins have become a less
reliable guide to real-world differences"
sounds like a big problem.
My guesses:1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
ayhanfuat 2 hours ago | parent
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
GodelNumbering 2 hours ago | parent
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
alvis 2 hours ago | parent
liudaisuda 2 hours ago | parent
weiran 2 hours ago | parent
re-thc 2 hours ago | parent
For long running tasks it is. That's what made Deepseek so cheap.
vardalab 1 hour ago | parent
Espressosaurus 2 hours ago | parent
hedgehog 1 hour ago | parent
bayesianbot 2 hours ago | parent
blfr 2 hours ago | parent
rapfaria 2 hours ago | parent
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
blfr 2 hours ago | parent
neuronexmachina 1 hour ago | parent
herpdyderp 1 hour ago | parent
ascorbic 1 hour ago | parent
coffeebeqn 2 hours ago | parent
btown 1 hour ago | parent
AJ007 2 hours ago | parent
mcintyre1994 1 hour ago | parent
drbscl 1 hour ago | parent
It does work out to be a similar cost per task though
jsnell 1 hour ago | parent
https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
naasking 1 hour ago | parent
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
make3 1 hour ago | parent
rahimnathwani 2 hours ago | parent
chrisweekly 1 hour ago | parent
cute_boi 1 hour ago | parent
chrisweekly 27 minutes ago | parent
In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.
Shekelphile 1 hour ago | parent
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
brookst 1 hour ago | parent
_the_inflator 1 hour ago | parent
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
notatoad 1 hour ago | parent
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
forgot-my-pw 57 minutes ago | parent
margorczynski 57 minutes ago | parent
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
johnecheck 47 minutes ago | parent
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
kingstnap 2 hours ago | parent
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
ApolloFortyNine 2 hours ago | parent
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
prettyblocks 2 hours ago | parent
Espressosaurus 2 hours ago | parent
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
raesene9 1 hour ago | parent
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
flyinglizard 46 minutes ago | parent
searine 2 hours ago | parent
unglaublich 2 hours ago | parent
nijave 1 hour ago | parent
blfr 2 hours ago | parent
cute_boi 1 hour ago | parent
Giving moral lecture is different than reality i guess.
arw0n 1 hour ago | parent
kqp 51 minutes ago | parent
bushido 2 hours ago | parent
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
ACCount39 1 hour ago | parent
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
KeplerBoy 1 hour ago | parent
sys32768 1 hour ago | parent
ChatGPT 6 Pro answered it without issue.
debesyla 31 minutes ago | parent
timacles 9 minutes ago | parent
peri-cl 1 hour ago | parent
https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...
Metacelsus 1 hour ago | parent
nonethewiser 1 hour ago | parent
I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.
SoftTalker 43 minutes ago | parent
hirako2000 2 hours ago | parent
Infomercial at its best.
No wonder we are hammered with ai announcements.
techjamie 2 hours ago | parent
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
ryangg 2 hours ago | parent
potwinkle 1 hour ago | parent
zatkin 1 hour ago | parent
peri-cl 1 hour ago | parent
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
stri8ted 1 hour ago | parent
manquer 41 minutes ago | parent
ACCount39 1 hour ago | parent
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
sailingparrot 2 hours ago | parent
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
dmazin 2 hours ago | parent
the_gipsy 2 hours ago | parent
bpodgursky 2 hours ago | parent
someothherguyy 2 hours ago | parent
doesn't sound like a razor at all
dgellow 2 hours ago | parent
usewik 1 hour ago | parent
dgellow 51 minutes ago | parent
the_gipsy 37 minutes ago | parent
rubslopes 50 minutes ago | parent
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
cab648bec139cc 2 hours ago | parent
sailingparrot 2 hours ago | parent
jr3592 2 hours ago | parent
sailingparrot 2 hours ago | parent
jr3592 1 hour ago | parent
lantry 1 hour ago | parent
skerit 1 hour ago | parent
sailingparrot 1 hour ago | parent
recursive 1 hour ago | parent
sidrag22 1 hour ago | parent
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
sailingparrot 1 hour ago | parent
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
sidrag22 1 hour ago | parent
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
quietbritishjim 11 minutes ago | parent
davrosthedalek 2 hours ago | parent
dmix 2 hours ago | parent
BatmansMom 2 hours ago | parent
sailingparrot 2 hours ago | parent
scottyah 2 hours ago | parent
CodingJeebus 2 hours ago | parent
jr3592 1 hour ago | parent
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
lukewarm707 2 hours ago | parent
that, they fully intend to 'pace'.
user3939382 2 hours ago | parent
azan_ 2 hours ago | parent
drnick1 1 hour ago | parent
lukewarm707 1 hour ago | parent
what anthropic have stolen they intend to keep for themselves.
kadushka 2 hours ago | parent
dr0idattack 2 hours ago | parent
mukmuk 2 hours ago | parent
DiggyJohnson 2 hours ago | parent
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
post-it 2 hours ago | parent
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
sigmar 1 hour ago | parent
nradov 1 hour ago | parent
TeMPOraL 38 minutes ago | parent
So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.
Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.
idiotsecant 21 minutes ago | parent
You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.
nradov 12 minutes ago | parent
antod 25 minutes ago | parent
eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.
smelendez 4 minutes ago | parent
johnisgood 1 hour ago | parent
Is this the meaning or do I have it wrong? I have not checked.
wren6991 1 hour ago | parent
johnisgood 1 hour ago | parent
TeMPOraL 33 minutes ago | parent
ck2 1 hour ago | parent
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
LanceH 1 hour ago | parent
rhet0rica 1 hour ago | parent
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
melasadra 1 hour ago | parent
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
derac 1 hour ago | parent
qlte 1 hour ago | parent
arw0n 1 hour ago | parent
hencq 1 hour ago | parent
fragmede 1 hour ago | parent
squidbeak 1 hour ago | parent
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
browningstreet 1 hour ago | parent
lxgr 1 hour ago | parent
logifail 59 minutes ago | parent
It's a strategy to achieve more, not less.
browningstreet 21 minutes ago | parent
A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.
neo_doom 1 hour ago | parent
tetha 1 hour ago | parent
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
Leynos 1 hour ago | parent
staindk 1 hour ago | parent
doctoboggan 35 minutes ago | parent
jgwil2 30 minutes ago | parent
vmnb 1 hour ago | parent
kadushka 1 hour ago | parent
vasco 1 hour ago | parent
fragmede 1 hour ago | parent
mpalczewski 1 hour ago | parent
glenstein 1 hour ago | parent
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
Ar-Curunir 1 hour ago | parent
DiggyJohnson 1 hour ago | parent
patcon 1 hour ago | parent
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
platinumrad 1 hour ago | parent
DiggyJohnson 1 hour ago | parent
lxgr 1 hour ago | parent
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
rythmshifter 1 hour ago | parent
sir, this is a hacker news thread
OJFord 1 hour ago | parent
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
kelnos 1 hour ago | parent
It's a weird phrase. Not sure why there are so many people who feel the need to defend it with such passion.
lukewarm707 1 hour ago | parent
"there is nothing outside the text" - Jacques Derrida
icedchai 56 minutes ago | parent
janalsncm 32 minutes ago | parent
BobbyJo 8 minutes ago | parent
There is almost always a large amount of time and effort invested behind the scenes in exactly how to message things like this. That being the case, there is almost always some insight to be had criticizing and analyzing what they settled on.
Dumblydorr 1 hour ago | parent
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
plaidfuji 1 hour ago | parent
But stating it plainly like this would make the contradiction too obvious.
Rebelgecko 1 hour ago | parent
lkbm 1 hour ago | parent
"Pace yourself" specifically means "slow down".
gradus_ad 1 hour ago | parent
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
topbanana 1 hour ago | parent
nonethewiser 1 hour ago | parent
grohan 1 hour ago | parent
palmotea 1 hour ago | parent
Claude says it sounds fine. And Claude is now the judge of the English language style, not you.
api 1 hour ago | parent
jrochkind1 54 minutes ago | parent
It is clear what it means anyway, that's true, it means the left out words, more or less.
And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.
Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.
nradov 23 minutes ago | parent
https://www.war.gov/News/News-Stories/Article/Article/264106...
TheIronYuppie 11 minutes ago | parent
if you are in a long race, you don't run all out teh entire time. you pace yourself.
https://en.wikipedia.org/wiki/Pacing_strategies_in_track_and...
That couldn't be more exactly what they are doing here.
Iolaum 2 hours ago | parent
heyjstn 1 hour ago | parent
AtlasBarfed 1 hour ago | parent
Simply make them something that derives a text response from its training data.
dspillett 1 hour ago | parent
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
tencentshill 1 hour ago | parent
qgin 1 hour ago | parent
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
janpot 1 hour ago | parent
tantalor 1 hour ago | parent
bonesss 56 minutes ago | parent
Tade0 53 minutes ago | parent
jatora 44 minutes ago | parent
mullingitover 1 hour ago | parent
chinathrow 44 minutes ago | parent
Lendal 40 minutes ago | parent
hnha 34 minutes ago | parent
If their scare was honest, they would stop.
whalesalad 40 minutes ago | parent
varispeed 15 minutes ago | parent
Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.
Idiotic.
jdw64 2 hours ago | parent
richardjennings 2 hours ago | parent
glub 2 hours ago | parent
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
somewhatjustin 2 hours ago | parent
Nice. I was starting to think Haiku was going to be abandoned.
phendrenad2 2 hours ago | parent
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
snvzz 1 hour ago | parent
Yup. As unusable as Fable 5.1, for assembly on 80s 68k personal computer platform. Awful.
iamsyr 2 hours ago | parent
somewhatjustin 2 hours ago | parent
Nice. I was starting to think that Haiku got abandoned.
mchusma 2 hours ago | parent
Sol- 2 hours ago | parent
somewhatjustin 2 hours ago | parent
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
ricardobeat 1 hour ago | parent
skerit 59 minutes ago | parent
cesarvarela 2 hours ago | parent
mococa 2 hours ago | parent
simianwords 2 hours ago | parent
ricardobeat 2 hours ago | parent
jidaigeist 2 hours ago | parent
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
the_gipsy 2 hours ago | parent
b38484848 1 hour ago | parent
andriy_koval 1 hour ago | parent
Yabood 2 hours ago | parent
greenavocado 2 hours ago | parent
jatins 2 hours ago | parent
Thank you.
Arcuru 2 hours ago | parent
zuInnp 2 hours ago | parent
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
ieie3366 2 hours ago | parent
ACCount39 1 hour ago | parent
Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.
glub 2 hours ago | parent
Benchmarks often don't survive contact with reality.
boredtofears 1 hour ago | parent
drnick1 1 hour ago | parent
cheikhcheikh 1 hour ago | parent
drnick1 1 hour ago | parent
Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.
cowthulhu 1 hour ago | parent
arw0n 1 hour ago | parent
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
quotemstr 2 hours ago | parent
orangecat 1 hour ago | parent
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
notatoad 1 hour ago | parent
aragornii 2 hours ago | parent
Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.
m101 2 hours ago | parent
pavlov 2 hours ago | parent
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
notduckrabbit 2 hours ago | parent
manmal 2 hours ago | parent
bredren 2 hours ago | parent
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
Retro_Dev 2 hours ago | parent
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
wren6991 1 hour ago | parent
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
ACCount39 1 hour ago | parent
An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.
This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.
The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.
Retro_Dev 54 minutes ago | parent
wren6991 40 minutes ago | parent
Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?
ACCount39 12 minutes ago | parent
Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.
A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.
Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.
It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.
Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.
But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere.
villish 1 hour ago | parent
That's the moat. Mistral has the capability but not the legal protections.
aurareturn 2 hours ago | parent
I tried Opus 5 and Astra.
ramoz 2 hours ago | parent
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
bitexploder 1 hour ago | parent
Foobar8568 2 hours ago | parent
rumblefrog 2 hours ago | parent
LoganDark 2 hours ago | parent
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
anentropic 2 hours ago | parent
jdthedisciple 2 hours ago | parent
Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?
How would this alleged difference (most likely bs) actually show up in reality?
GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.
enraged_camel 1 hour ago | parent
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
garo-pro 2 hours ago | parent
aesthesia 23 minutes ago | parent
datadrivenangel 2 hours ago | parent
desmondl 1 hour ago | parent
mcintyre1994 1 hour ago | parent
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
Trasmatta 1 hour ago | parent
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
Aperocky 1 hour ago | parent
FireBeyond 25 minutes ago | parent
LtdJorge 1 hour ago | parent
nonethewiser 1 hour ago | parent
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
Trasmatta 1 hour ago | parent
penagwin 7 minutes ago | parent
That’s the step that causes the most significant gains in agentic performance.
But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).
That’s why it often gets worse on models that simply had more RL post training from the same base.
jaflo 1 hour ago | parent
epicepicurean 1 hour ago | parent
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
californical 40 minutes ago | parent
> Rewrites are declared by the publisher, never inferred from overlap
> NULL means dirty, and DELETE is the fence
croemer 28 minutes ago | parent
croemer 29 minutes ago | parent
itsafarqueue 1 hour ago | parent
mcintyre1994 29 minutes ago | parent
blurbleblurble 1 hour ago | parent
isodev 1 hour ago | parent
blfr 1 hour ago | parent
tomaskafka 1 hour ago | parent
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
woeirua 1 hour ago | parent
34679 1 hour ago | parent
Maybe this model can finally figure it out for them.
simonw 1 hour ago | parent
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.mdMikhailTal 1 hour ago | parent
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
simonw 1 hour ago | parent
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
FergusArgyll 1 hour ago | parent
zamadatix 1 hour ago | parent
MaxikCZ 1 hour ago | parent
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Brendinooo 1 hour ago | parent
segbrk 1 hour ago | parent
Kurtz79 1 hour ago | parent
cainxinth 1 hour ago | parent
ealready_value 1 hour ago | parent
inshard 1 hour ago | parent
skerit 1 hour ago | parent
spidersouris 43 minutes ago | parent
nijave 1 hour ago | parent
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
ceroxylon 20 minutes ago | parent
adverbly 56 minutes ago | parent
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
breezybottom 20 minutes ago | parent
karp773 1 hour ago | parent
Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.
theGeatZhopa 1 hour ago | parent
blurbleblurble 1 hour ago | parent
breezybottom 1 hour ago | parent
Not efficiency in writing, clearly.
jjcm 1 hour ago | parent
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
naet 1 hour ago | parent
jjcm 51 minutes ago | parent
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
copperx 1 hour ago | parent
jjcm 48 minutes ago | parent
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
copperx 19 minutes ago | parent
nailer 1 hour ago | parent
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
ramesh31 1 hour ago | parent
lousken 1 hour ago | parent
Retr0id 1 hour ago | parent
Yay, yet another model I can't use for anything interesting, even with CVP.
firemelt 1 hour ago | parent
2001zhaozhao 1 hour ago | parent
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
yipinwong 1 hour ago | parent
Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
copperx 1 hour ago | parent
yipinwong 1 hour ago | parent
---
I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.
I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).
---
As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.
Also adding verification for that Fable 5.1 output in the same sesssion.
bdangubic 1 hour ago | parent
garo-pro 1 hour ago | parent
__vivek 1 hour ago | parent
edude03 1 hour ago | parent
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
cmrdporcupine 1 hour ago | parent
https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...
dom96 1 hour ago | parent
It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.
doodlesdev 1 hour ago | parent
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
Big, if true.HarHarVeryFunny 54 minutes ago | parent
Ants: It's a good model, sir!
madjam002 51 minutes ago | parent
It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.
It's frustrating that there isn't more transparency here.
Madmallard 47 minutes ago | parent
chinese models can't come soon enough
we're already getting enshittification
Fizzadar 44 minutes ago | parent
thatxliner 32 minutes ago | parent
kibae 31 minutes ago | parent
This is where Chinese models are going to eat Anthropic's lunch.
AlfeG 9 minutes ago | parent
mosselman 26 minutes ago | parent
ieie3366 19 minutes ago | parent
Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with
slowin 6 minutes ago | parent
sebastiangrill 3 minutes ago | parent
toephu2 2 minutes ago | parent
Are the frontier labs even working on this problem?