165 points whiteros_e 5 hours ago 129 comments
dada216 3 hours ago | parent
freakynit 2 hours ago | parent
gpugreg 1 hour ago | parent
embedding-shape 3 hours ago | parent
They must have hit really hard scaling limits if the prices were hiked so much so quickly.
broodbucket 3 hours ago | parent
lompad 3 hours ago | parent
And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.
chobbledotcom 3 hours ago | parent
pyrophane 3 hours ago | parent
bbor 3 hours ago | parent
For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.
Daviey 3 hours ago | parent
disiplus 3 hours ago | parent
world2vec 2 hours ago | parent
Can I ask where are you using all those tokens?
wartywhoa23 2 hours ago | parent
p2detar 1 hour ago | parent
tokai 2 hours ago | parent
world2vec 2 hours ago | parent
disiplus 2 hours ago | parent
world2vec 1 hour ago | parent
embedding-shape 46 minutes ago | parent
_0ffh 2 hours ago | parent
rubslopes 1 hour ago | parent
buckle8017 2 hours ago | parent
Daviey 1 hour ago | parent
embedding-shape 47 minutes ago | parent
Daviey 25 minutes ago | parent
I now exclusively use https://omp.sh/ as my harness:
I set it up so it never works in the main branch so subagents etc don't step on each others toes, and only merges back when complete: https://github.com/Daviey/mario/blob/main/.omp/hooks/pre/wor...
A good AGENTS.md is essential: https://github.com/Daviey/mario/blob/main/AGENTS.md
I then provide specifications for what I want, making sure it is unit tested.
asp_hornet 3 hours ago | parent
andy_ppp 3 hours ago | parent
asp_hornet 3 hours ago | parent
criley2 2 hours ago | parent
Also why Meta gets a +1, just charge less money on the training path.
orf 2 hours ago | parent
If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.
These are not equal.
andy_ppp 39 minutes ago | parent
Havoc 2 hours ago | parent
Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate.
>They must have hit really hard scaling limits if the prices were hiked so much so quickly.
Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.
probst 1 hour ago | parent
_aavaa_ 1 hour ago | parent
The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate).
Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow.
They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed.
They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).
Schlagbohrer 1 hour ago | parent
That is shocking. Is it per-token I wonder?
workbreak 53 minutes ago | parent
_aavaa_ 31 minutes ago | parent
I’m getting 97%.
bbor 3 hours ago | parent
jensb1 3 hours ago | parent
bbor 3 hours ago | parent
In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.
In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.
I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(
TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.
[1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.
Bluestein 2 hours ago | parent
dgellow 2 hours ago | parent
> use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS
From my European point of view the same risk/concerns apply when using US providers
jLaForest 2 hours ago | parent
podocarp 1 hour ago | parent
pjc50 1 hour ago | parent
Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".
If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.
tuesdaynight 1 hour ago | parent
lelanthran 47 minutes ago | parent
I mean, if they get to distill other's IP, why can't others distill their IP?
woadwarrior01 2 hours ago | parent
phoghed 2 hours ago | parent
butterNaN 2 hours ago | parent
pjc50 1 hour ago | parent
No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.
_aavaa_ 1 hour ago | parent
lelanthran 49 minutes ago | parent
HarHarVeryFunny 48 minutes ago | parent
I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?
Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?
rob74 3 hours ago | parent
Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...
fxwin 2 hours ago | parent
Also I feel like the obvious way to read the very first sentence is that GLM is a language model
> As we develop GLM, the model sometimes exhibits capabilities that surprise us
jbonatakis 2 hours ago | parent
drbscl 2 hours ago | parent
Come on now
Also, why would they introduce themselves on their own blog?
Mashimo 2 hours ago | parent
Where GLM-5.3-Flash is the newest "small / fast" model.
peri-cl 2 hours ago | parent
bogdan 2 hours ago | parent
rob74 1 hour ago | parent
tokai 1 hour ago | parent
bogdan 20 minutes ago | parent
Appreciate the clarification. For me it was the "F" in "WTF" that tipped me. Other than that, it's more than fair for you to not know what GLM is. Things are moving so fast that I would be surprised if anyone can keep track of it all. Cheers, have a grand day!
HarHarVeryFunny 1 hour ago | parent
Why would you be reading their corporate blog posts if you don't even know who they are?!
Argonautlabs 2 hours ago | parent
One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2.
Method and numbers: https://github.com/argonautlabsai/argodrive (built on antirez/ds4).
tipsytoad 2 hours ago | parent
Havoc 2 hours ago | parent
GLM has in the past been more technical rather than speculation about future development on RSI etc.
Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.
dude250711 2 hours ago | parent
Schlagbohrer 1 hour ago | parent
HarHarVeryFunny 1 hour ago | parent
Just like the rest of the world, including the US (Intel, Micron), SMIC are currently using ASML lithography equipment (DUV, not EUV), but Shanghai Aishengna are now moving into early production with their own DUV machines, with SMIC and CXMT as early customers.
There is also a state sponsored Chinese EUV development underway.
jonstewart 2 hours ago | parent
HarHarVeryFunny 1 hour ago | parent
All I can recall reading from OpenAI about what they have actually done in the name of "RSI" is using one of their models to help automate the training process.
zicohacks 1 hour ago | parent
HarHarVeryFunny 1 hour ago | parent
In addition to Huawei who make the Ascend series that Ziphu are using, there are also at least a half dozen or so other Chinese companies also making their own AI accelerators.
0xbadcafebee 58 minutes ago | parent
ipsod 27 minutes ago | parent
Seems to work for them.
freakynit 25 minutes ago | parent
Almost everyone knew that these sanctions would backfire within a few years. You can't really put sanctions that have noticeable negative effects on bigger economies. They only work for small to medium economies. I believe sanctions on any economy in top 10 would fail.
HarHarVeryFunny 18 minutes ago | parent
Now, the US is left out in the cold with little influence left, themselves now the ones with an anti-missile shortage.
menaerus 58 minutes ago | parent
> Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.
stogot 48 minutes ago | parent
chung8123 1 hour ago | parent
tokai 1 hour ago | parent
gpugreg 55 minutes ago | parent
> Why would I pick GLM over Claude?
To support the company that makes their model weights available for download, while Anthropic lobbies to restrict access.Bawoosette 52 minutes ago | parent
menaerus 28 minutes ago | parent
GLM: 80 USD (pro), 168 USD (max) -> with "limited-time event" discount this becomes 56 USD and 117.6 USD
I also don't understand why are they so much costlier, and I would also like to give it a try.
ipsod 18 minutes ago | parent
GLM's "Max" plan is (was?) equivalent to 3x Claude's 20x ($200) plan.
throwa356262 1 hour ago | parent
"We implemented a series of aggressive memory optimizations, including..."
This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing.0xbadcafebee 59 minutes ago | parent
But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you seen how bloated and slow Python is? Do global locks really sound like a strategy for fast dynamic computation?
wolttam 38 minutes ago | parent
esseph 28 minutes ago | parent
Yes, but it's calling C code.
kamranjon 8 minutes ago | parent
HarHarVeryFunny 8 minutes ago | parent
If you look at how many years the whole NVIDIA and CUDA ecosystem has been evolving, it's certainly impressive how they've just stood up and optimized this CUDA-free 100,000 node cluster in just a few months.
KronisLV 54 minutes ago | parent
konart 52 minutes ago | parent
And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.
9cb14c1ec0 31 minutes ago | parent
a012 25 minutes ago | parent
Because they don’t have to. Most of the time money would buy you newest and/or more hardwares so there’s low/minimal interest to optimize the code or approach.
vblanco 20 minutes ago | parent
esafak 12 minutes ago | parent
Signed, a customer.