482 points krackers 17 hours ago 138 comments
wolttam 16 hours ago | parent
krm01 16 hours ago | parent
kibae 16 hours ago | parent
Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
jwpapi 15 hours ago | parent
I’m saying who has a million dollars for me, so I can make my own model?
Bolwin 7 hours ago | parent
I still opus 4.6 though not for code
nikcub 9 hours ago | parent
thehamkercat 16 hours ago | parent
medlazik 16 hours ago | parent
atemerev 8 hours ago | parent
stymaar 4 hours ago | parent
brookst 1 hour ago | parent
I don’t think Trump changed them, but Trump is absolutely a symptom of larger social collapse in the US, and that collapse has affected Altman and Amodei. We’re not even pretending that truth matters or that the wealthy can ever suffer consequences, and those two seem quite liberated by that.
b3lvedere 4 hours ago | parent
rozab 16 hours ago | parent
bayindirh 16 hours ago | parent
Keeping the garage door open, or at least making the door translucent. It's always cool.
jampekka 16 hours ago | parent
brookst 1 hour ago | parent
liuliu 16 hours ago | parent
lucrbvi 16 hours ago | parent
jampekka 16 hours ago | parent
I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
https://en.wikipedia.org/wiki/Training,_validation,_and_test...
liuliu 16 hours ago | parent
SwellJoe 16 hours ago | parent
esafak 15 hours ago | parent
kingstnap 15 hours ago | parent
It's not the direct feedback loop of RL but its not far.
brookst 1 hour ago | parent
It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.
Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.
nodja 15 hours ago | parent
sspiff 1 hour ago | parent
They result of the benchmark does not feed back into the training, it simply serves to provide a measurement of progression over time.
ProfessorLayton 16 hours ago | parent
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
joelwallis 16 hours ago | parent
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
james2doyle 16 hours ago | parent
I always found that those Mimo models to be really good at tool calling and following instructions
walrus01 16 hours ago | parent
girvo 11 hours ago | parent
walrus01 10 hours ago | parent
girvo 9 hours ago | parent
It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it
jonsoft 9 hours ago | parent
And I am not a web developer! It's an extraordinary model.
(Mouse and keyboard required)
flexagoon 13 hours ago | parent
rapind 12 hours ago | parent
I’ll give Mimo a try.
trollbridge 11 hours ago | parent
UltraSpeed was absolutely awesome. I miss it.
DS 4.1 Flash is amazing. Well worth the extra cost.
miyuru 7 hours ago | parent
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
baxtr 7 hours ago | parent
ignoramous 6 hours ago | parent
API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.
ehsankia 4 hours ago | parent
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
pelagicAustral 3 hours ago | parent
epolanski 1 hour ago | parent
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
epolanski 1 hour ago | parent
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
levocardia 16 hours ago | parent
SwellJoe 16 hours ago | parent
ricardobeat 14 hours ago | parent
conception 11 hours ago | parent
ricardobeat 3 hours ago | parent
jambutters 13 hours ago | parent
cpcabbge 7 hours ago | parent
iammrpayments 7 hours ago | parent
impulser_ 16 hours ago | parent
Where is the cool shit from the US labs?
culi 15 hours ago | parent
noir_lord 15 hours ago | parent
In the short term, true.
In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
impulser_ 14 hours ago | parent
dlisboa 12 hours ago | parent
The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.
reddec 54 minutes ago | parent
[1] https://www.amd.com/en/ecosystem/oem/supermicro/amd-instinct...
hsbalanxvxjsmab 10 hours ago | parent
impulser_ 8 hours ago | parent
Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?
Why do you care?
hsbalanxvxjsmab 2 hours ago | parent
atemerev 8 hours ago | parent
esafak 15 hours ago | parent
fzysingularity 15 hours ago | parent
dr_dshiv 14 hours ago | parent
dzonga 13 hours ago | parent
now we r just noticing the grave getting dug deeper.
skybrian 12 hours ago | parent
rapind 12 hours ago | parent
ijidak 10 hours ago | parent
Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.
epolanski 1 hour ago | parent
SlightlyLeftPad 12 hours ago | parent
passive 14 hours ago | parent
I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.
The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
ricardobeat 14 hours ago | parent
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
Cookingboy 13 hours ago | parent
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
arcanemachiner 3 hours ago | parent
brookst 1 hour ago | parent
markasoftware 4 hours ago | parent
ehsankia 4 hours ago | parent
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
buffalobuffalo 1 hour ago | parent
ernsheong 14 hours ago | parent
dr_kiszonka 12 hours ago | parent
hsbalanxvxjsmab 10 hours ago | parent
ttul 12 hours ago | parent
rao-v 11 hours ago | parent
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
gtirloni 10 hours ago | parent
thenews 11 hours ago | parent
hsbalanxvxjsmab 10 hours ago | parent
Retro_Dev 9 hours ago | parent
hsbalanxvxjsmab 2 hours ago | parent
ssn2000 9 hours ago | parent
jstummbillig 8 hours ago | parent
wg0 7 hours ago | parent
Google had this GPT long go and a wise man within Google noted:
"We don't have any maot neither does anyone else."
The AI bubble burst is guaranteed and is only delayed by IPOs.
user43928 4 hours ago | parent
Open models have not yet caught up with February's Mythos checkpoint.
Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
wg0 3 hours ago | parent
Alifatisk 6 hours ago | parent
monneyboi 5 hours ago | parent
singularity2001 3 hours ago | parent