25 points stikit 1 hour ago 49 comments
mariocesar 53 minutes ago | parent
For more long work, I now use Fable to create a PLAN.md. I tell it to make a plan that will be executed by other models, and most of the time it ends up choosing Opus or Sonnet.
I didn't start doing this recently. Before that, I would just use the top model for everything. Splitting the work across different models depending on the task has helped a lot. They run faster, and I usually get much better results
sourcecodeplz 33 minutes ago | parent
mariocesar 26 minutes ago | parent
Here is the script https://github.com/mariocesar/dotfiles/blob/main/common/.loc...
notnmeyer 19 minutes ago | parent
bellowsgulch 52 minutes ago | parent
most engineering tasks don’t require frontier llms
when they get stuck, then i consider moving up to more capable models
purchasing a claude plan seems widely unnecessary to me
the tasks they do better than the average engineer cut both ways: unless you have an existing portfolio of well written and designed work done pre-llms, it looks like you’re producing slop that pretends to be well designed
poor typography choices despite using the mode,
poor layout choices despite using popular CSS frameworks
etc
bad engineers will always be bad engineers
tools don’t make up for it
edit: a follow up to this— everyone is using eyebrows in their layouts and have no fucking clue why it was done to begin with
everyone has a status pill floating above their front page hero display text and its not fucking status related
so gross
admiralrohan 24 minutes ago | parent
jraedisch 51 minutes ago | parent
alstonite 50 minutes ago | parent
ewindisch 38 minutes ago | parent
Astra low on the $200/mo "20x Pro" plan gets me through a single day.
jbonatakis 27 minutes ago | parent
vallerie 50 minutes ago | parent
- like a fancy auto complete (here are some stub methods, they should do X, fill them in)
- using fairly detailed plans and test harnesses, so blowing up the world is hard
The 3.X Flash family have been fairly capable models, and the selling point for me is just raw speed. Gemini is noticeably faster than the competition, about 3-4x, and I just get work done faster with it.
That said I'm keeping an eye on Open Weights. DS4 Flash was good until price hikes, and finding a provider that serves at high speed and without quantisation at the prior price is tricky.
martythemaniak 20 minutes ago | parent
There's no Gemini pro model currently, so you gotta pair that with a 20 openai plan for access to more advanced stuff if you need it.
traverseda 46 minutes ago | parent
hmokiguess 42 minutes ago | parent
herpdyderp 41 minutes ago | parent
- Preferred: Claude Code with Opus 5 Medium
expedited123 39 minutes ago | parent
philbo 38 minutes ago | parent
o_m 37 minutes ago | parent
I also don't want to use the Claude Code and Codex agent harnesses. The good thing with Codex subscription is that it can be used in other harnesses, unlike Claude. As far as I know, only Anthropic has this restriction.
r_lee 11 minutes ago | parent
It's really strange because when Opus 5 released, there were some that pointed this out, but a bunch simply said it was the best and as good as Fable etc etc.
but for me, it caused me to get very demotivated and avoid interacting with the model, at least when using Claude Code.
sourcecodeplz 36 minutes ago | parent
unbeatable price/intel ratio per M tokens:
$0.10 (input)
$0.20 (output)
$0.002 (cached-input)
jinnko 34 minutes ago | parent
w22oop 33 minutes ago | parent
the__alchemist 32 minutes ago | parent
Granted, yesterday I threw a few tasks to Astra which the former 2 botches; it produced clean, correct solutions quickly, so pending further eval, this may take over.
IMO unless it's a mechanical tasks, it's worth it to use carefully -crafted queries on the more expensive models, than iterate through messier solutions on the cheaper ones.
codazoda 31 minutes ago | parent
For my personal stuff, I'm on a small $20 plan, so I need to use tokens conservatively. I was very rarely exceeding limits until I built a Dark Software Factory. It's not as efficient at token use. So, I use Sonnit over Opus here.
At work I have a $100 plan that I rarely exceed so I use Opus. I have access to Fable too, and I did use it a lot while it was new, but I don't find it improves most of my work by too much. I do mostly bug fixing across several hundred repositories with hundreds of thousands of lines of code, mostly written by humans over the past 20-years. These projects interact with each other so I run claude from the root of my sandbox (I was nervous to try this but I'm not looking back now).
I also use Sol as a secondary for my personal work. I pay for it because I like to talk to ChatGPT on the web. Since I already have the subscription, I let Sol write plans for me. It does a better job at certain tasks and it saves me some Claude tokens. Maybe I should consider Terra for the task, but I don't run up against my usage limits for the little bit I use it.
I'm trying to use Gemma 4 12b for some workloads but I haven't mastered the model yet. It's still very experimental for me. I can get work from it but it takes a lot of hand-holding. For local models, however, it's all I have the RAM for.
pighive 20 minutes ago | parent
Havoc 29 minutes ago | parent
...and then sprinkle in some other models when i think a second opinion will help
oduis 27 minutes ago | parent
amelius 24 minutes ago | parent
tiluha 15 minutes ago | parent
Privacy policy is not great with deepseek api, but you are always just trusting their word with any hosted llm and in theory i could at least self host the models i use from deepseek.
time0ut 23 minutes ago | parent
I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it.
I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others.
Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do.
Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out.
shelled 21 minutes ago | parent
A day ago I activated Google AI Pro free via Google's tie-up with a local company (I do pay for this company's product though and it's anything but costly). Now I will use this too.
No other reasons to pick these, or not picking anything else.
rpmisms 21 minutes ago | parent
donatj 19 minutes ago | parent
behole 19 minutes ago | parent
rsyring 18 minutes ago | parent
- LLM expense budget
- What type of dev: work, personal, real time spaceship thrust vectoring, html contact forms for family, etc.
- human in the loop with short as possible turns, software factories that can run for days, or something in the middle
Just off the top of my head. I'm sure there are others I'm missing.
hgoel 17 minutes ago | parent
Used to pay for a Claude 20x plan and did everything in Opus, but I hate how it talks now and recent events (OAI scooping, Anthropic's spying, third party Chinese model hosts stealing and selling credentials) have really pushed me towards local AI for personal needs. Am not allowed to use Chinese models for work even if self-hosted so not much choice there.
cyanydeez 17 minutes ago | parent
seanmcdirmid 16 minutes ago | parent
conradludgate 14 minutes ago | parent
For work, given we pay API pricing anyway, I've been happy experimenting with K3-medium as my default, and now I'm planning on trying GLM 5.3 as well. Sol-medium is my fallback for my work purposes but I occasionally use Opus 4.8 as an additional reviewer.
I like the open weight models because much of my time at work is spent on security hardening (specifically, hardening my own service), and Sol/Opus keep snitching on me and blocking my prompts.
m0rde 14 minutes ago | parent
Defaulting to Sol Medium/Light for most planning and implementation I think is tricky or want more care in.
Luna Extra High for everything else (implementation, tedious take over my browser and do stuff).
I read all of its output tokens and lots of thinking tokens to understand the general flow of things, but only minimally look at code these days. I can't grok what Claude models speak and it's gotten worse. OAI models speak my kind of tech language I guess.
Light human review, some automated review.
Most work is for internal use.
seabrookmx 14 minutes ago | parent
I'm not as up to date on the other vendors' models, but when I last used Gemini my feelings were similar between Pro and Flash.
beej71 13 minutes ago | parent
ghosty141 13 minutes ago | parent
montroser 9 minutes ago | parent
For what it's worth, here's a take on its speed vs cost vs intelligence: https://artificialanalysis.ai/models/deepseek-v4-1-flash
I can go all day and night with this thing with multiple sessions going, and I spend like $2 per day retail. With opencode-go, that fits within the $10/mo subscription, so that's what it ends up costing in real life.
Cakez0r 9 minutes ago | parent