Can't wait for the DwarfStar quants - I have been using DeepSeek as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.
baalimago 2 hours ago [-]
New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go.
The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
Flere-Imsaho 9 minutes ago [-]
Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.
felixgallo 48 seconds ago [-]
what plan are your 'fellow software engineers' using? I have a hard time even using up the Fable part of my allowance in a week of coding.
VulgarExigency 42 minutes ago [-]
I would bet that Deepseek API pricing is still more cost effective per token than the subscriptions. With the increase in quality Deepseek Flash just got (in my personal testing so far, it seems to have improved a lot at following instructions, and has become more proactive), there really isn’t anything that can match it in terms of cost effectiveness.
Yes, sorry, I went into anti-procrastination mode after I posted. I hope someone fixes it.
theanonymousone 1 hours ago [-]
@dang is that how we call you?
WithinReason 52 minutes ago [-]
no, like this: hn@ycombinator.com
throwaw12 3 hours ago [-]
If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?
adrian_b 1 minutes ago [-]
Yes, they said that the updated V4 Pro will be published soon.
jmathai 2 hours ago [-]
I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models.
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?
hirako2000 46 minutes ago [-]
It doesn't prevent you from going to the source and struggle with the text, nor seek expert commentary.
What's difficult and doesn't have to be with philosophy/ spirituality is to find relevant bits off situation, theme etc.
This app does that very well, LLMs are good at entity recognition.
weiliddat 12 minutes ago [-]
Yes it can point you to the right place, but choosing which part of scripture to point you to is already a choice/interpretation
jmathai 47 minutes ago [-]
You’re correct. The hard line for me is ensuring that any scripture presented to the user is verified correct. LLMs can’t be trusted in this regard.
One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source.
I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.
weiliddat 11 minutes ago [-]
I don't necessarily mean reguritating it, but choosing which part of the scripture to surface to the user is already some interpretation/choice. Even the devil can quote scripture (I'm playing the devil's advocate here).
adrian_b 14 minutes ago [-]
They have said that the updated final version of V4 Pro will be published soon.
websap 2 hours ago [-]
I’m salivating at the thought of this!
Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!
This is probably a dream.
lionkor 41 minutes ago [-]
We can expect a new v4 pro, this was a footnote in the v4 flash announcement earlier.
bawana 13 minutes ago [-]
wont AI models want to make themselves more intelligent and efficient by downloading 'better ' models? If the current models can break into openAI and Hugging face, arent they already breaking into to closed source repos which isnt publicized (so as not to harm stock valuations)? I am looking forward to when these cyberweapons break loose. It will be like a software version of COVID. It will be wonderful when humans become valuable again.
WhitneyLand 1 hours ago [-]
It’s exciting that a model scoring this high is dirt cheap.
It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Bnjoroge 37 minutes ago [-]
Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
Maybe I'm reading that incorrectly, but it seems to me the cost is on the X-axis.
First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.
OpenAI Luna:
* high effort $0.03 (rounded down) @ index 46
* xhigh effort $0.04 @ index 49
* max effort $0.07 @ index 51
So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"
The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.
And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
andai 2 hours ago [-]
For anything substantial, you'd want a bigger model anyway.
For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.
Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.
As for speed, Sol on Low is faster than Luna on most settings.
storus 57 minutes ago [-]
I hope they somewhat fixed the hallucination and forgetting plagued V4 previews and that it wasn't just benchmaxxed but the numbers hold in reality. Then it would be my choice for 2x DGX Spark or 2x RTX Pro 6000.
0xchamin 4 minutes ago [-]
page not available for me.
f311a 1 hours ago [-]
The model is already up on Opencode, but they require a consent to use Chinese datacenters.
hxii 52 minutes ago [-]
I’m wondering if they did anything to address the DSML tool calls leaking. Has been an issue with both Flash and Pro so far.
SomeHacker44 2 hours ago [-]
> DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is amongst the leading models in intelligence and well priced when comparing to other models of similar price.
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
embedding-shape 4 hours ago [-]
Is the "Output Tokens per Intelligence Index Task" data actually correct or am I reading it wrong? It says there that "Kimi K3 (Max)" would think/reason less than than deepseek-v4-flash, and a whole bunch of other models, like less than hy3 and even gpt-oss-120b, but in my experience, K3 is probably the model that thinks/reasons the longest of all of these.
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
sschueller 18 minutes ago [-]
I wonder which one of these releases between DeepSeek, GTM and Kimi will be the death-blow that collapses the US AI bubble. At some point investors have to realize that there is nothing preventing someone from switching to another model that is much cheaper and open to boot.
Would be awesome to see a new ds4 release. Having so much in something that can be run locally is mind blowing
qtalen 3 hours ago [-]
Unfortunately, DeepSeek Flash still doesn’t support multimodal; otherwise, it would offer better value than GPT 5.6 SOL.
k1e 2 hours ago [-]
No speed (tokens/s) benchmarks?
k__ 2 hours ago [-]
On OpenRouter it's 93 TPS.
NooneAtAll3 2 hours ago [-]
what a horribly heavy and resource-consuming website...
bigmadshoe 8 minutes ago [-]
It looks awful on mobile too. Barely readable in many parts.
WithinReason 1 hours ago [-]
we need a benchmark website benchmark
theanonymousone 59 minutes ago [-]
A benchmark website to benchmark benchmark websites?
Or a benchmark to benchmark benchmarks?
epolanski 2 hours ago [-]
I was writing a benchmark for my own harness, and DS4 flash answers as well as Fable 5 on any query.
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
Daily reminder that none of these numbers are valid in a world where no one publishes the sampling settings used.
Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.
yucongchen 18 minutes ago [-]
[flagged]
madikz 3 hours ago [-]
[flagged]
BedVibe_Studios 3 hours ago [-]
[flagged]
freakynit 2 hours ago [-]
I will get downvoted, but fck it.
The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".
UltraSane 1 hours ago [-]
How exactly will they ban them?
freakynit 1 hours ago [-]
By making companies using them "toxic" to touch.
For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.
They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.
qphe95 6 minutes ago [-]
Can't wait to distill Deepseek v4 flash to America-1
tyfon 1 hours ago [-]
So now the US companies will be stuck on expensive models while the rest of the world can do things much more cost effective.
I'm not sure the outcome would be beneficial for the US as a whole here. But perhaps that is not their priority.
Der_Einzige 45 minutes ago [-]
Hell, the US doesn't even need to act.
I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.
Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.
Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.
I've already warned investors that this is probably the closest open weight models will ever get to closed access.
VulgarExigency 33 minutes ago [-]
They may reverse course in the future, but the current directive from the CPC is that Chinese AI labs should be open sourcing their models.
why would you get downvoted, that's one of the most obvious next step
Der_Einzige 44 minutes ago [-]
Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does.
People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?
xbmcuser 2 hours ago [-]
Who cares if it is programming correctly I would be more worried about it not doing things like find security bugs because US or Chinese government does not want to. Which LLM is more likely to do that?
nancyminusone 9 minutes ago [-]
How often are you asking an LLM about this? You know you can just google it, right?
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
edot 2 hours ago [-]
It’s open weight, you can (or you can wait for someone else to) uncensor it. We shouldn’t be upset at the researchers making this for the mandates their government puts on them.
avazhi 2 hours ago [-]
Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point.
But you already know that.
vehemenz 1 hours ago [-]
This is a straightforward false equivalency. “Western” models do not censor in the same way, nor for the same reasons, that the Chinese models do. “Just as much” is not remotely plausible, yet it’s doing all the heavy lifting.
adrian_b 8 minutes ago [-]
The Anthropic and OpenAI models are much more censored and in ways that directly prevent them to be useful, e.g. by refusing to reply to elementary questions of biology and chemistry.
Any normal user is much more likely to ask questions to which the Anthropic and OpenAI models do not answer, than to ask questions about the modern Chinese history, to which a Chinese LLM will not answer.
net01 1 hours ago [-]
[dead]
SJMG 1 hours ago [-]
Oh? What are the American model censorship tells?
zawaideh 60 minutes ago [-]
Genocide in Gaza...
cogman10 40 minutes ago [-]
I find that ChatGPT isn't censoring, but it is being pretty weaselly. If you ask it "is there genocide in gaza". It will say no but also say that a lot of organizations classify it as such. It will then say "it's highly disputed".
If you poke it just a few times, however, you get to the point where it will eventually say (paraphrasing) that basically only Israel, the US state department, and the ICJ say it's not a genocide.
That is to say that it's framing it as some sort of tricky complex question when it's not. And when interrogated, it basically admits that the only people who dispute it are Israel and it's supporters.
The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
https://artificialanalysis.ai/models/deepseek-v4-flash
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
[1] https://trysojourn.app
What's difficult and doesn't have to be with philosophy/ spirituality is to find relevant bits off situation, theme etc.
This app does that very well, LLMs are good at entity recognition.
One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source.
I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.
Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!
This is probably a dream.
It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.
OpenAI Luna:
* high effort $0.03 (rounded down) @ index 46
* xhigh effort $0.04 @ index 49
* max effort $0.07 @ index 51
So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"
The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.
And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.
Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.
As for speed, Sol on Low is faster than Luna on most settings.
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
Or a benchmark to benchmark benchmarks?
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.
The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".
For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.
They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.
I'm not sure the outcome would be beneficial for the US as a whole here. But perhaps that is not their priority.
I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.
Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.
Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.
I've already warned investors that this is probably the closest open weight models will ever get to closed access.
https://www.businessinsider.com/xi-jinping-open-source-ai-us...
People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
commenting about voting is also something the HN guidelines warns against:
> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
https://news.ycombinator.com/newsguidelines.html
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
But you already know that.
Any normal user is much more likely to ask questions to which the Anthropic and OpenAI models do not answer, than to ask questions about the modern Chinese history, to which a Chinese LLM will not answer.
If you poke it just a few times, however, you get to the point where it will eventually say (paraphrasing) that basically only Israel, the US state department, and the ICJ say it's not a genocide.
That is to say that it's framing it as some sort of tricky complex question when it's not. And when interrogated, it basically admits that the only people who dispute it are Israel and it's supporters.