Hacker Newsnew | past | comments | ask | show | jobs | submit | jug's commentslogin

And without this harness it scores about 62%+, a dramatic improvement over even Fable 5.1 at 30%. I thought it just bears saying for context.

Political? It's using a Chinese LLM trait. If it only refused to talk about a particular kind of soccer, we'd probe it with that instead. The goal is not to discuss politics, the goal is to find out which model it is. What's political here is the language model.


From what I'm seeing I do believe TikTok is worse.


I really like the combo 5.6 Luna & Sol for price and performance and would be perfectly happy if they stayed here for a moment without mucking about with sidegrades that I think AI evolution has often felt like lately.


This is true and is only becoming more important the more they improve. I am already moving to checking so they're at least somewhat following the status quo and otherwise prioritizing price and platform. I think this will be an emerging way of viewing AI in 2027 and the winner will probably be open models and China.


I think this likely plateaus and we all just get the smartest intelligence humans need running locally…


Don’t think it’s happening any time soon for most people. My Mac has stayed 32gb for many years now. I don’t think I’m moving into 128gb territory any time soon with all the price hikes.


If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. cost efficient models aren't necessarily invalidated early by progress that matters.


Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost efficient way, and especially now!


Kimi K3 is fairly cheap per token but thinks like a madman with poor self esteem.



In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.


There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue down this path of optimizing reasoning for other models.

https://bottlecapai.com/post/thinkingcap-qwen3-6-27b/


Yeah, they thing forever and doubt everything "wait but" for 200k tokens for almost any question.


On the flip side, I really like being able to inspect its reasoning chain thoroughly, as opposed to the "black box" that Anthropic models are now.


Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open


That's a summarized and filtered view of the actual reasoning.

OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.


Older models did show the full unredacted thinking trace, but I don't think Opus has ever shown full CoT.

Here is an archived version of Anthropic's API docs saying that Sonnet 3.7 (only) has unredacted CoT on API: https://web.archive.org/web/20260324051339/https://platform....


That’s cool. I knew o1 hid it since launch, so I assumed Anthropic would also have never shown it.


Got it. Thanks


Right I was going to say, no way of knowing whether these issues are unique to Chinese models.


Depends on whether the models report the correct amount of tokens.

5.5 Sol repors 10x fewer reasoning tokens than Kimi k3. If it is correct, than it unlikely has those doubt issues.

At the same time, I feel like their reporting is incorect and we are now paying per "intelligence", not actual tokens. We can't verify it anyway..


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: