Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
LUmBULtERA
64 days ago
|
parent
|
context
|
favorite
| on:
Claude Sonnet 5
That's yet to be determined. I think a lot of open-weight models are benchmaxxed and their usefulness for many tasks are not represented by those.
enraged_camel
64 days ago
[–]
Yes, this has been my experience. They all struggle with long-horizon tasks and eventually start going in circles.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: