Hacker Newsnew | past | comments | ask | show | jobs | submit | haritha1313's commentslogin

users have complete control of what a good output looks like. we have seen that adding criteria that is specific to what the expected output is, including at trajectory level, helps a lot to avoid false positives


This just seems like lousy testing. Why was the guardrail just an understanding with the third party and a prompt and not actually tested for edge cases? Starting to wonder if its because of agents monitoring agents' work and yolo-ing it.


?


openai just signed the letter


sam must have read my comment.


I keep buying the pretty notebooks and collecting them. I hate writing on them because they are too beautiful. I've finally started using random company merch notebooks (that are not pretty) while I also carry around a pretty one for the pretty thoughts (which I rarely use because nothing ends up deserving it)


The reality is most people building their own models and providing that alongside SOTA ones don't really care about how great these models are. They just prove that 'hey we are smart enough to build our own models so you can trust us instead of going with a single provider like Claude via Claude Code', also a cheap alternative for cost sensitive/free users - at least this was the case for Windsurf, not sure if Devin Desktop still has that tier. They just need to hillclimb the benchmarks and show something reasonable enough there.


Ah this is so cool! Wish I had it in my research days.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: