Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Even more important in a local context is the difference between token generation and prompt processing speed. We tend to focus on the former, but for multi-turn/agentic workflows the latter can dominate.


Yeah definitely. I've recently commented on that: https://news.ycombinator.com/item?id=48557890




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: