Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Use nvidia hardware instead and use a larger cluster serving many more users concurrently. Easily 10x–20x higher token rate per GPU with public solutions like dynamo and sglang.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: