Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I tried to get Qwen3.8 27B working properly under high concurrency, and while the quality level is spectacular for the size, the performance wasn't the best, even with MTP. Unless you have a very big infrastructure, it's difficult to run a dense model concurrently with high throughput.I suppose that's why almost all large models are now MoE. On the other hand, 1500 tok/s is an impressive speed, and that speed is very important for agent tasks, so a service like this instead of local infrastructure might make sense, although it also depends on your busines constraints.
 help



Completely unusable last time I tried it. You got hit with rate limits after the first few minutes of using it. It's fast but cannot sustain its claimed speed before it immediately hits its rate limit. What good is it if you can't finish a task? It's like having a sports car that can only drive 180 mph for 5 seconds every minute and then has to cool down for an hour.

Borderline fraudulent to advertise to developers when it's completely impractical to use. Their only support is their company Slack channel.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: