Hacker Newsnew | past | comments | ask | show | jobs | submit | Kostic's commentslogin

Sure but that holds true for most new tech. Expensive and slow but it can only get better or stay the same. I'm thinking it's only going to get better.


Your video is showing Kimi 2.6, not 3? Once the weights are released, there might be providers that serve it without the censorship filter.


A very beautiful website and a machine. Oxide folks should be proud, you can see the love that went into it.


Absolutely


Time to copy&improve the niche solution.


70% of the Steam users have a PC that's weaker than the Steam Machine[0]. These things will sell out, that's not the issue. The issue is that Valve probably won't have enough power to secure hardware deals to fulfill all the orders, thus limiting the growth of their hardware side.

[0] https://www.techpowerup.com/342970/valve-claims-steam-machin...


If they have a PC weaker than the Steam Machine, they were obviously not willing to pay for a stronger PC even when components were cheaper before, so how is the situation better now?


For personal needs I connected VSCode with llama.cpp running Qwen 3.6 27B or Gemma 4 31B and it's good enough to cancel my cloud subscription.

Qwen running on my 1st GPU at q4@176k context from 70 to 50 tok/s with MTP, pretty good for coding.

Gemma on the other hand is using both GPUs, running q8@64k context, doing document sentiment analysis, summarization, proofreading and translating, at consistent 25 tok/s. Somewhat slow but usable for batched workflows. Might get some more once llama.cpp starts supporting MTP with tensor split mode.

Still using frontier LLMs at dayjob since I'm not paying it and those are obviously better. Hopefully we'll have a Sonnet 4.6/Opus 4.5 level 30B model in a year or so.

EDIT: Prompt processing starts from 800 t/s and drops to 400 t/s. In most cases my starting prompts are around 16k-24k of tokens and require from 60 to 90 seconds to be processed. Not great but acceptable.


What extension do you use in vscode to connect it to local llama.cpp? Or do you auth with github copilot and then point to localhost? Or something else?


Auth with Github Copilot and then point it to localhost[0]. Hopefully the auth to Copilot requirement will be dropped for local models at some point. Would love to use a fully open stack (VSCodium and everything) in the future. My config:

``` [ { "name": "http://127.0.0.1:8888/v1", "vendor": "customendpoint", "apiKey": "llama.cpp", "models": [ { "id": "gemma4-31b", "name": "Gemma 4 31B", "url": "http://127.0.0.1:8888/v1/chat/completions", "toolCalling": true, "vision": true, "maxInputTokens": 65536, "maxOutputTokens": 8192 }, { "id": "qwen3.6-27b", "name": "Qwen 3.6 27B", "url": "http://127.0.0.1:8888/v1/chat/completions", "toolCalling": true, "vision": true, "maxInputTokens": 180224, "maxOutputTokens": 8192 } ] } ] ```

[0] https://code.visualstudio.com/blogs/2025/10/22/bring-your-ow...


i made this specifically for use with vscode/llama.cpp: https://github.com/khimaros/mortar


"I'm sorry" is not gaslighting but an admission of fault it learned from our texts. And if an LLM managed to delete your database, it's time to slow down the vibe train and put up some guard rails.

LLMs are awesome but not without supervision.


Hard agree on the guard rails bit.

Would it be less sucky if an intern accidentally deleted the database? If not, take some steps to make sure no one can delete it without jumping through visible, noisy hoops.


Taalas showed that you could make LLMs faster by turning them into ASICs and get 10k+ token generation. It's a matter of time now.


Actually pretty interesting to think: in a few years you might buy a raspberry pi style computer board with an extra chip on it with one of these types of embodiment models and you can slap it in a rover or something.


For now. These will be pretty cool Linux machines if Asahi starts supporting them at some point.


It's going to take 3 years+ and it will be a 8GB RAM linux machine.


Still would be a good ssh terminal or presentation laptop.


I would not go below q8 if comparing to sonnet.


Yeah. Q2 in any model is just severely damaged, unfortunately. Wish it weren’t so.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: