This is straight up dystopian level spying. Streaming microphone inputs all the time, knowing everything that is displayed on the screens via object recognition and stuff. Scraping your local network. If I’m not mistaken there is the possibility for the tv to connect to public/open networks to send this stuff without you knowing. Magically syncing with your phone. It’s truly terrible, and a class action is all but guaranteed to come. Gamers nexus has a great video on this.
No, you're not up to date with the latest news. There's a separate feature that records you all the time, transcribes it to text and uses it to select ads for you
I believe this is big news. Speculations here on out: I imagine this being baked into consumer products, greatly increasing the local token capability for consumers. They will suck the cloud-oriented companies' milkshake. Most users do not need extremely capable models, they just need some automation to do better web-searches, and get simple facts etc. If it can do simple coding tasks too, but at thousands of tokens per second, in stead of tens or hundreds, the development will benefit so much. It will benefit AMD in other ways too. I imagine they can start selling physical chips, usb-drive like devices, that just does llm. If you want a newer, better, model, you simply go to a store and buy one. Need more capability, buy more drives. Similar to physx back in the day, but with usb-c and a smaller footprint.
I’m not saying zed necessarily should be the one to do this, but in regards to "why not git, jj.."; if we don’t explore the fringes, how do we know if we are at local or global optimum with current solutions?
The magic is in the fact that they essentially have an ASIC llm device. There is no other trickery. The problem they will face is that it is actually locked in silicon, so upgrading models will be difficult, and likely require new hardware each time.
At some point I imagine you’d add a software layer on top that holds more current training and can be called as needed trading off for slower responses. There’s already work out there splitting models across networks. You could have the base on silicon, some stuff in memory on the machine, and another frontier tool in the cloud.
Taalas actually already support LoRA, basically doing exactly what you say.
The other thing that I think is really interesting about all of this, is that LLMs are already perforce behind the times with their knowledge cutoff, so adding an additional ~3 months for bake into silicon isn't such a huge deal, I think, for the ~10x more efficient and faster you get.
Yeah, really a fascinating time in computing. Once these chips start to become more common I'll be curious how people find ways to use them. One can imagine a world where really simple inference is available on dirt cheap chips found in toys and other low cost consumer electronics.
There are already studies proving that a "stupid" model with a good harness + tool calling will outperform a "smart" model.
Things like this give me hope for a system that can be fully local and private, but also with the ability to be almost infinitely extendable with tools.
Is the size of the ASIC limiting factor or can they infinitely tensor parallelize? If thats possible then it would make it only a matter of economy of scale, there are good enough models already for people to invest in that kind of platform
If this all works, and becomes widely available, I wonder if people would get a lot of health anxiety. Now they can see stuff that is normal, but strange or unexpected.
I’m no doctor by any means, but what if, as an example, an organ fluctuates in size, or composition, naturally. A medical professional would know these things, but a random person off the street might get stressed out and start to panic, or perhaps overcompensate with their diet or something.
I think more data is generally good, but data without context or insight can be problematic.
reply