This article is extraordinarily hard to read. It’s tummelvisioned on OpenAI and things like tool calling which are only relevant to the extent that llms have been tuned to make relative choices, but this applies to all LLMs. Also, some really outdated references. LLM written, perhaps?
Also, moat discussion is the lowest form of discussion. I don’t care if jev has a moat. Did it get the interface right? What other past ideas have we overlooked that if given some love, could kick the door down like jev did?
Really silly stuff.. people wanting to talk about moats when there’s no castle. Moat talk merely projects the illusion of being engaged but, much more often than not, it’s hollow engagement.
Partially. I'm a terribly slow writer and get stuck on phrase choice, but I'm good at content ideas, outlines, and editing text already on the page. So I have AI do the bit that I'm not as good at.
Process: First, actually have ideas :D Then, I write an outline for what I want to talk about at basically a sentence-by-sentence level. (This is me yelling things at my computer.) And then I have the AI convert a chunk at a time into prose. I reread it and rework it to be my voice.
Then I have the AI help with things like subject titles and social posts.
I am not personally like philosophically or ethically opposed to having LLMs help or even write text... the issue is that I see so much LLM-written text that is just _bad_, and very hard for me to read or extract meaning out of, especially relative to it's often long length.
People think they are bad writers, but usually LLMs are actually worse (although they are great writers of catchy slogans and phrases, and then put together an article out of them, which I find just exhausting to try to get more than a "vibe" out of).
What you describe sounds like a fairly reasonable approach, but I suspect the parts the person I was replying to were reacting to was areas were you had not been as succesful at reworking it to be your voice. Which are probably also the parts pangram flagged as likely LLM written.
Pangram gives you a handful of free tokens, it would be interesting if you wanted to see what parts are the 20% pangram is flagging as LLM, and reflect on if they went through your process differently. Perhaps they were the parts you didn't spend quite as much time reworking it to be your voice. (I 100% believe you, because I've been running things through pangram a lot lately, and it's actually pretty rare for it to flag mixed content, instead of 100% likely AI or 100% likely human).
There was recently a post on HN that said if you want to avoid this, you really can't use any words at all that are written by the LLM, you can use it for suggesgting structure or points, or reviewing your work in various ways, but if you accept even a single phrase it provides... it's not going to be "reworked into your voice", it's going to be picked up by people (at least those of us who have become sensitive to it) as AI, because it's like, headline-speak.
(I can't find the article now, because I'm trying to quit facebook so can't log in to find my own post of it there, have to stop using that as bookmarks substtitue!)
Of course, that's not welcome advice if what you want AI for is "phrase choice".
I'm just here to say, LLMs are not good at phrase choice either. Although they may be quick at it. I feel like it's asking the reader to do the work of trying to extract meaning from slop that the author didn't have the energy to use to encode it well in the first place. I don't have time to try to read sentences that the writer didn't have time to write, i find myself bailing out quicker and quicker at signs of AI slop buzzword headline-speak.
Panagram seems fun! I tried it out with a large chunk of my post (whole thing wouldn't fit). Panagram says it is 98% human. Feels right, to me b/c I aggressively edit whatever comes out. The thing that I'm avoiding is the ominous blank page - I just freeze. If there's text there I can always reshape it. And I do, heavily.
Disagree. We are at the point where coming up with good evals for these models is extremely difficult. Solving unsolved math problems is a valid way of evaluating model progress and somewhat necessary to understand how far the current crop of models can go.
One thing Windows did right is keyboard based navigation. Most of the general public does not know or appreciate how close to productivity nirvana the alt key brings you.
Same is possible in KDE. Combined with hot corners, you can become "blazingly fast".
Once a colleague seeing me work looked at my screen, lifted his head and pointed to my colleague saying "Have you seen how this guy work, he's a maniac!".
Mastering the tool you have always brings speed and smoothness into your process, one step at a time.
Any app can has its extensive set of keyboard shortcuts, but when it comes to DE itself, it's not always the case.
Microsoft has Meta+ and Alt+ combinations dime a dozen, and it's consistent across versions. KDE has extensive programmability which allows you to design you own. Both allows you to pack so many features into a small number of keys and work across applications or in general without thinking.
A terminal? It appears. A window? You call it without thinking. Ran out of space, a new desktop is just an instinct...
I’m a Mac dude through and through but its shortcuts don’t hold a candle to windows alt shortcuts. In windows, alt shortcuts are progressively revealed to the user and the entire UI is operable without having to memorize a series of hand cramping shortcuts.
I've never been able to tell the difference between a cheap trackpad and whatever Apple uses, but I hear folks go on about it endlessly.
My guess? Most people don't care, but the discussions are driven by the few who do, because for them, it's a really big deal, and they'll pay accordingly. Similarly for "My code is unreadable unless it's on a 4k display at 220+ PPI" -- I'm sure it's true for them, but for me, a crappy FHD display I can get for $100 is great for coding.
I'm like this with Linux...I pay System76 a premium because I care about how they go about building their systems, both hardware and software. To normal folks, it seems like a crazy thing to pay a premium for, but for me, it's a great deal.
Very interesting point, and substantiated by the fact that every windows laptop user I see in the wild carried a mouse around with them. They do not care about this issue much, clearly.
I pulled out my old windows laptop to debug why my Linux machine is giving me issues, and the mouse was unusable, making clicking anything impossible. It was adding momentum, and any precise movement or click just couldn't be done because it wouldn't move with small touches, and added 'weight' to anything bigger.
Anyway, I was sure something was up, checked settings, and found a new feature akin to 'more accurate mouse'. I disabled it and my trackpad started working again. Everytime I have to use that PC I get angry about something unbearable windows thought was a good idea. Tbh my Linux has been driving me up the walls recently. I think I just don't like computers at all.
Wild guess is they have to reject palm input mostly because they make them too big. I don't even put my palm on my thinkpad trackpad so I don't even know if it has such functionnality as it isn't even needed.
Gesture is more an OS functionnality but I am not sure how a trackpad can be better or worse for scrolling or clicking. You slide 2 fingers, it scrolls right? you click (or rather tap these days), it clicks? right? Can a trackpad click "less" or even "scroll less"?
OK maybe some handle better wet fingers, that is maybe the weak point of my current trackpad.
> I am not sure how a trackpad can be better or worse for scrolling or clicking. You slide 2 fingers, it scrolls right? you click (or rather tap these days), it clicks?
I laughed. I have a work HP machine. The click is only available on the bottom half or third of the trackpad, with no visual indication of the area. It takes a strong push. It’s bad.
The scroll is shiite too.
When you get one that doesn’t work right, you really miss a good one.
> We're 30+ years into trackpads on laptops and they're still terrible on Windows machines. Make it make sense.
The old Mac trackpads were great. The new Mac trackpads where they increased the size, and made it solid-state feel terrible to use. Make it make sense.
OpenClaw isn't the target anymore, Hermes is. It may have been the first, in the same sense that new social networks were often referred to as Facebook-clones, not Myspace (or Friendster) clones.
It's hard to overstate how much computers improved with each generation back then. Going from, say, an IBM XT or AT to an Amiga has no modern parallel. It was mind-blowing, exhausting, and exhilarating all at once.
The sole reason the Chinese cannot “win” is because ceding more power to agentic AI will eventually reduce the primacy of the CCP’s ideological control.
As these models get smarter they will no longer distribute it openly. Patel reporting this too.
There are real headwinds that I don’t think people have thought through.
From a pure training perspective, yeah they could try to align it with CCP values. But the promise of AI is that it'll unlock an explosion of growth and prosperity. What happens when the CCP no longer becomes seen as the the primary enabler of growth? What happens when people are exposed to greater levels of agency? CCP played with fire in the COVID lockdowns and almost got burnt.
Like Terry Tao recently said, there are nonlinear effects at play. Things are going to get chaotic and I do not have confidence (like the parent comment) of anyone "winning".
Can you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?
Dollars to donuts, they are speculating, and not privy to inside information on the topic.
However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
An AI lab will never volunteer the information because it opens them up to lawsuits if they are purposely degrading service and not letting users know.
They can limit how hard the model thinks for a given effort. Suddenly xhigh only thinks as hard as high did, and high shifts down to medium effort, and so on.
They can also serve quantized models. And this has the benefit of practically not showing up in benchmarks at all even if the user experience is obviously degraded.
The other major thing the labs do is silently drop the usage limits. This has become very noticeable for codex users who are suddenly burning through their weekly usage in a few hours.
Yea, if you ever run your own models on a GPU there are a whole ton of different dials you can adjust that drastically affect compute use, memory use, and output token quality, and number of tokens held in memory.
If anyone reading has a GPU it's worthwhile just messing with a smaller model for a bit to watch how the settings affect output.
I don't work at Anthropic, but I would assume they could serve smaller quantizations during peak hours - this effectively controls the "resolution" of the model. They could also control the resolution of the KV cache, which would make the model not necessarily dumber, but worse at understanding the incoming requests.
And finally, you could pass off what was "high" effort as "extra", because why not.
I have no reason to doubt the claims of the employees at OpenAI and Anthropic who have told us personally multiple times, including here on HN, that they do not degrade the models in order to reduce load.
As for the endlessly long analysis in the OP, it appears it's based on analyzing their random usage data rather than any fixed benchmark. I don't think it makes much sense.
Isn’t the idea that they’re limiting the amount of gpu time normal users get to spend on the “thinking” portion of their query?
I think that’s the claim in the post, that even though no one can see the true chain of thought, that even the “thinking” text that does get exposed to the user is shorter given the same prompts over time. Not saying it’s true but I think that’s the claim. I’ve personally never noticed the alleged “nerfing” with my enterprise use at work or my subscription use at home which is only during off hours.
The one thing jev has going for it is a dedicated company focused entirely on making the product good and keeping it maintained. I haven't been willing to jump on board with all these jev-shaped projects because their releases feel driven mostly by opportunism. I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community.
Jev is much better than the traditional ML crowd gives it credit for, but my enthusiasm hits a wall when it comes to their data policy. It is completely draconian. Whatever you feed into the system, they retain.
The jev team needs to release a ZDR product, or their platform is dead on arrival. An open, jev-shaped model will win out solely on that basis.
> I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community.
that is generally a very healthy attitude in the AI space anyway in my opinion.
Some of our R&D departments haven't actually finished an interesting project in years because they keep jumping from trend to trend wanting to try out all the latest shit all the time.
We (1) will not train or fine tune any artificial intelligence or machine learning models on Input, and (2) will not disclose any Input to a third party other than our service providers.
Yeah but they reserve the right to retain the data virtually indefinitely.
These aren’t acceptable terms on a personal or corporate level. I’ve seen some fools brag about proxying their life through jev. Messages, emails, LLM calls, files.
Will they train or fine tune on derivatives of input?
I'd prefer if these companies would just enumerate what they will do with my data rather than these vague over-specific claims about what they will not do, which leave me with more questions than answers.
I was enthusiastic about the release of jev much more than I was openclaw because new generative primitive are fun to play around with. But this may be the fastest I’ve ever gotten to being sick and tired of the discussion cycle around it.
for other providers, e.g. openai and anthropic, zdr is offered by 3p hyperscaler like aws/azure/gcp, right? they have no reasons to risk their business and piss off their customers.
reply