Maybe less than we might hope. Not only is the knowledge distributed over the hundreds of people involved in designing a car, there may be conventions/habits of design and manufacture whose purpose has become hazy but are nevertheless important, or details of a design iterated upon whose original motivation is less understood.
The classic worry is https://en.wikipedia.org/wiki/Nuclear_winter: the soot thrown up by the ensuing urban fire storms causing temperatures, rainfall to plummet, and hence food production. On reading the scenarios the odds of total human extinction are lower than I recollected, but 80% of world pop dying of starvation dwarfs the direct death toll (usually reckoned in the hundreds of millions)
The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.
It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data?
What surprises me is they're not more boldly/plainly lying about it.
How would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.
Unless they know exactly the researcher’s account, they may not know in their end if he had the setting to let them train on his chat logs. They also probably don’t know if he had any correspondence on any forum where he may have discussed this and it got picked up by scrapers.
I’m not saying they didn’t do anything unethical. I’m just saying even if they were ethical, there’s plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
Agree model providers can't take much margin, a few percent plus maybe a bit extra from those willing to pay for a "better" model. Upstream is more concentrated: TSMC, NVIDIA, AMD could maybe raise prices and capture more of the value, which would affect open model providers.
LLMs don't "understand bitwise noise at a native level". The unit of perception for an LLM is a token, a soundly superbit level. They would have the same difficulty with bits as the number of "r"s in "strawberry". Yes they can be post trained to deal with such difficulties, but it's no more "native" than a human who memorised the ASCII table.
The degree of nativeness of the skill is orthogonal to the raw power of these systems. A frontier model has access to dozens of GPUs, each moving terabits of data, and performing trillions of matrix multiplication operations, per second.
Gradient descent and reinforcement learning algorithms don't really look much like evolution, unless you squint so hard that everything does (ie, squinted so hard that you've closed your eyes).
They're messaging each other by jamming strings in a constrained (unauthorised) side channel. Hence the lack of spaces. Unclear how much else of the weirdness is just from those constraints
And here I at first thought the direction the original post was going was "here's how I got a LLM to do this automatically". Seems like it would not be so difficult. Might still be be a win over their innate dispositions.
> BoVeX gives us a controlled tradeoff between these two states. By changing how much it costs for the text to be semantically wrong, we have a dial that allows us to smoothly interpolate between Lorem Epsom and Donald Knuth.
reply