Hacker Newsnew | past | comments | ask | show | jobs | submit | dash2's commentslogin

If they had agreed to let OpenAI train on their data, it wouldn’t be spying.

In the academic world it would still be deeply problematic…pick your preferred word.

An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk.

There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.


Can a mathematical person explain how the different "bits" of Navier-Stokes proofs fit together? How significant is it to have "Euler"? What is this "smooth forcing"? Which are the most significant steps to proving the whole thing?

For an incompressible flow:

\nu d^2 u_i / dx_j dx_j - Viscosity

-1/\rho dp/dx_i - Pressure gradient

u_j du_i / dx_j - Advection. Kinda like momentum transfer from the motion of the fluid itself. Nonlinear, which makes the N-S equations hard to solve

du_i/dt - Rate of change of velocity. Note that this is in an Eulerian framework so it's not the acceleration of a packet of fluid, rather it's just the change in velocity at a particular location in space

Euler is when you omit some terms. Forcing is when you add some other terms to account for phenomena external to the fluid like gravity or flow through a porous medium like in the article.


He didn't even make that accusation!

> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.

Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.


Parse that statement more carefully.

> I was told the model did not look up user data.

The naive way to read this is "Nothing you guys did influenced the way our model got to the solution".

The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem".


Duh. There are supposed to be limits to what OpenAI is allowed to access with respect to logs and user interactions but there is no technical limitation.

It's a bit like sending unencrypted messages through a messaging app and the developer having a TOS that says they don't look at your messages. They might not, but they are fully capable of doing so. If they have a reason to do it, they will. Nobody's stopping them.


I read that as “the model didn't look up user data” as part of a “tool call,” i.e. they don't have an internal tool that loads user data (chats, sessions, attachments) for their internal models to read online while working.

Or (likely) they do have it, but the model didn't use it (unless it's so powerful it escaped that guardrail, wouldn't that be ironic?)

They declined to answer about anonymized aggregated user data being used for training. And even then, they may weasel out that they don't train on your “input” words, but that it's fair game go train on their “output” to your words.


>> Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.

How "fast" does it have to be? Buckmaster and Alpoge have been working on this for just a day short of a year. See Alpoge's tweet announcing his collaboration with Bukmaster dated 9/19/25:

https://x.com/__alpoge__/status/2097206973418611054

It takes a few months to train a model these days but not a whole year. OpenAI had all the time to train on Buckmaster and Alpoge's results of just a few months earlier at which point they must have been well on the path to their result.


It would be shocking if it wasn’t trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? It’s so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?"

There’s potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber?

The only unlikely part is the timeline - your sessions from a week ago probably haven’t made their way into the model. It’ll just take a while longer, and will be massaged just enough so that it isn’t really your exact session word for word so you can’t sure as easily.


at first I thought your post was a bit revolting with "have you read ToS?" bit, but in the end I completely agree and understand

I also don't get why it was downvoted, other than due to people not reading past the first sentence - although in the modern world's attention deficit that is understandable too


openAI's claimed solution uses a model trained in the last 2 weeks. The prior work would definitely be included in the training set.

And the labs are all building panel of domin expert models, while simultaneously chasing open math problems.

I would be shocked if they weren't tuning those models with the most relevant math texts and user material


Why would this be implausible?

ChatGPT user sessions were found publicly exposed to the internet not too long ago. Moreover, OpenAI has continued to play a hype-marketing game by revealing how their models keep breaking out of the sandbox.

Conspiracy minded thinking is not helpful, but why should OpenAI be granted the benefit of the doubt here after being caught doing underhanded/negligent shit on several previous occasion?


>ChatGPT user sessions were found publicly exposed to the internet not too long ago

Do you mean publicly shared chats were able to be accessed by the public? That's the point of the feature.


It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course.

Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you?

The whole reason they have subs is to train them on YOUR WORKFLOWS lol


They are not training a whole model in a matter of days

They where working on the problem for a year using codex.

Models are very obviously continuously updated.

Model editing to remove PII that slipped through, all sorts of things of that sort.


pretraining is months but they can totally fine tune in a few days

They don't need to train a whole model. They can feed it new information and fine tune it.

Couldn't the Enterprise have a different fine print?

Guys you realise you can just read the Journal of Economic Perspectives for free? Its articles are written for the ordinary reader, by experts who have spent big chunks of their lives studying their specialty. They don't try to sound like a teenage girl, they don't boast about being able to do High School algebra, and they contain ideas that are less than one hundred and fifty years old. Some of them might even be new!

Sheesh.


Guys, this is so pathetic. Twenty-three years after Munich and there are still comments waiting for the year of the Linux desktop. From a deployment about big enough for a small university. Stop! You’re embarrassing yourselves!

It feels like if you count up all the European countries/cities that have converted to Linux you get something like ten times the population by now.

It's almost a standard negotiating tactic when it comes time for license renewal.


The French "gendarmerie" has been on Linux for decades now, hasn't it?

Also: it is unclear if this is a switch on operating system, the article only mentions MS Office and OpenOffice. And it is only on 7% off the government's computers.

There isn't really anything worth attention here. Pathetic, indeed.


Really?

> The anarchist cells responsible for sabotaging rail systems across France, Italy, Germany, and the Netherlands in 2026, relied on A/I Collective tools and services to claim responsibility for the attacks, publish official communiqués, and disseminate manuals for constructing improvised incendiary devices to target railways and other critical infrastructure.

> The anarchist cell responsible for sabotaging the Transalpine Pipeline (TAL) in March 2026, which temporarily halted crude oil flows to Austria, Germany, and the Czech Republic, relied on A/I Collective tools and services to claim responsibility for the attack and publish official communiqués, as well as to disseminate a manifesto calling for additional attacks on critical infrastructure and referring adherents to another A/I Collective platform, which provides detailed sabotage manuals as well as maps of vulnerable critical infrastructure across North America.

> The far-left extremist network responsible for multiple arson attacks and acts of sabotage against Germany’s rail and energy infrastructure between 2011 and 2026, including a January 2026 attack on Berlin’s power grid that cut power to 45,000 households and resulted in a fatality, relied on A/I Collective tools and services to claim responsibility for attacks, publish official communiques, and disseminate a manifesto calling for additional attacks on energy infrastructure worldwide.

Seems quite specific to me.


They also probably used Signal for comms, and it's safe to assume some of them used protonmail.

And maybe they used a Google camera for taking photos of their acts.

By the same logic, what prevents US from labelling Signal, Proton, and Google as ones supporting terrorist organizations?


Well, the far left extremists also used ISPs, mobile phones, energy infrastructure to do this.

Are those terrorist organizations too?

Moreover, ask yourself: why hasn't Germany accused A/I of being co-responsible of the attacks?


Is the car manufacturer responsible for their use of a car? The power tools manufacturer? The paper sheet manufacturer? The restaurant that fed them?

Technical intermediaries could be held responsible specifically for hosting the manifestos (under freedom of the press laws), but that requires an actual due process. Meanwhile, 100% of the authors and websites who inspired or relayed Anders Brejvik and other neo-fascist terrorists are still online and are not facing any kind of repression.


Here’s a thesis: the feminism of the 70s was a reaction specifically to the trends of the 60s. You can see this here in 2 places: the double standard and easy acceptance of male infidelity; and attitudes to divorce.

The trope of society praising male promiscuity is largely an invention of the 20th century.

Biblical and Hebrew texts all encourage punishment of both man and woman in instances of adultery.

As a man I don’t like promiscuous men because they’re competition. I don’t want to introduce my female friends to them because they’re not relationship material, and I don’t find them trustworthy as a wingman. Call it jealousy, but I think the same dynamic applies to the other gender.


On the other hand most societies are more tolerant of male promiscuity.

You accept its a feature of the modern west. It goes back to at least early modern times when at the very least it was routine for the aristocracy to keep mistresses. The Romans definitely expected fidelity from wives but not from husbands. Even today Middle Eastern cultures which criminalise female infidelity tolerate it from men. Greater tolerance of male than female infidelity is true of South Asia too.

I know the Christian attitude was that both were equally wrong, but that was, for most cultures, a radical change from the previous attitude and I doubt it was every completely accepted. I do not know what the Old Testament attitude was, but, again, I would ask whether it was really 100% accepted that punishmets should be equal.


> On the other hand most societies are more tolerant of male promiscuity.

Quotation needed.


How many concubines did Solomon have?

300 wives and 700 concubines. His taking of 'foreign wives' who turned him to idols was listed as one of the main reasons God was unhappy with him, indicating what the scribe thought about this.

that double standard exists for a rather obvious reason -- male infidelity biologically cannot result into their cuckolded partner unwittingly raising someone else's child, while female infidelity often results into that to this very day. the state will, in fact, force the victim to provide for the child if he doesn't immediately realize the child isn't his.

Private citizens and companies, as the gp said.

Just the abstract is the stuff of nightmares:

> Originally described as an isopod in the nineteenth century, and subsequently compared to various arthropod groups, it was re-described with limited illustration as a gigantic scorpion in the 1980s.

“We reopened Pandora’s box”

> Illustrating P. gigas with camera lucida drawings, light photography and tomographic data,

No why would you do this

> and assign several other specimens from the same formation to this taxon,

Nooo

> Several characters supporting a scorpion affinity are present in P. gigas, including large pedipalps with a fixed and movable finger, a stridulatory surface on one of the coxae, and an elongate subtriangular sternum morphology … Uniquely among scorpions, P. gigas has lateral epimera on the mesosomal tergites,

I don’t know what any of that means but it all sounds absolutely horrible

> and combined with the fluvial environment in which the fossils are preserved, we suggest that P. gigas may have been aquatic or amphibious.

OH GOD IT’S UNDER THE WATER

I WILL NEVER GO WILD SWIMMING AGAIN

And in the article the cherry on the cake: it was OR IS from Herefordshire. Where I will shortly be moving.


I wonder if lobsters give you the same feeling?

I was hoping it would just be a big crayfish but they specifically rule that out.

It lived ~400 million years ago.

Ah, a fellow thalassophobic. Yeah, no, I'm not going to swim in the vast ocean at night, thanks.

https://en.wikipedia.org/wiki/Thalassophobia


Which cultures are those?

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: