Hacker Newsnew | past | comments | ask | show | jobs | submit | uh_uh's commentslogin

There are trade-offs here:

Give up too early -> users will get annoyed because the task would have been solvable if the model pushed harder.

Give up too late -> collateral damage while completing the task A.K.A. misalignment.


Asking for the user input isn't giving up


It is comical at this point. Some people just can not stand the thought of AI actually delivering and are trying to find whatever ways to discredit it.


This isn't really about delivering - it's more about helping to understand the shape of problems that AI can solve right now. If they took 1000 problems and threw the model at it and it solved these ten, is there something we learn about these ten problems and the kinds of things that current AI is good at? That's very different from picking ten problems _at random_ and solving all of them successfully, which would suggest a much less bumpy capability surface. It's interesting and it would be good science to release it.


That's totally disjointed from anything in this thread. The main accusation is that openai is cherrypicking math problems and we should be against these results. As if a mathematical proof stops being provably correct because it was cherry picked

And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.


It's disjointed?

The post that started this sub-thread asked:

> 1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?

I think it's an extremely relevant question to ask, because it helps us better understand the current state of AI being able to handle math, for exactly the reasons I outlined. I was arguing against the idea this is just a reactionary anti-AI kind of question to ask. It's not! You can be very impressed by what AI is capable of in math (I am) and still think those are really interesting things for OpenAI to disclose (I do).

OpenAI specifically called out a $2000 per problem average, which implies something that's probably not true ("if you throw $2k at us we'll solve an open problem for you"). It would be cool to know what the actual number is.


It just feels silly to haggle about the price here. It doesn't even matter because it's going to drop by an OOM quickly.

If these 10 problems were solved by humans, it would be pretty impressive, even if it took a large number of researchers! Yet when AI does it, HN commenters suddenly feel the urge to play accountant.


The interesting question is "Which problems are LLMs good at solving, and which problems are LLMs bad at solving?", which could also be restated as "Which problems are cheap for an LLM to solve?". So cost is relevant here.


> In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it).

But that's the start of math research, not the end.

The point is to get practice and experience doing research.

Did ChatGPT learn anything from these proofs, that it can build on?

Part of what's annoying people is that ChatGPT is churning though problems that are meant to be motivating. They are problems that aren't worth the effort of human professionals (usually because they are incredibly computation-hevy, so better suited for a computer than a human), so they are good for students to work on.


It's not about discrediting AI. We know LLM is a commodity technology like electricity at this point. If somebody in 1900 claimed they had a setup at home where they feed in electricity and cool air comes out the other end (meaning they invented AC), obviously people would want to know what the setup is, so everybody can have AC.


Sure, but I don't really understand what the argument is to _not_ be transparent about methodology, since if the models are so powerful, then doing so would easily support the claims and put these concerns to rest. People are right to be skeptical given what is being implied and the orientation of the narrative

I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."

By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?

I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?

To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations

I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though


Entirely comical too that some people can not stand the thought of people poking very big comulent holes in the claims of AI delivering what it claims to deliver. As if having skepticism is some how a way to discredit a person.


Yes, but if their machine god really is as good as they say it is, why are they constantly resorting to statistical sleight of hand at best and outright lies at worst with every public statement?

That's not normally how people act when they're confident in their product


AFAIK there is no prophylactic for pork tapeworms. I'd love to be proven wrong.



These dewormers treat the condition after the individual was already infected. I don't think they prevent the disease. Additionally, taking Albendazole without anti-inflammatory meds can be fatal in case of an active infection. Reason is that if the cysts start dying the brain, the toxins released might cause swelling which can lead to neurological damage or death.


What kind of apps are you writing where GC spikes matter?


trading, networking, gaming, ai, realtime, almost anything with hard response requirements.


> It's what having such senses feels like from the inside; the first-person view.

The hard problem is that there is such a feeling at all.


It's not hard at all when you acknowledge that such senses exist in the world, and that you (like others) possess them. As an aside it tends to foster a certain tendency towards empathy.

In essence, you're asking why there's an inside to being a self-modeling system. But "inside" isn't something extraneous, something additional -- rather, it's what "self-modeling" means.

Really the "hard problem" has a very easy answer, but it's a physical/functional answer, and dualists and obscurantists simply don't like it.


It's embarrassingly silly to say but I've frequently just boiled down the hard question to the question of "where is the experience of the color blue stored in the universe?" Even as a non-dualist, I still haven't found much of an answer that I like. I'm all ears if you've got a book recommendation.


The question presupposes that "the experience of the color blue" is a discrete object that needs a storage location. But that's the dualist picture in disguise. On a functionalist view, blueness isn't stored; it's what certain neural activity constitutively is when you're that system observing that blue.

As an aside, isn't it more weird that violet and purple look indistinguishable despite being physically so different? It's said that this is because our L-cones (red-sensitive) have a secondary sensitivity peak at short wavelengths. So violet light triggers S-cones + a bit of L-cone. Purple light (red + blue) also triggers S-cones + L-cones. Similar activation pattern = same quale. It's all functional/physical.

Read Tom Cuda "Against Neural Chauvinism." Also Daniel Dennett.


> On a functionalist view, blueness isn't stored; it's what certain neural activity constitutively is when you're that system observing that blue.

Why should there be anything a certain neural activity is when making an observation? This is adding something additional to functionalism. You're just sneaking the hard problem back into the picture without realizing it.


What is mysterious to me is why and how chemical reactions in a certain part of my brain create an experience of blue.

Yes some chemical change happened there, but so what.

These are not very unusual chemical reactions. They happen and are happening everywhere. Does all the chemical reactions going on generate an experience to some experiencer?


I think the flaw in your reasoning is the assumption that chemical reaction is causing the sensation of blue.

But imagine if the consciousness and what it senses cannot be separated. So the consciousness sensing blue and the chemical reaction happening in the brain, are just correlated. One did not cause the other.

One can ask where that correlation came from. I think that the such correlations are inherent in such worlds where consciousness is possible.

I think everything that we observe as physical laws, causality etc, are just such correlations.


That is an interesting thought.


This is where these questions take me. Since the experience is the only thing I can be certain of, I'm less drawn to "everything is physical" answers and more drawn to ideas from phenomenology and Bishop George Berkeley. And since I'm not super religious, I'm not really comfortable with those "answers" either.


>where is the experience of the color blue stored in the universe?

It is not stored anywhere. It is part of the consciousness that experience it. In other words consciousness comes bundled with everything it will ever feel.


So you say that the hard problem of consciousness is explained by the fact that we appear to be conscious?


The kneejerk response would be: Are you not conscious at this present moment? If we were to modulate your spatiotemporal senses with drugs or a lobotomy, do you doubt that you would be very differently conscious, or perhaps entirely unconscious?

I mean, there is a credible first-person answer to that question of yours, which each man can answer for himself.

But considered more seriously, the "hard problem" is an artifact of treating experience as a separate thing that needs to be generated. If you accept that self-modeling systems bounded in space and time exist, you've already accepted that experience exists -- because experience is what such a system is, from the inside. There's no second step where experience gets added. The question "why is there experience?" is exactly akin to "Why is there an interior to four walls and a roof?" The interior isn't a separate thing; it's necessarily constitutive.


> because experience is what such a system is, from the inside.

There being an inside to self-modelling systems bound in space and time is the hard problem.

> The question "why is there experience?" is exactly akin to "Why is there an interior to four walls and a roof?" The interior isn't a separate thing; it's necessarily constitutive.

That's given from three dimensions of space. This is not the case with subjective experience. Functional and physical terms don't have an inside where experience lives. It's what makes the p-zombie argument potent.

Let's put this another way. Functional terms are abstracted from experience to model the world. See Nagel's What It's Like to Be Bat paper on science being a view from nowhere, which is really about the fundamental objective/subjective split. Or Locke's primary and secondary qualities.

You can't get experience out of abstract terms. Experience doesn't live inside abstract concepts. We can model the world with them, but experience was left out at the start.


>You can't get experience out of abstract terms.

Would you agree that you are conscious at this point?

Would you agree that there are some set of physical laws, an initial state, and a set of random events to the universe that we inhabit?

Would you agree if we simulate this initial state on a computer, and step through it using the set of physical laws, and the random events, we will see the eventual emergence "you", who we know is conscious?

So are you saying that the entity inside the simulation is a zombie who is not actually conscious?


> Would you agree that you are conscious at this point?

Of course, I'm having a conscious experience replying four days late.

> Would you agree that there are some set of physical laws, an initial state, and a set of random events to the universe that we inhabit?

We inhabit a universe modelled by laws physicists have arrived at to describe observed behavior. That's as far as I'm willing to go ontologically.

> Would you agree if we simulate this initial state on a computer, and step through it using the set of physical laws, and the random events, we will see the eventual emergence "you", who we know is conscious?

No, I don't think computation is conscious. It's abstract symbol processing.

> So are you saying that the entity inside the simulation is a zombie who is not actually conscious?

Yes, it wouldn't be me. I don't think simulating the world is the same thing as the world itself, despite all the science fiction stories to the contrary.


I'm not a dualist or anything. I'm in the "it's weird and I have no idea what the answer is" camp. And yes, I've read Dennett. I'm trying to understand your views. Lots of questions follow, but don't feel like I'm barraging you unnecessarily. Just trying to figure out your view with what seem to me like interesting questions that I myself can't really answer.

I'm using "consciousness", "subjective experiences", "senses" and "qualia" as synonyms here, but if you see a difference, please mention it. Obviously "consciousness" has many definitions that have nothing to do with the "hard problem of consciousness", so I'm using it in this sense here. I'll use "qualia" as it's the word that relates most to the hard problem of consciousness. You can substitute it with "sense"/"senses" if you like.

1. Do you view qualia as an emergent property? Of what exactly? What is a self-modeling system? Is a human one? Where would the boundaries be; would they even be defined? The human body or the brain only or the nervous system? Or whatever neurons activate when a certain thing happens, like seeing blue or feeling pain? What about animals - pigs, dogs, rats, snails, ants, bacteria? What about AI, current and theoretical?

2. Could there be a set of minimal self-modelling systems in some abstract space that are the boundary of what has qualia and what doesn't? Like, these 1000000 neurons arranged like that qualify, but if you take 1 out, they don't? Or is it a fuzzy boundary somehow?

3. What kind of statements could be made about the qualia of yourself and of others? Not sure what kind of answer I'm looking for, but how objective or truthful would those statements be? Maybe "qualia is nothing really, we only have the set of equations that govern physics and everything else is an abstraction"? Like an apple isn't anything really, it's just a badly defined set of atoms and energy. There is no "apple" or "chair". Or is it something else?

4. What are your views on meta-ethics and ethics in general? Should we care about it at all?


How do you know they (and others) possess them?


Isn't there a magical moment needed still when a single qubit "touches" the rest of the universe?


It touches you, and you are just as quantum as the bit.

So two entangled versions of you follow, one entangled with each state. (Actually as many quantum versions of you that touched the qubit times two.)

Which is what happens, as we know from experiment when any one qubit interacts with another independent qubit. We get the product of entangled states, each now correlated. But different entangles states are now in superpostion with each other.

So correlation/entanglement happens and is experienced, despite no collapse of superposition. No information was destroyed or created.

Each of you thinks, wow now the qubit only has one state. But that is because there are two versions of you, correlated respectively with the two uncollapsed qubit states.

Complete conservation. That is the "experience" of collapse that needs no explanation, because it is a predicted experience not requiring an actual collapse. Just as spherical Earth models don't need a special explanation for the appearance of locally flat Earth, because spherical models predict a local flat Earth experience.


I'd say we are confused about both the lowest (quantum) and highest level (consciousness) phenomena of the known Universe. Quite humbling.


We have a theory whose plain reading matches experiment at all scales.

Consciousness is something else. It is tempting for humans to pair mysteries up, pyramids and aliens, or whatever. But there isn't any factual basis for linking the experience of self-awareness with quantum mechanics.

Is there a factual reason we know digital minds couldn't be conscious? Where quantum effects have been isolated from the operations of mental activity. That seems like a premature constraint to assume.


I wasn't trying to link the two. Just pointed out that there seems to be a lot of unknowns on the map.


Would you be similarly pedantic if a high-schooler did the same?


Yes. Someone making one contribution among many to a paper clearly does not deserve anything like sole authorship credit of the entire paper, which is what the title from OpenAI implies to me. I don't believe I'm being pedantic at all. And, by the way, high schoolers or college students make co-author-level contributions to real papers quite frequently in the US at least (I was one of them).

The text of the post is much more honest. The title is where the dishonesty is.


Hi, I'm an author on the paper. It was definitely a human-AI collaboration, but it is also true that the final simplified formula, Eq. 39 in the paper (which is what we had been seeking, without success), was conjectured and proved by GPT. So it derived a new result in theoretical physics. I'm genuinely puzzled by your complaint.


OK, but don't you see where this is going? The trajectory that we're on?


How so?



Its kind of a suck up that more or less confirms the beef stories that were floating around this past week.

In case you missed it. For example:

Nvidia's $100 billion OpenAI deal has seemingly vanished - Ars Technica

https://arstechnica.com/information-technology/2026/02/five-...

Specifically this paragraph is what I find hilarious.

> According to the report, the issue became apparent in OpenAI’s Codex, an AI code-generation tool. OpenAI staff reportedly attributed some of Codex’s performance limitations to Nvidia’s GPU-based hardware.


There was never a $100 billion deal. Only a letter of intent which doesn't mean anything contractually.


> OpenAI staff reportedly attributed some of Codex’s performance limitations to Nvidia’s GPU-based hardware.

They should design their own hardware, then. Somehow the other companies seem to be able to produce fast-enough models.


> They should design their own hardware

They made a deal with Cerebras for fast inference.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: