Hacker Newsnew | past | comments | ask | show | jobs | submit | chrisldgk's commentslogin

The difference here I think is that the company (Anthropic, not OpenAI) was specifically selecting rare books for training (and subsequently destroying). I think the value proposition is a much different one in this case. In your example the books are still available in high enough amounts that replacing it is cheaper. If the book is rare enough to not have been digitized for training data yet (which is why Anthropic is interested in it), keeping that book around is much more important.

If Anthropic subsequently released the digitized version of it we’d be having a different discussion, but for obvious reasons they’re not.


I have heard allegations that they might be destroying non replaceable books, but only as a hypothetical.

Replacibility is the standard tbat matters in this which I mentioned in my post. I haven't heard of a specific instance of an occurance of destruction of an irreplaceable book, but I'm willing to look into cases if you have a link.


> I have heard allegations that they might be destroying non replaceable books, but only as a hypothetical.

> Replacibility is the standard tbat matters in this which I mentioned in my post. I haven't heard of a specific instance of an occurance of destruction of an irreplaceable book, but I'm willing to look into cases if you have a link.

I think you can assume they're destroying rare or even irreplaceable books, unless they can show they've implemented careful processes to avoid doing that.

But given they're all basically SV startups, it's very unlikely they're doing anything except the minimum effort to get what they want now.


For me it’s just as unbelievable that the second hand book stores would sell rare or irreplaceable books at bulk prices. If they do they are really not good at what they are doing and I wouldn’t blame any startup big or small to destroy it, especially after preserving it digitally.


> If they do they are really not good at what they are doing and I wouldn’t blame any startup big or small to destroy it, especially after preserving it digitally.

1. That's fucking bizarre logic. "It's ok to do bad thing if someone didn't stop you."

2. These startups aren't "preserving" it digitally. The destroyed physical copies would have likely outlived whatever datasets they're creating.


Rare != valuable.

They are "betting" pennies in their books.

Rarity equating monetary value is a common logical fallacy often found with stamps and other collectable cultural artifacts (details would disgress too much here, see Pokemon cards or whatever).

And rarity vs quality ir utility is even harder.

But value is 1) not determined by scarcity and 2) neither by "unique information content" (because that's by definition not measurable without a panopticon) and 3) nor by "creative value" per se 4) it is indeterminate whether some detail in a work might prove useful later, in general. Even for scientific papers.

They just can bet on all fronts and for training material it's 100% rational to do it.

Apart from that, to briefly disgress into the discussion of "cultural value": many cultural works we hold as very important today were sold for cheap, if at all, when the author died.

Music scores are good examples.

But back to the broader "value"/rarity angle:

The historical value of some odd self-help or DIY instructions book from the 80s is hard to determine! It might happen to be able to provide a secondary source for something, or even act as primary source for some other fact of the time that is tangential to its subject.

None of the qualities that make trivial literature interesting to, for example, researchers, must correlate with market value (at the time of writing or at the time of research?).

Sure, there are truckloads of "unique" books from even 200 years ago that are neither culturally valuable nor sellable right now.

But these things can change quickly.

Imagine you want research details about old retail products.

Or the usual street price of something in some city at a specific time.

Or the inhabitant of a certain rentable space at a certain time.

Or some specific detail about replacing the parts of a certain type of car that's no longer around (is this valuable then?).

Humans create endless complexities, and monetary or evolutionary value is not coupled anymore to step-changes such as "how to light a fire using X,Y,Z".

We're deep down the rabbit-holes of our own ephemeral shaping of our world.

What word was commonly used to describe a sandwich in 1925 in Hungary.

Or what problems did people commonly encounter when repairing car doors in the 1940s in England?

Whatever. I can't make up a good example right now, but market value does not imply quality; and even quality (according to some cultural norm) is not universally measurable.


They did it due to some weird copyright argument about copying and destroying the original seeming like a transfer and not copyright violation, which is insane and apparently is mot always interpreted in said way. So they are literally taking a copyright gamble they might lose.


It is fair use to train on books you buy. The anthropic case established that. The destruction of books is just cutting the spine so they have individual pages to scan.


Destroying a book to argue fair use is not accepted as a fair use act is what I was saying, so they're needlessly destroying rare books.


They are not destroying any books to make any claim. They are doing it so the pages fit in the scanner.

Without an example of an actual rare book that is shown to have been destroyed, all there is is the is the claim they have the capability to do that.

The closest I have heard is the possibility of rare books purchased in bulk orders of cheap second hand books. This is a similar risk to melting down someone's wedding ring from a batch of scrap metal. It could theoretically happen, but the knowledge of the items presence is absent and it was not requested.

What other allegations have there been? I'm prepared to look at actual examples of you have them.


I think the important part of books is the knowledge in them. Not the binding or anything else physical. With some exceptions of course and sure there are things to be learned about physical attributes that don't translate to the digital version, but 95% is in the content.

> If Anthropic subsequently released the digitized version of it we’d be having a different discussion, but for obvious reasons they’re not.

I don't think you can scan books and just release a pdf. Someone correct me if I'm wrong. Even if you can, I'm sure it's a gray legal area.


It’s sluggish, it’s incredibly bloated with all the services that Docker (the company) wants to sell you and the UI is badly designed.

OrbStack is my go-to when it comes to running docker on Mac. It also uses macOS native containerization APIs so containers will be a bit more performant than they would (at least use to) be on Docker Desktop.

Another good approach if you’re fine without a GUI or want to bring your own is just ‚brew install docker‘.


The main gotcha I encountered with orbstack was with pulling down internally built amd64 images and trying to run them on apple silicon. I just remember constantly running into issues and going back to docker desktop as a result, even if their solution is janky in its own way.


I'll check it out. Yes, no UI needed. I run Docker Desktop but I don't think I've used the UI more than on ce.


I feel like people are focusing on the wrong thing here.

He starts off at 2027, where he’ll have a headband that can read his thoughts. That’s in 4-16 months. Correct me if I’m wrong, but last I checked we‘re quite far from being able to read people’s thought from a headband.

Do these clowns actually forget that part?


> But non-invasive methods are getting much better. The hardware is improving and getting cheaper, though I apologize for being vague about the particularities of our hardware.

The statement implies that the company has some secret hardware that can, or at least, is close to doing that. That would blow my mind, but the author seems to believe it.

As an aside, I believe the author is a she.


> Do these clowns actually forget that part?

No, they are aware. You and I are not the audience for this post. People with capital to invest are the intended audience.


To be fair, that is a thing mostly done by the political elites and not an opinion that most citizens (that I’ve interacted with) hold. Most of us would much prefer all of the money that’s being put into weapons for Israel to be invested in our country instead.

Then again, there’s currently a big political shift to the right and the US/UK flavor of identity politics seems to have made its way over here, so that’s probably also not changing anytime soon.


I remember at the ad agency I used to work at (somewhere around 2018) we used to have a happy hour every Friday where from 5pm onwards all drinks (including booze) in the cafeteria were free, which led to people of all departments including C-Suite to get together and drink. I remember it as a good time though, I don’t remember anything ever happening that led to people having repercussions for employees. Then again, this is German drinking culture and the stiffest drink you could have was wine, so people wouldn’t usually get wasted, just tipsy.

The Christmas parties were a different breed though. HR definitely had a lot of work to do after those.


When I was still in my early 20s, at my first office job, drinks after hours (beer, wine) were free from the company kitchen.

Only took part once (because I was new, first week and all), saw how openly employees spoke to the directors (luckily nothing came of it), and never again in my life ever consumed alcohol in front of a boss.

As GP said, it's a full on soft power thing - always keep your wits about you when in the presence of colleagues and bosses.

There is no safe space at work; you can be fired!


Japan here and there is a word for it "nominication" for nomu (to drink) and communication. Instead of Christmas party we have a year end party (bonenkai). Otherwise reads the same. It's common and many people actually like it. There can be negatives, like the pressure to drink too much. But I never heard of anyone getting fired in all my social circle.


Outside of Japan, the idea of bureikō is not necessarily similarly part of the culture.

https://en.wikipedia.org/wiki/Nomikai#Nomikai_etiquette


I've been to a work Christmas party where one of my coworkers got so drunk that she couldn't talk and kept grabbing my arm and walking with me. I don't drink at all, so it was a bit surreal, and, compared to Greeks, Brits are on just a whole other level on the normalisation of alcohol consumption.

Obviously there were no repercussions, and when I mentioned something the next day, I got "you're not supposed to talk about that stuff in the office", which was interesting context for me. The drunk colleague and I went on to become really good friends.


Here in Canada I worked at a company that had free alcohol on Fridays, and it got to the point where people would invite their friends - non-employees - to come and drink. They'd all get so blackout drunk they'd pass out in the lounge and the customer service staff would have to climb over them on Saturday morning to get a coffee.

That particular perk was removed after a couple of months.


That’s because he is. None of these people are your friends and all of them will fuck you over if it means getting richer and more powerful.


At this point, why not just write the code yourself? Defining exactly what the product is supposed to do is the hard part, writing code is the easy part. Write your specs as code and you have your product - why let your LLM do the fun part?


Just because I am capable of "writing all that code", doesn't make the option preferable to defining a vast majority of spec up front and having an LLM generate an implementation. I am already going to spend the brain power on reviewing the code. I am already going to spend the brain power on pontificating edge cases, external module interactions, and next steps. Why not fast forward to that point and save 80% of the time (and brain power/attention/motivation to boot)?


If you can define the spec up front, this is probably true.

For anything large, the spec becomes increasingly more complicated. Look at software schedules in the old waterfall days of the 80s/90s: the spec / planning period was maybe 30-70% of the project.

Unless you’re working on pretty routine stuff, the real problem is that the customer (which might be you) almost never knows what they want. The spec will change the minute a customer gets something to play with.

This was the real value of agile in my mind: letting a customer change their mind as early as possible.


> I am already going to spend the brain power on reviewing the code.

Very few devs are actually reviewing any generated code.

> Why not fast forward to that point and save 80% of the time

If you are saving 80% of time, you aren't actually reviewing the code.


> Very few devs are actually reviewing any generated code.

Just because very few devs are qualified at doing their fucking job, it doesn't make someone trying to use AI properly wrong.

> If you are saving 80% of time, you aren't actually reviewing the code.

The idea is that if you spend time in specification ahead of time, reviewing and validating will be easier and less time consuming later.

I haven't tried it myself, but the idea rings true to me.


I've never seen a spec survive first contact with implementation. The spec is refined while writing the code.

Hell, you probably couldn't even build a simple bike shed from plans without having to revise them while building, so I am skeptical that without writing you are going to pinpoint the problems in the spec.

Reading only gets you a short way towards learning.


> I've never seen a spec survive first contact with implementation. The spec is refined while writing the code.

Neither have I. This does not make the spec useless. I don't spec hoping that it will be the source of truth, I spec because planning more often than not allows me to spot inconsistencies and ambiguity ahead of time, not halfway through implementation.

> Hell, you probably couldn't even build a simple bike shed from plans without having to revise them while building, so I am skeptical that without writing you are going to pinpoint the problems in the spec.

I think you are using specification and design wrong.

It's not supposed to be a bible that implementation can't deviate from. It's a plan, not law. It's okay for the plan to be adjusted in contact with reality.

It's still useful to know ahead of time constraints, expected output, assumptions, premises, etc.


> It's not supposed to be a bible that implementation can't deviate from. It's a plan, not law. It's okay for the plan to be adjusted in contact with reality.

My point is that without writing, you can't surface the type of problems you usually surface. The AI isn't going to surface those problems for you.

It's rare when reviewing that you think "Oh shit, this approach is totally wrong, we need to throw it away", while it's common when writing code to have that reaction.

If you aren't writing, you aren't having that reaction, and you aren't going to get it from reviewing code that has thousands of green "passed" lines in the testsuite.


> It's rare when reviewing that you think "Oh shit, this approach is totally wrong, we need to throw it away", while it's common when writing code to have that reaction.

That's not my experience. It is actually very common for specifications and design to be reviewed and improved.


> It is actually very common for specifications and design to be reviewed and improved.

I think there may be crossed wires here - specs and designs are reviewed, but I've never seen a code review result in a spec+design review, while I always see spec+desiogn review happen during the "writing code" phase.

In short, reading code does not result in a spec+design review, writing code does. If you are not writing code and only reading it is unlikely you will trigger a spec+design review.


I use a multimodal approach to defining my spec: different layers of criteria for how the software looks, behaves, what it produces, and under what constraints.

For the literal code:

• A healthy cocktail of /WX + /Wall, plus clang-tidy with very few suppressions

• An extremely opinionated mix of clang-format and LLM-generated bespoke formatting that AST-based tools cant express

• Hungarian notation; all stack locals pre-hoisted, declared in order of appearance, and separated from subsequent assignments

• Enforced dataflow: all memory accesses are bounded independent of branch resolution, with only data-oblivious indexing

• Functions have a single point of return

In a C89 workflow, this pushes agents to produce code where wrong business/domain decisions are unmistakably obvious, while eliminating the vast majority of bug classes before I ever read it.

So yeah, Ill reassert 80%, if not more.


i spent an inordinate amount of time thinking about this comment yesterday. i’ll cut to the crux of it in a second, but wanted to preface what i’m going to say by clearly stating that i don’t intend for this to be snarky or putting anyone down. that’s not my intent. ready? here we go…

is it possible that you might be in a job that’s not right for you?

it sounds like you want to cut out 80% of your job. if you want to cut out 80% of your job, maybe you’re doing the wrong job? y’know?

like, i read this comment and my mind goes to project/product manager who has real experience of coding. going from a spec (tickets / design docs / customer feedback notes / epics / stories / whatever) to a working implementation (team of engineers build it and you don’t have to use brain power). it sounds like you’re describing turning your job into the job of a PM.

we need more PMs like that in the industry. good PMs are few and far between, good PMs who know what’s it’s like to code — even fewer. so maybe have a think about why you’re working this way? i don’t really care if you do or don’t, but future you might be glad if you took some to think about it.


I'm with you all the way here. I derive zero pleasure from simply typing out the code once the spec is clear. Having a fast forward button to skip that phase is a pure win in my book.


I do get pleasure from typing out the code in some languages (and not in others; hello javascript, java!). Similarly, I love writing text with a calligraphy or fountain pen. However, I can't dedicate too of the much work / business time to whatever is more pleasurable.

So, I "doodle" some text / ideas / planning with a calligraphy pen, and type in some code, occasionally, both mainly for the fun aspects. There are side benefits to both, too. Writing some plans slowly and "beautifully" drags them out and I get to think longer on them, so the sporadic "nice looking plans" are often more well thought. And doing the coding all by myself stops my brain from losing the ability. I was initially in the 100% AI-writes-all-code camp for a while and noticed I am getting notably slow in some personal coding skills. It is too early to treat specs as the new code and old languages as assembly (but I admit we might get there some day).

In other words, I think AI doing 90-99% of the coding, depending on the language verbosity and AI accuracy for the code at hand, is quite reasonable.


Personally, this is an experience I thought about first before writing my comment. I think in the days pre-AI coding assists, I believe you describe the intrinsically human experience that's requisite to write code by hand. The wonder, the joy, the frustration, the confusion, the elation--the discovery. These days, the things I wonder about lie deeper and deeper behind more and more lines of code, through journey's that provide less and less joy, and thusly becoming more and more unreachable as I'm human, bound by an excess of things in addition to time. AI has helped me rediscover some of this sporadic creativity demonstrably due its ability to prototype recreational ideas on a whim

Professionally, I'm employed writing safety-critical avionics software. Superfluous amounts of cogent tooling putting guardrails on agents has enabled me to spend heaps more time to think deeply about how the software should work at a systemic level. The code by definition must be heavily criticized and battle-tested before it can go out the door to begin with. Albeit a beautiful part of coding, those sporadic bursts of creativity drive the code leaving my desk less and less, and I feel strongly that has made its quality paradoxically better since I'd spent much more time on broader implications and interactions.


100% the opposite here. I derive all the pleasure from writing code, which is why I'm still writing code.


a spec can be wrong until you prove it is right..


Not a developer would come to mind.


I do this because I'm wagering that LLMs will keep getting better. I'm wagering that specs will maintain value while code will degrade in value (become commoditized).

Code lacks the surrounding theory that situates the code in the world [1]. My specs contain the theory that the code lacks, which makes specs more valuable in the future. Specs are proprietary data. Data holds value in a post-AGI world, not code.

I am defining specs to be more than just an architectural spec, to me it's more like I'm writing a booklet about a subject, and I'm using it to teach the LLM via in-context learning. It might need a different word than "specs".

[1] https://pages.cs.wisc.edu/~remzi/Naur.pdf


Piggybacking here; I'd describe it like a fish ladder. Instead of "teach" I'd say "orient." LLMs are a force whose magnitude is undeniable and increasing, but it's up to us humans to provide the theory that exerts the magnetic forces to naturally encourage them in the right directions.


> I’m wagering

So isn’t that gambling, not engineering?


It's a reductive inverse corollary, but highly skilled Blackjack players are known to hesitate hitting on 18


>> Defining exactly what the product is supposed to do is the hard part, writing code is the easy part.

There is a massive difference between a spec, which defines what the product should do, and code, which defines exactly how it should do it. Moving from the former to the latter is not "the easy part". Anyone who genuinely believes that either works on easy and straightforward problems, or is some sort of programming god. Because translating specs to code can still be difficult and exhausting.


> Defining exactly what the product is supposed to do is the hard part, writing code is the easy part.

> There is a massive difference between a spec, which defines what the product should do, and code, which defines exactly how it should do it.

He states: The difficult part is figuring out the details so LLM doesn't save much time. You state: If LLM is able to correctly assume the details that saves you a lot of time.

Case 1: Part of the spec describe some basic feature based on a popular framework and industry standards, everything is trivial. You are right, he is wrong.

Case 2: Part of the spec describe some niche feature and/or uses some not popular framework and/or require deviation from industry standards and/or cutting edge performance/latency requirements and/or uses a bunch of proprietary non-googlable data. You are wrong, he is right.

The more senior engineer are the less time they spend on case 1, those are easy, they don't spend much time on it, it is the 2nd which is much more time consuming.


One thing that I can’t seem to parse from the article is why the researchers assume that this is an unresearched part of ADHD and not a different disorder entirely. I’m sure they have their reasons, but I don’t think it’s written in the article.

To me it seems that if it’s not „treatable“ the same way ADHD is, I’m not sure if it’s useful to categorize it as such. On the other hand, I’m happy if kids with this disorder can get a diagnosis and treatment that actually helps them sometime in the future due to this research.


You have a set of diagnosis criteria, and matching those criteria gets you the ADHD diagnosis. This study takes people who fit the diagnosis, and says there's a test you can do to split those people into three groups.

But yes, once they have a better understanding of what that difference means, the next step might well be to split the ADHD diagnosis into two separate disorders, or even that, like cancer, ADHD is actually a whole range of separate but related conditions.


Diabetes has a similar issue, with type 1 and type 2 having very different causes and pathologies.


> why the researchers assume that this is an unresearched part of ADHD and not a different disorder entirely

Whether or not the extreme dysregulation is a different disorder in its own right or not, is not relevant here. They are grouping ADHD matches; clinically recognized ADHD presentations plus MRI recognized ADHD which have a distinct brain sub-pattern occurring in people that have the same distinct behavioral traits. ADHD frequently has co-occurring conditions.

"Identifying “specific subtypes” of ADHD will make it easier to treat these children effectively". Having a more objective way to diagnosis for things like that seems to be the focus of the approach. They expect it to keep evolving, so I wouldn't say they are assuming anything about absolute labels -- just grouping what they now know to be true, that certain external traits match certain distinct brain patterns that are within the larger adhd brain structure.

Also, I think it's not that it is "not treatable" as ADHD, it's that ADHD can be treated in many different ways and currently the wide variety of responses to such is still a black box. Adderall instant release could briefly make me tired, I would sometimes break off a small piece and use it as a sleep aid. Some other `treatments` (I prefer societal alignment coping aid) resulted in what seemed like an expensive joke. Subtypes may eventually be able to show which options work best for which types and to start there first, instead of the current default iteration.

This link adds more about their research. https://medicalxpress.com/news/2026-03-distinct-adhd-biotype...


It is a good point and I also struggled with that bit somewhat. It is different in so many ways, have different symptoms, does not respond (as well) to the same medication, and affect different parts of the brain. The jump from there to "subtype" was not too logical for me ...


I would suppose it interrupts the page load after streaming the HTML and before loading and/or executing the cookie banner‘s javascript, meaning the content is there but the cookie banner will never open.


Per their own docs, D1 is primarily meant for things like Auth DBs that you have frequent read/write access to but that store limited amounts of data. If you need more storage, running Postgres somewhere else and querying via Hyperdrive is probably what you want to do instead.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: