How many tokens/second you getting? I have the same CPU but a 5070 and only 64 GB of ram. I just got llama.cpp built and am now hitting a whopping 5-6 tokens/s.
DX11 is much more high level while DX12 implements direct control over the GPU. 12 is much more like vulkan.
I guess since this is emulating the device instead of just passing the commands to vulkan it is easier to do DX11, but I could be wrong about this particular implementation.
DirectX 12 isn't just much more like Vulkan, it's so shockingly similar, it's basically Vulkan. It's jarring when you read it for the first time because you ask yourself why it even exists, but a second later, you realize it's so Microsoft can own their own graphics interface, too.
I could be remembering this incorrectly, but that's the impression I remember.
DX12 was over 2 years old by the time Vulkan was released. People in the industry also use DirectX extensively and have for decades (Vulkan is fairly marginal in comparison); many modern graphics API features (mesh shaders, ray tracing) were designed and developed in DirectX, as Microsoft worked with closely with ISVs on these features before they spread to wider audiences and published specifications.
Honestly if you are creating games, with the rise of Proton and vkd3d, there is seemingly less need to target Vulkan today than ever before. Just target x86_64 Windows with DirectX11 or DirectX12 and you will even get Linux support (and some day, probably macOS) "for free". That's just for games, though.
Nope, thanks to AMD giving Mantle to Khronos, Vulkan actually happened, otherwise there would be no Vulkan now, they would still be arguing about OpenGL vNext on those committee meetings.
And only exists in first place because AMD offered Mantle to Khronos, otherwise they would probably be discussing how OpenGL vNext would look like to this day.
While flohofwoe certainly knows this, for the others,
"Khronos is free to take as many pages as it wants out of the Mantle playbook, and AMD will impose no restrictions, nor will it charge any licensing fees."
I run it on my milk-v duo s risk sbc. The risk-v version is maintained by a few dedicated people, and the milk-v specifics I had to fix myself. But it works :)
I knew there's an unofficial version for ARM. I meant more that the vast majority of AUR packages target x64 only because that's what 'official' Arch supports.
That being said, didn't know there's a RISC V effort as well now, so TIL.
Articles regarding my own harness or how I set up llama.cpp etc?
I really just iterated over the harness over and over for about two weeks with opencode until I was sort of satisfied (still lots to do there :).
For the llama.cpp I asked claude fable to optimize it for my hardware and iterated a few times. In the end I landed on the following: https://pastebin.com/2PpJFUC0
I half expect Nvidia to have buyback contacts like Ferrari with the larger customers to prevent a price crash when they all upgrade and to keep them scarce.
I hope not though, perhaps I can pick up a H100 in a few years if they get sold on the open market.
I worked at an org that had a substantial on-prem GPU datacenter. We transitioned to <Big Cloud Provider> with a substantial negotiated discount rate, with part of the contract being we would sell them all of our hardware and not purchase any more.
If you're planning to run in clouds, committing to not buy hardware (during the contract term, presumably) isn't a big imposition. Maybe you switch to a different cloud, and you wouldn't buy hardware for that.
If you want to switch back to on prem, there's probably a way to structure acquiring hardware so it doesn't break the contract. Maybe you lease it, maybe the purchase happens through a related company, maybe there was no way for the contracted cloud to find out...
I don't know if this was IT shrugging me off or if it was something real but some IT person at this big ISP I worked at told me that they cannot just buy an SSD — my windows box at work was running off of a hard disk in 2019 — and that there was some contract that said any computer hardware we bought had to be through HP or something like that and it takes many months it something like that.
> I don't know if this was IT shrugging me off or if it was something real but some IT person at this big ISP I worked at told me that they cannot just buy an SSD — my windows box at work was running off of a hard disk in 2019 — and that there was some contract that said any computer hardware we bought had to be through HP or something like that and it takes many months it something like that.
Plausible - the hardware might be leased, and so you have no right to modify it. You have to pay them to modify it.
Same as if you leased a car, you cannot do the services yourself, you have to pay them (and an approved agent of theirs) to do it.
We could only book our business travel through American Express. It must have made economic sense at the c suite level, I wasn't seeing it at my level, the flights were consistently more expensive.
Was it a HP Finance thing? Because there are a shit ton of people who can transact through them now. We had HP finance, and it became a game of sorts to see what crazy shit we could get with it. Theres a local mob who sell computer parts who were happy to use HP finance. That said it definitely wasnt everyone and if you didnt do the legwork you could definitely be trapped.
The usual very short term corporate thinking that maximizes quarter profits while bankrupting the company in the long term.
IMO a very shortsighted decision, if not downright stupid.
What is short sighted is parroting this over and over. It's not insightful, interesting or adding anything else to the discussion than "mean people are bad".
It would be very hard to find current examples of companies doing that harmful excessive focus on the long term. Only Valve comes to mind and they are anything but bankrupt.
Quarter profits maximization will create those loops of doing and undoing the same thing.
If you check the successful companies (Apple, etc.) their activities are mostly year cycles.
Where folks (often) get lazy is the resulting math over what the real bean-counters care about (but are too lazy to check often).
In a past life, I worked on costing models for a Cable/Fiber contract house, to help the company decide 'what was profitable to keep in house' versus 'what do we subcontract' (sometimes that could even mean we just 'rented' a machine and had a qualified operator using it, based on that employee's hourly rate and expected L2R for taxes... so many spreadsheets...)
And from from my 'I don't know all the factors for this but I've seen how people screw up the big ones' view
(and frankly, I'm guessing a lot of us have seen and dealt with the same category of 'bad math' around outsourcing IT work...)
An on-prem data center means:
- You need to account for electricity costs
- i.e. CA vs midwest electric rates.
- cooling and power backup capability
- Smaller factor but real
- personnel cost
- e.x. there's probably cases where a smaller org could be better off with 'on-site' server admins that have other roles based on local wages. Kinda case specfic but it's a case.
- whatever the 'space' holding the stuff costs
- Sardonic take :Hey, let's have another unused meeting room instead! (e.x. In the case of on-prem shops that simply fled to AWS in their migration from VMware)
- the cost of licensing whatever is running
- In defense of this, In one of my earliest IT lives, AWS handling the Oracle licensing for a DB was a *huge* win as far as making it as easy as possible to ensure whatever was going on we couldn't have the Oracle licensing folks 'ding' us on whatever infraction occurred between reviews (that could not be understood by the majority of the company, often including the accused. I was never guilty but I saw it happen to others.)
- OTOH I know lots of folks who just want to be lazy about what they have to document.
Still, IMO a lot of orgs don't do the right math around these decisions, or just buy into the 'Well trends can change' as though they can decide as an org they need to suddenly triple capacity in a month and it would be able to organically happen in the first place.
Frankly, the orgs that 'might' need that either have their arch set up where they are in cloud, or they are onprem but can scale to cloud if needed in interim.
The only reason a datacenter would ditch their H100 is if it becomes uneconomical to run them, with newer silicon providing much more power efficiency. When that happens they'll look like a used V100 looks today: horribly inefficient, lacking modern data types and engines, requiring screaming server fans with weird adapters to not melt, way beyond end of life in terms of cuda support. Almost completely damn useless unless you really have no other alternative.
Honestly, the primary reason you would ditch a H100 is because the ML models have moved onto a different data type, e.g. nobody uses fp8 anymore and everyone jumped onto fp4 or some custom block float thing.
There's going to be a golden age of GPGPU compute in the next few years once A100/H100 are fully obsolete for running frontier models efficiently and the price plummets
It will be perfect for stuff like GPU-accelerated query engines, "classical ML" and every other CPU-based workload that could conceivably be offloaded to GPU
You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.
I've looked at this some but I already have a GTX1070 which is only supported upto CUDA 11.9.
That's precludes some interesting modern optimizations out of the box. I've spend a lot of LLM tokens backporting some things, but I'm really not sure the hassle is worth it.
New hardware is just better. I think in maybe 5 years when supply and demand are back in equilibrium we are going to have some killer technology for decent prices, and 15yo H100s won't look attractive.
> You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.
In a recent Gamer Nexus video with Level 1 Tech, they mention V100s are also quite useful for many applications that use FP64: stuff four in a workstation, and many PhD candidates would be quite happy with the throughput they can get for certain scenarios.
Linux hackers will be finding all kinds of crazy uses for hardware that now costs $100,000 and in 5 years will be available as scrap.
This is assuming that there is no big, big disruption to the semiconductor industry (e.g. TSMC getting attacked), in which case... well, I am gonna treat each stick of RAM I currently have like it's a faberge egg.
GPGPU? General Purpose GPU? If embarrassingly parallel CPU algorithms weren't offloaded to the GPU previously, why would the A100/H100 price drop make a difference? We had cheap GPU in the past and we still left plenty of performance on the table with CPU programs because they were easier to build.
Is the idea that previously maintaining GPU programs was expensive whereas now AI makes it cheap? If so, I could buy that line of reasoning.
Maybe relatedly, I expect (hope) the hardware manufacturers will ramp up supply in the meanwhile which would also put downward pressure on GPUs. Right now though this hardware crunch is making me sad, not even because of GPUs but also because of general memory / disk.
BC250 shares memory with system. 16 GB GPUs have long been available to consumers for a small premium. What's never been readily available before is 40+ GB of HBM.
> I, for one, welcome the coming age of the post-LLM-datacenter-overinvestment-bust-fueled backyard GPU supercomputer revolution.
> The Big Question is…
> Who is cultivating the option to snap up and repurpose vapourised datacenter investments at fire sale prices, soon as the "datacenter debt" cometh calling?
They're pretty specific to the datacenter use case with no outputs and they need to be cooled externally, principally through the very loud high speed fans used in data centers. I suppose you could strap a fan to one and put it in a normal case or maybe make a dedicated after market cooler (like the water blocks made for water cooling cases).
I think there's a market for a home AI server that can run an LLM or video gen model behind a web frontend. Not literally a raspberry pi strapped to an H100, but something with lopsided enough specs that people joke it is.
And precisely because it's such a huge headache to do yourself, I think a small company could make a nice business wrapping up used datacenter cards in that sort of server.
I'm doubtful they can, based on current inference prices. An H100 draws about 50W while idling, which is ~$8/month at average US electricity prices. They also draw ~400W while active.
The electricity prices are relevant because if you paid $0 for your H100 and didn't use it a single time, you could buy millions of tokens in inference just on the electricity it draws while idling. If you can't keep that thing saturated through the night, you're probably underwater overnight. Likewise, it's too small to run even the frontier open source models so you need to be able to live with worse models.
Max power matters because you aren't going to run many of those H100s before you blow breakers in most houses. Newer houses in the US are 15A service to non-kitchen breakers, so 1650W (that might be peak rather than continuous, not sure). If you're plugging that into an existing run, you could maybe run 2 before you start blowing breakers? You can't just plug 4 H100s into the wall in a normal house.
Maybe I'm wrong, though. I'd be curious, it'd be neat to run my own inference for something more than what'll run on a 3080.
The median home age ranges from 26 to 63 depending on state, and 100 amp service was common into the 90s. 100A service is probably more common than you think, I'd bet we teeter around them being equally common.
My home still has 100 amp service (as does almost the whole city, most of it is a historic district), but it was built in 1910 so that's not shocking.
> A electrician can plop in a electric car charger for instance, that is a 40-50 amp circuit.
Oh yeah, I was mostly talking about having to upgrade the service to the house from the pole. A new circuit with a single outlet isn't horribly expensive, but needing to upgrade service is pretty rough.
Newer houses yes probably have the overhead. We put in a charger for my wife (because she drives 100 miles a day we needed a pretty beefy level 2) and had to replace our whole panel because 1) we had 125A so failed the load calc and 2) we had the infamous Zinsco panels and the electricians refused to squeeze one in even if we did pass the load calculations. Luckily we had a basement so they only had to extend the circuits a little instead of rerunning them all over the house.
I'd imagine that you can have the power to the board shut off when not in use.
But I never assumed it was to be cost competitive at current token rates. I assumed it was for the same reasons people might use open source hardware. Freedom to tinker, etc.
Basically sounds like a hobo version of PCI-e hotswapping which is a thing on enterprise motherboards but not many consumer boards. Also depends on if the board can actually power off the slots which I'm not sure many are setup to do.
A typical oven plug can offer 9.6kW (so your 24-node H100 can keep your kitchen nice and warm). My furnace was 16.5kW, which once failed in an odd way and was stuck on for about a week. The people living in the house simply opened the windows until I got around to fixing it. I estimate it wasted $400 on the power bill.
Helicopter. I live in Yeovil, Somerset, UK - there's a helicopter factory just down the road. I had a IBM "AS/400" or whatever they are called now in our computer room rack for a customer and it made nearly as much noise as everything else put together. It was clearly tuned for start up noise to impress because they would fire up in sequence, rise to a crescendo and then slow down in sequence to just a din instead of painfully loud.
A switch or PC server on boot will normally run up cooling fans instantly to max as a default protection mechanism until the "OS" has started and sensors read and then the fans will slow down to deal with the actual thermal load.
If you switch off your air conn, it gets noisy, quickly. Recently in the UK we are seeing routine temperatures around 30C and we broke 200 odd year records for temperatures a few weeks back. I know its even worse elsewhere but our infrastructure is not designed for this. Here we are at the same latitude as Calgary AB!
That's called staggered spin-up. One does that for sets of HDDs, too.
Anyway, what is it good for? To not overload the PSU for one, because spin-up needs more energy, until it has reached it's intended range of RPM. (for fans & HDDs)
Another reason to spin-up high initially, to fall down to slower when up, is to overcome friction in the bearings, to get stuck things unstuck, and shake dust off. (for fans)
Tell that to my switches and PC servers. They all spin up their fans in one go as soon as power good is received and the initial BIOS thingie has stubbed out its fag on the floor.
The PSUs are painting their nails when the fans go off initially. The GPUs, RAID controllers and co are still waking up.
A further counter example is an Equallogic PS6500. That has 48 3.5" spindles in it and three PSUs plus a few fans and two controllers. The whole thing boots in a couple of minutes.
My workstations are usually tower servers, which are the same design as their 4 and 5u rack counterparts.
Until the thermal management kicks in, they sound like jet planes. When the thermal management starts and assesses the required cooling, it’ll throttle down the fans to reasonable levels.
That is, until the moment you push the machine to its limits. When then happens, you might get back to the same levels of the boot time, but it’ll require you to push everything to the max - CPU, memory, storage (all 24 bays) and so on. For a normal user, there is a lot of room and it’s virtually impossible, even with a dozen of Teams windows open.
If you check eBay for V100 SXM variant, you will often see them sold in combination with a cold plate, together with a PCIe carrier board. There’s definitely a market for them.
You don't need much of a fan if you are prepared to pump a lot of water.
I mean, sure, if you're trying to cool the thing with 100ml of water, then yeah, you need a fan.
OTOH I have an unused car radiator in my garage, a quiet and cheap pump, and space to hang that radiator outside the window. It's barely an afternoon's worth of work if the H100s already have liquid-cooling intakes/exhausts on them already (I dunno, I have never seen one in the flesh).
I'm pretty certain 10+ litres of constantly circulating water (antifreeze, in my case) through a car radiator would be sufficient to cool down 2x H100s. Hell, add in 10 RPM used car cooling fan and I can probably cool 20x of those running at full-bore.
If anyone wants to ship me a bunch of H100s, I'll happily build it all out, take pics and videos and post it back here. Just sayin' :-)
The GPU alone has a TDP of 700W, together with everything else (CPU, RAM, storage, fans) you're looking at 1500W+. Depending on the country, that may be enough to saturate your home's electricity uplink...
In North America we are on 120V, making a standard 15A outlet only 1500W max, and something like 1200W sustained. To use higher wattage appliances, we have to upgrade our outlets to 20A (2000/1600W) or up our voltage to 240V, but that carries a different set of plugs and outlets as well.
Usually the home service is 200+ amps these days which is 24 KW total across all circuits, though you're right that most individual circuits are only 10-15 amp.
1,800 watt for periods less then 3 hours and 1,440 watts for continuous loads over 3 hours.
Realistically, a 15A breaker won’t trip on a 2,300 watt load for at least a few minutes, and often not until a few hours. A listed breaker is expected to trip in several seconds to 3 minutes for a 3,600 watt continuous load.
Take all those numbers and go up by 33% for a house or apartment with 20A circuits (quite common). Go up by 667% for an oven plug.
As anyone who takes care of rentals during winter knows, a typical circuit can tolerate two 1,500 W space heaters without tripping very much (although this is unsafe and a bad idea).
Sorry, I was doing conservative napkin math, better safe than burnt to a crisp, as grandpa NEVER said... I'm currently having to rewire a lot of the house that he built and I bought so he could retire. There have been a lot of electrical issues that the inspection never found.
He hated outlets, and hardwired almost every appliance directly into whatever circuit was closest that fit the amperage, including literally cutting the plug off a cord, to nut the wires directly into the Romex.
I still haven't found what circuit my range hood is on. I turned off all 120V and it remained blowing, and knowing grandpa, he powered it off a single leg of a 240V. And I know he didn't use junction boxes, so the splices are likely all inside the walls.
In the US they have a 120V system so it might actually saturate a normal socket over there. My PC room has a 16A fuse and 230V though, so should be plenty :)
The US has a 240V system for residential power, not a 120V system.
Most homes in the US have 200 amp service. That's 48 kW of power available to the house, though code states that you can only pull 80% of that continuously, so 160 amp/38.4 kW.
The 120V misconception comes from the fact that it is delivered as split phase on two 120V legs. Any competent electrician can run a 240V, 50-amp (40 amps/9.6 kW continuous) circuit to any room in a house. Plenty of homes have these circuits for electric ranges, EV chargers, or RV power outlets, and there's nothing stopping anyone from having one of these circuits installed in whatever room they'd like it installed in.
Heck, my old landlord and I installed one ourselves to provide an outlet for an EV charger in the garage of a house I lived in about a decade ago.
400 amp and even higher service is available, though it is fairly uncommon. Plenty of large houses will have a 400-amp service, though, which means you can double the power numbers available that I mentioned above.
gotcha. I've never heard of 16amp service, in Canada we have 15 amp.
But its very common to have many 15amp and 240V lines run in a house in Canada.
My kitchen alone has 4 different lines. i don't think your scenario is all that uncommon given that almost all new builds will have the same lines run in their home.
That's actually less bad than I assumed. High-end gaming GPUs are already almost at 600W with peak usage above that. I assumed it was much worse than that; that seems absolutely feasible for running at home.
> […] you're looking at 1500W+. Depending on the country, that may be enough to saturate your home's electricity uplink...
1500W is the power of the typical American microwave or (tea) kettle at 120V. (Convert over to a NEMA 6 plug and dual-pole breaker and you can get 3000W.)
3000W is nothing special in the rest of the world with >200V wall plugs.
Never heard electrical uplink before but to those wondering Italy, India, and Japan all have requirements around 3kw. I was slightly surprised by this, but in the end if your running this in a tiny space with that low of power your already probably not buying used H100s.
It’s common for car companies when they enter a new market. It removes uncertainty from the second hand market.
By doing that, you know upfront what the value of your used hardware will be at the time you decommission it. It removes a lot of the risk for buyers in a volatile market.
Its restrictive, but not that different from other types of deals.
For example, when I worked at a VFX software company, we were exclusively with one hardware partner. This unlocked something like a 50% discount across all our infra needs (they were big enough to provide switches, racks, servers, storage)
The hyperscalers signing these contracts have decent legal departments. Think about Oracle for example - I'm pretty sure they know every trick there is about beneficial contract drafting.
I don't think they need some special protection against this kind of contract.
so are buybacks. you choose to sign the contract. there's no way they didn't have an escape clause, although likely it meant not using the cloud provider anymore
Yes. "You choose to sign the contract" is not a reasonable argument, for two reasons:
1. There is one supplier, so you have no choice.
2. Even if you had a choice to sign the contract, this still means that it's not the same as a trade-in, because trade-ins are always voluntary, but once you have signed the contract, a right of first refusal is not.
In general, the "you chose to sign the contract" argument is a poor justification for bad contracts. If the contract is bad, it is bad regardless of whether you chose to sign it.
Yeah and now they have lost 98% of their profits in the last year. How's that catering to the ultra-wealthy working out for them? They used to build attainable cars that were nearly as cheap as a Corvette. Now they're double the price, and most certainly aren't double the performance.
From what I’ve read the only part of their business that is holding up is the 911 and particularly the high-end variants. It’s everything else that’s struggling.
> pretty much nobody cares about ML processing for analysis.
I work in a bank and a can tell you that the customers absolutely hate ML when it rejects their loan application. Over the pond in the US, I have an impression that the fico score is not exactly popular either, but I have no first hand experience.
In the US, FICO scores are mainly unpopular with "credit criminals" who have low scores. The score performs quite well in predicting how likely a borrower is to repay a loan. The problems that arise with credit scores in general are when they're used for purposes they were never designed to serve, like screening job applicants.
For home mortgage lending, FICO scores are now being partially replaced by VantageScore.
WoW has become quite soulless after the changes where everyone gets the same end game equipment but you have to grind upgrade currency to improve it.
My wife and I used to play quite a bit but it doesn't really engage anymore. Perhaps we are getting too old, but we should be in the correct customer segment (mid 40 years old).
Now I have spun up a local wotlk server with player bots powered by ollama which is actually a bit fun again.
I agree, that approach to game development is soulless. I don't really play any new video games anymore. In the rare instance I do play games, it's something from 15+ years ago. The games were designed to be played, beaten, and that was it. It wasn't an effort to keep you locked in forever. It became that. Then it became a soulless effort to keep you locked in forever.
Unfortunately, soul doesn't matter. The quarterly statement does.
Edit: I have a pro subscription now but it has also been on a free tier level for perhaps 1.5 of these years.
reply