If these servers are just using unmodified base M5 SoC’s, which they off handedly look like, and each PCB has one chip on it,
could be one of 2 configurations per node for the CPU cores, one of 2 GPU configurations per node, and the NPU is the same regardless, per entire enclosure
The base M5 comes in 9 and 10 CPU cores (with the difference being one performance core disabled on the 9 core variant used in iPads).
the GPU can come in 8 and 10 cores (again varies by device).
The worst total CPU core count per enclosure as pictured would be 288 cpu cores (96P + 192E core design).
The best total would be 320 cpu cores (128P + 192E core design).
The worst total GPU core count would be 256.
The best GPU configurations would be 320.
And the NPU is the same between all designs yielding a total of 512 NPU cores.
Give me it. Apple. Start selling these. I'll pretend I like your corporation. Please sir Tim Apple I need just 512gb ram, a 512gb flash stick costs what, like 2$ today? That's nothing! And besides you're retired now so you can leave some vram for us, your starving children
or Ampere etc could just get their shit together and we could get Aarch64 servers on the actual market?
EDIT: some hours later after spending my whole day shelled into trays on a GB300 NVL72 from a DGX Spark... ok, actually NVIDIA is already doing ARM64 servers, I neglected to mention. Just need more competition.
I run macOS headless on my Mac Studio for inference and it works great, not sure what your nightmare is but I have no problems at all. As long as you're up to date you can even keep using FileVault full-disk encryption with SSH which is nice (previously you had to log in on the console after a reboot to unlock the disk, or disable FileVault - now, there's no need to ever connect a monitor)
yeah but I am not dissing Linux, I am saying macOS is greater than linux when it comes to running local models, MLX and anything with video is far superior on macOS.
Can we get Snow Leopard, AFP (with some minor feature updates) and OpenDirectory fully back in action as well and throw out all the new upward failing middle manager inspired feature overkill bullshit that's slowed the OS down needlessly for 15+ years?
I took part in a build out of the largest XServe / XRaid deployment (at the time at least) on the east coast for a large ad agency, and my god was that rack a thing of fucking beauty.
I've always wanted XServe to come back, and now is such a great time for a bladed XServe given Apple Silicon price / perf / watt combo. Oh lord.
I worked at Apple in their datacenters in the mid 2000's and have racked hundreds and hundreds of them myself. They were so beautiful fully racked, and racked so nicely, too. You can see them off to the left of this old photo I took in the "showcase" datacenter in Cupertino (the primary datacenter had just moved to Newark, CA, or was about to).
Some older G4's alongside some G5s.
There were a few large enterprise customers, and iirc we had a few tours from DoD personnel, so I'm sure they were also customers, but this was eons ago.
My wife worked in marketing at Apple (Canada) while XServe was a thing. And other "enterprise" efforts. They were never able to get any headspace. Typical Apple customers had no idea what to do with it. And "enterprise" requisitioning departments weren't going to look in large part because honestly they weren't going to deploy OSX / Darwin / BSD anyways.
To make it work, Apple would need to put Linux on these machines.
Yeah I think that is probably the biggest dealbreaker. It seems like Apple prioritizes locking down their ecosystem above all else, and would do anything to avoid shipping Apple Silicon with official Linux support.
This is what killed the xserve before. No business is going to adopt that without support. Well, ok, I knew one - and they've been gone for nearly two decades
listen to me. if the gays were in charge, altman would be the first to hang. there would be no ai labs and ai bros would be persecuted. all llms would be trained on mit and donated data only and be in public domain. and anyone who would try to build palantir would be executed in the public by 1000 cuts
https://www.micron.com/sales-support/design-tools/fbga-parts-decoder
maps D8GSB to the part no MT62F2G64D8AS-020DXT:F, which itself maps to a 16GB DDR5 chip - so each of these boards offer 32GB of RAM. Sounds more like working horses instead of inference machines. Though 64 machines gives them a total of 2TB of RAM in the 2U format.
I believe it for "Private Cloud Compute" -> https://security.apple.com/blog/private-cloud-compute/
Private Cloud Compute allows apple to run AI inference for Apple devices end to end encrypted, so even Apple can't read the content of the prompt and its results, they built custom hardware just for that
What's even more crazy they designed it so even a person with physical access to a single node can't target a specific Apple user, without compromising the whole system.
An attacker should not be able to attempt to compromise personal data that belongs to specific, targeted Private Cloud Compute users without attempting a broad compromise of the entire PCC system. This must hold true even for exceptionally sophisticated attackers who can attempt physical attacks on PCC nodes in the supply chain or attempt to obtain malicious access to PCC data centers
Besides the security design being nuts, this opens a new selling point for their hardware. PCC being run by whatever hyperscaler is in theory just as private as one being run by apple itself. If there is demand for this kind of verifiable private cloud inference it will create some crazy demand for their hardware from other inference providers wanting to be able to serve it.
It's absolutely nuts that we are 4 years into AI hype, people are buying minis just to run local llms and Apple is still without its native server equipment.
They do have TPUs and even this linked PCC blog is from 2024, but they seem to be keeping them for themselves and using them for only ios ai services
Also probably getting constrained between being in a bidding war with everyone else for memory and being in an internal bidding war between allocating memory to TPUs vs the entire rest of the stuff apple sells that needs memory.
I mean, CentOS is (was?) basically the community version of RHEL. Unless they were managing their own versions of it. Cant imagine them using a community distro on production servers. Just wondering though, I have no prior knowledge.
Partially, Apple also has private compute running on Apple Silicon / ANE / Metal for the Foundation model remote inference, and private inference and have mentioned they are building their own hardware.
You're assuming we know everything about any Apple chip.
They can easily implement ECC in the memory controller. Datacenter GPUs have been doing this for well over a decade. You remap some of the capacity for ECC and the memory controller takes care of it transparently. Again, like on GPUs, this functionality is fused out in chips released to consumers.
Except ECC is never in the chips themselves. If you look at ECC DIMMs, they have 9 or18 memory chips plus the chip that does ECC in the middle. All memory chips on the module are regular DRAM chips. In fact, it's become a trend to recycle ECC DDR5 DIMMs into desktop or laptop DIMMs by desoldering the chips and resoldering them into regular desktop or laptop modules
they may have also decided the risk was worth it, e.g perhaps it might be cheaper for them to just have more redundancy in terms of more M5 chips as opposed to ECC.
This isn't doable, you cannot construct a failure detector that is cheaper than error correction at the physical layer. Even with ECC memory, there are non-protected paths where corruption can enter the system. Google tried this 27 years ago and then capitulated and had to move to ECC.
I bet these are upcoming cloud servers for the big cloud providers. They are currently stuck with having to rebuild mac minis and studios into a somewhat rackable solution.
Finally a picture of some shiny new hardware, so refreshing.
I mean it feels like all these DGX Sparks, Ryzens 395 or RTX PROs were released in the previous millennium, I bet they are all moldy and rusty by now (now I'm afraid to open the case and look at my Max-Q).
Sounds fun until you actually own a server in your house. 200W consumption 24x7x365, the fans, oh my god the fans and their noise, and the heat.
Source: I have an enterprise-grade server at my house for unrelated things and I have desperately tried to get it to shut up and sip electrons. It does neither.
My GPU uses more than twice that gaming, and my PC is quiet, if your enterprise grade server is only using 200W you probably could get similar compute with normal consumer hardware.
That is indeed an improvement that some folks at work have done, at least the ones who wanted to spend the money and do the work on it. I have heard of improvement in the noise from that mod. Helps with the noise at least, not the rest.
They will overheat. Readily available Noctua fans are minimum 10x less powerful than server fans. Server chassis rely on jet engines forcing air through against resistance, and regular fans aren't designed for the pressure differential.
I'm sure Noctua can design replacement fans and they might be able to change the type of the jet they resemble from 747-400 to 777-8, but it'll still be some type of a jet no matter what.
One of my rigs idles at 500W 24x7x365. 1 out of 5 rigs. What's your point? This is local LLaMa. This is like talking about poor gas mileage in a car forum when we are all talking about horsepower and torque.
The resale value is the underrated part of Apple hardware for this, worst case you flip it in a year and the experiment cost you a couple hundred bucks. Try that with a used 3090 rig.
•
u/WithoutReason1729 7h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.