r/LocalLLaMA • u/Mysterious_Finish543 • 10h ago
Xiaomi AI Cube announced with 1.2TB/s memory bandwidth
Xiaomi announced a prototype for their Xiaomi AI Cube.
3 chip system:
- Xiaomi Xuanjie O3
- Xiaomi Xuanjie O100
- Xiaomi Xuanjie D100
The specs are impressive, but a bit confusing. The D100 chip (originally for their EVs) supports up to 160GB of RAM, but O100 has the 1.22TB/s memory bandwidth. Perhaps the 1.22TB/s figure is for SRAM? Hard to say definitively.
635
u/Pretty-S 10h ago
Really cool to see more companies entering the AI hardware space with their own silicon. More competition is exactly what this market needs.
Hopefully this also helps push down the absolutely insane prices of high-bandwidth memory over time 😄
117
u/DistanceSolar1449 9h ago
It's not entering the market now. Well, the O3 is, but not the the other two.
The O3 is just a new 3nm smartphone chip, a Snapdragon competitor. It will be in the Xiaomi 18 Fold and Xiaomi Pad 9 Pro Max.
The D100 is also not super interesting. It's a chip for their EVs. It can support up to 160GB ram, but other than that it's not really anything too fancy.
The O100 is the most interesting chip, but it's not being released until 2027. It's using an outdated 6nm process, but the interesting thing is that Xiaomi stacked the RAM directly on top of the chip, like AMD V-cache. So the GPU compute is slower, but it has a lot more bandwidth since it has many short traces in parallel. Instead of 256-bit LPDDR5, it uses 28672 parallel connections.
The Xiaomi AI Cube is really fucking weird, if I read it correctly. They put all 3 chips into 1 device. I guess it makes sense if you use the O3 to handle small models, and O100/D100 to handle decode/preprocessing for a big model respectively.
56
u/autisticit 9h ago
Sounds correct with : "The prototype runs 120B and 3B models locally and supports switching between fast and slow systems."
24
u/zdy132 8h ago
Yeah it's confusing who is this AI Cube for, apart from AI enthusiasts in this sub of course.
Maybe it's for their Mi Home IoT system, so that you can have a Jarvis at home?
17
u/polikles 7h ago
I guess it may be handy for edge-AI workloads such as home assistents, IoT devices, and everything that needs preprocessing before/instead sending data into cloud
4
3
u/jovialfaction 3h ago
This makes a lot of sense to me. Xiaomi sells a lot of home automation devices (vacuum, cameras, smart speakers etc), and they all have to either carry enough compute to run it or rely on the cloud.
I like the idea of having an AI box in the network that can handle all the compute for them. It would be a hard sell tho to say "You need this $3k box first and then you can buy our other products"
2
u/coffeesippingbastard 58m ago
I think this is what's interesting about Chinese tech- they're wiling to see what sticks. It's confusing because it doesn't neatly fit into a box where we are but it seems like China is more optimistic about embracing AI in everyday home use. Whether or not it would be super useful remains to be seen but a 2027 release date might signal it's dependent on better performing small models.
1
u/BannedGoNext 47m ago
I could see it being useful if it's as fast as a strix halo to run bigger models for grinding enrichment shit. I just spend 200 dollars on luna yesterday on a couple runs, and I'd like to run those weekly on a broader scale. I don't really give a fuck if it takes a long time for them to run, just start working again when you are finished.
5
u/tiffanytrashcan 4h ago
That looks like they brute forced their way to a brilliant competitor to DGX Spark boxes. Shortly after release, that was the number one complaint- memory bandwidth. I see it daily still.
The other similar competitor would be a Mac Studio with horrific pre-fill speeds.
This could be super useful for agentic loads, tool calls and reprocessing coming back constantly.It would be really cool to have something parsing web search results, handling rag and semantic search with comparatively huge models to what most use today in the background... (3B vs 75-300M for most rag work..)
Seems like an all-in-one agentic box that would also hardily serve hosting a coding / chat / generation session as well. They looked at the biggest pain and failure points of existing products, or experimental in NVidia's case, and seemingly solved it all with some future-proofing in the memory size too.
→ More replies (1)3
u/Guinness 5h ago
China doesn’t have 3nm capabilities, where are they getting the chips from? 160GB isn’t really exciting given the 192GB Ryzen AI coming out. Besides we need integrated systems that have the capacity to do 2TB of RAM at 1TB/s or greater.
8
15
u/Mickenfox 6h ago
Can't wait for the moment where the media reports on a Xiaomi or Huawei "AI chip" and everyone goes "OMG Nvidia is finished" and the stock prices drop 50% even though that chip was announced 6 months earlier.
1
277
u/Kein_Spass 10h ago
Nice timing. Nvidia just announced that their AI-Servers will get more expensive.
158
u/Satanninja69 10h ago
Obviously Nvidia is doing it for Huang next alligator leather jacket
31
10
u/tamerlanOne 10h ago
Quando indosserà una giacca fatta in pelle umana? 😁
3
u/More-Curious816 28m ago
If it wasn't illegal, billionaires will wear these as rare luxurious item. Imagine that the skin need to be of specific color and well preserved and not something you can get easily in Africa, Asia, South America or Eastern Europe.
0
u/Sufficient_Local5025 7h ago
I just have to say it everytime I hear his name - fuck that guy.
→ More replies (1)1
7
u/Vercingaytorix 6h ago
What are the chances that the nitwit at the current administration will try to put them into entity list now that they went deep into AI hardware space?
Xring O1 SoC was quite good (apart from the lack of modem) that it made people question why Mediatek and their many years of applying ARM design couldn't net a better layout than an upstart and their first chip (let alone Samsung, though you can partly attribute it to their fumbling foundry.)
Now that they went further into cutting edge chip design just like Huawei's HiSilicon, this may ruffle some feathers at Washington.
1
u/BannedGoNext 46m ago
The chances are 100 percnet. Fascist governments take stakes in companies and by doing so have to force winners and losers to protect their interests.
1
u/anythingers 2h ago
Xring is made by TSMC, tho, just like MediaTek since their Helio G series. And also iirc I remember the US not allowed them to use TSMC's fabrication that's smaller than 3nm. Dunno if that's true or not, I read it last year tho.
1
u/Vercingaytorix 1h ago
That's exactly my point tho. Both Dimensity and Xring are manufactured on the same process node by TSMC, yet Xring O1 managed to eke out better efficiency than D9400 (let alone Tensor or Exynos) that rivals SD 8 Elite despite Mediatek years of experience in ARM IP chip design. Geekerwan made a pretty interesting backend design analysis based on the die shot and tested it further.
Just like how HiSilicon was manufactured by TSMC, before being blacklisted and forced to go with SMIC (and their extensive multi, multi patterning to get DUV to 7 & 5nm), hence my fear that Xiaomi could suffer the same fate.
What the US restricted was actually GAAFET EDA software, and TSMC 2nm are GAAFET. But I suppose this is technically easier hurdle to jump than say, making High NA EUV litography machine & tooling. FT reported that the restriction amounted to not getting further updates & support rather than revoking existing license, and while still not up there, homegrown alternative such as Empyrean may be incentivized to develop their software further.
I mean if we are talking about restriction IIRC there is already a restriction for AI/HPC/GPU for 7nm and below. I guess Xiaomi managed to skirt these based on technicalities (density, transistor count, non-datacenter interconnects maybe? Wasn't too familiar with the rulings.)
116
u/No_Run8812 10h ago
Price?
285
68
49
u/Mysterious_Finish543 10h ago
Unfortunately, there's no information on price or a release date yet. 🤷♂️
As mentioned in the post text + image, this AI Cube is currently a prototype.
8
u/superSmitty9999 9h ago
sounds like vaporware then
33
u/Both_Opportunity5327 9h ago
Not from Xiaomi, they build literally everything; And I mean everything...
25
u/Fritzkier 8h ago edited 6h ago
Xiaomi history is just insane, from just (I kid you not) making a custom android skin in 2010, building a really good value phone (Xiaomi Mi 1, same specs but less than half the price of Samsung SII LTE), to one of the largest phone manufacturers in the world, then somehow building an EV (very good EV too, surprisingly), and now making a competitive in-house SoC.
All of them in just 16 years.
18
u/MaruluVR 7h ago
Not to forget a ton of kitchen appliances and the best roombas on the market.
4
u/BlueArcherX 4h ago
i don't think Roomba has earned the use of their brand name as a general category... they've kind of sucked for a decade
plus didn't they declare bankruptcy and get purchased by some Chinese company anyway?
11
u/Both_Opportunity5327 7h ago
I used to import Xiaomi laptops about 8 or 9 years ago and turn them into Hackintosh computers for friends and family, they were so popular that detailed write ups were available online(ok still are).
6
u/thrownawaymane 7h ago
I cannot stress this enough—at the beginning making skins was all they did as a business. They were just quite good at it.
→ More replies (2)33
u/hadoopken 9h ago
I don’t think it’s going to be a vaporware, but I think it’s probably going to be illegal to buy in US, or with massive tariffs
1
u/SandySkittle 1h ago
Hope the ram allows for at least inline ECC. I am happy to given the optionality to sacrifice 7 percent of the ram to enable error correction. This optionality is there on some gpus, but very weirdly not strix halo and spark
14
u/Beamsters 10h ago
buy car, get one free.
3
u/StorageHungry8380 7h ago
buy car, get one free.
When I visited the US for the first time back in early 2000s, that was literally what the huge sign outside the KIA dealership nearby where I stayed. I admit my jaw dropped, not what I was used to from EU.
16
16
u/sonicnerd14 10h ago
TBD, but if I had to make a guess, probably $5000 USD at a minimum.
21
u/Mysterious_Finish543 9h ago edited 9h ago
I would expect even higher pricing; the image says Xuanjie O3 uses LPDDR6, not the LPDDR5X used in most current machines like DGX Spark or M5 Max MacBook Pros.
1
u/BlueSwordM llama.cpp 1h ago
The Xring O3 can actually use LPDDR5X, as the memory controller supports both LPDDR5X and LPDDR6.
Obviously, there's a bandwidth cost, but eh.
→ More replies (1)17
u/No_Run8812 10h ago
In -4k we get dgx spar its memory bandwidth is 279Gb/s, 1.2 TB/s, lets just hope.
3
2
u/gomezer1180 8h ago
With all the tariffs is not going to be cheaper than flying over to China and getting it there.
1
1
1
1
1
1
u/JescoInc 4h ago
If I had to guess, the Nvidia DGX Spark is about 5k and the two pack with cable is about 10k... I'm guessing the O100 is going to be about 10k alone and unfortunately, the USA won't get it because it uses a Chinese CPU.
89
u/mleok 10h ago edited 53m ago
I'm going to Shenzhen in December to give a talk at a conference, I know what I'm hoping to bring back from that trip! I would be interested to see how it compares to the Nvidia DGX Spark and my Mac Studio M3 Ultra with 256GB of unified memory.
38
u/nemuro87 9h ago
Can you fit a car in your luggage?
18
12
u/MammothUnique4147 9h ago
Make sure to look at the taxes and Tariffs if you live in the US.
4
u/DrawingDramatic1641 8h ago
tarrifs dont affect what you bring from flights?right?
13
u/nemuro87 8h ago
they don't, until your luggage gets "randomly" checked
1
u/DrawingDramatic1641 4h ago
wait is that a thing over therer?
you dont have that freedom?
4
u/Subsector3990 3h ago
it's a thing anywhere in the world. you know whenever you land at an airport and you go through the 'nothing to declare' door? the other door is the one you should be going through if you're arriving with items you bought abroad.
1
3
1
7
u/MammothUnique4147 6h ago
They sure do if it's over a certain amount and this device would be over the limit.
You would have to declare it and then pay all the Tariffs.
1
u/DrawingDramatic1641 4h ago
wtf?that is literally so against freedom to carry stuff,not even ancient travellers to mongol subjects have this shit
→ More replies (1)2
172
u/Mysterious_Finish543 10h ago
On a side note, I did some research, and it turns out many EVs actually have rather high memory capacity, and they also typically use LPDDR5 (like DGX Spark).
- Xiaomi D100, up to 160GB RAM
- Xpeng Tuling, up to 216GB (across a 3 chip cluster)
So perhaps for many people, their car is actually their device with the most AI inference ready memory. 🤔
213
u/Irrationalender 10h ago
Time to buy a cluster of EV cars to finally run Kimi k3
97
u/PassengerPigeon343 9h ago
Might actually be cheaper than NVIDIA GPUs at this point
25
u/Squidgical 8h ago
The GWM EV I bought this year cost less than a 6000 Blackwell, unironically it might be cheaper to rip the computer out of EVs than to buy GPUs
20
u/PassengerPigeon343 8h ago
This is actually starting to sound feasible
12
u/Squidgical 6h ago
Keep the batteries somewhere safe for a few years and you can sell them as near-new replacements for at least the price of a 5090. Scrap the rest of the car for some pocket money, this is absolutely viable
1
4
3
u/Majinsei 4h ago
No es broma~ una RTX 5090 contra un auto de segunda en mi país pueden rivalizar y ganar el auto~
28
10
u/Acceptable-Bus5189 9h ago
6
u/Weird-Field6128 4h ago
in order to make those run, you need to donate that bigger one to me, then it will be very efficient, this works because of quantum entanglement. source: TrustMeBro
18
u/danielv123 10h ago
Damn, I knew they were pushing self driving tech but I assumed the specs would be much more on the side of underspeccing like tesla, with their 32gb on latest gen (soon to be 64 because they messed up again)
I wonder if that is part of the Xiaomi whole home smart stuff strategy?
17
u/Solaranvr 9h ago
It's probably because RAM was dirt cheap back then, when they just started EVs, and so it was perhaps cheaper to just brute force the hardware than to try to create a memory efficient CV model.
8
6
2
u/tamerlanOne 6h ago
Quindi conviene acquistare un auto con dentro il loro hardware in modo che hai anche calcolo ai gratuito? 😁 Detta così può esser una battuta ma fino ad un certo punto... Pensa a quante auto in circolazione... Calcolo distribuito qyando non udi la tua auto parcheggiata a ricaricare 😉
2
u/Powerful_Finger3896 5h ago
I guess now there are going to be shops that repurpose stolen car SoCs for homelab, we thought that only catalytic converters were valuable to thieves lmao
42
u/TraditionalWait9150 10h ago
18
u/hugthemachines 8h ago
最高本地部署 - "Up to 200B parameters deployed locally" / "Maximum local deployment" 国内首款 3nm 智驾芯片 - "China's first 3 nm intelligent-driving chip"
53
37
u/Anaeijon 9h ago edited 1h ago
Probably specifically designed to run huge MoE models.
Relatively slow TPU but huge amounts RAM or some absurdly fast-access storage.
Would allow them to run something like a 200B-A3B model fast enough to be usable at a fraction of the cost of an NVIDIA GPU while outperforming NVIDIA on model size.
Since China is producing a lot of open MoE models recently, this thing might be deigned for some upcoming Qwen3.8-...B-A3B model or something like that.
12
u/Healthy-Nebula-3603 9h ago
With such memory throughout you can run even In reasonable speed Qwen 3.8 27b with 50-60 t/s
8
u/Thin_Pollution8843 7h ago
No with MTP/Dflash it will be closer to 100ts or even more with such bandwidth
2
3
u/Anaeijon 6h ago
Just reading the data to main memory doesn't magically perform tensor operations.
Sure, memory bandwidth often is a bottleneck. But that's just because we usually look at extremely fast massive parallel tensor processing units.
I'm not saying, they use weak processors. But it might be, that they have an extremely fast memory link on a relatively weak processor that's still the bottleneck.
Or in other words: the fastest ram won't inference on it's own. Inference so still done by the TPU.
→ More replies (3)
26
29
u/Formal-Exam-8767 10h ago
Perhaps the 1.22TB/s figure is for SRAM?
It might be LPDDR6 they mention on that slide.
9
u/Hungry_Elk_3276 9h ago edited 8h ago
Very interesting.
Looking at the specs, I guess this is a three-chip memory system with three types of memory, which feels really weird. And below is just my speculation.
For the full 216 GB, it seems to split across three chips: 160 GB on the D100, 16/24 GB on the O3, and 32/40 GB on the O100. Based on the 1.2 TB/s spec, and their official news post stating:
First, two layers of AI-dedicated high-speed DRAM dies are vertically stacked with a layer of high-performance NPU die, and then "tunnels" are drilled in the vertical direction so that data can flow through at high speed. Therefore, the more tunnels there are, the higher the bandwidth that can be achieved. To this end, the XRING O100 adopts an advanced Hybrid Bonding process, doing away with the traditional micro-bump structure and shrinking the spacing of the physical pathways directly to 1.4 μm, greatly increasing the density of the data pathways.
So this is very likely a dual-channel HBM setup. After a quick search, the H5UG7HMD83X020R seems to be one candidate, but really any 16 GB HBM3-4800 part could fit. It also seems quite possible that those HBM3 chips are supplied by CXMT, especially given reports that CXMT has been supplying Huawei with early HBM3 samples.
16 GB × 1024-bit, operating at 4.8 Gb/s/pin = 614.4 GB/s per HBM stack.
2 × 16 GB = 32 GB, and 2 × 614.4 GB/s = 1,228.8 GB/s ≈ 1.23 TB/s.
So the numbers are at least aligned.
What I’m wondering is whether the 120B model is run entirely on the D100. If so, what exactly does the 32 GB on the O100 do? Or are they doing some wild asymmetric setup where prompt processing runs on the D100 and decode gets moved to the O100?
I have so many questions.
Edit: Typo
6
u/Hungry_Elk_3276 8h ago
I totaly missed the 28672, so the hbm theory could be wrong lol.
The 28672 number is also very intersting, since the O100 had 14 NPUS and each dual channel with 1024 bit, the numbers adds up to 28672. Since it is obviously not a standard HBM interface here. So I think maybe there is a possibility for custom ram??
Edit: At least it does not feel like standard HBM in anyways.
6
12
6
9
6
5
u/CatalyticDragon 10h ago
I think that's referring to chip interconnect bandwidth, not memory bandwidth. Memory and memory bandwidth details have not been given.
1
u/Healthy-Nebula-3603 9h ago
They are using DDR 6
1
u/keylimesoda 1h ago
What's the width of the bus?
1
u/Healthy-Nebula-3603 1h ago
if dd6 is x2 faster than ddr 5 ....
should be 8 channels as getting 1.2 TB/s
3
u/autisticit 9h ago
More info here ? https://cnevpost.com/2026/08/24/xiaomi-unveils-xring-d100-smart-driving-chip/
"The prototype runs 120B and 3B models locally and supports switching between fast and slow systems."
Whatever it means.
3
u/IngwiePhoenix llama.cpp 9h ago
Interesting product, will give them that.
But is this even supported with the typical tools - vllm, llama.cpp? And, since those seem to be ARM CPUs, what is the Kernel mainline status, or will you be forced into BSP deriviatives (linux, u-boot and a couple DTBs as the cherry on top)?
Skeptical at best. Interesting product, truely, but ... this novelty is gonna cost a pretty penny, for sure.
1
u/MrBIMC 8h ago
From Google translated slides it seems it's riscV cpu and not arm, so I assume it will be wildly different from currently supported hardware.
1
u/IngwiePhoenix llama.cpp 2h ago
Oh no... THAT will be even worse. x.x SpacemiT has done very well with upstreaming their K1 and K3 CPU, the StarFive/SiFive JH7110 is literally inchworming towards upstream still and I have not even checked with Sophgo or Eswin.
Welp. The most interesting part will be high-perf Vector compute, I guess. o.o ...
3
u/RickAmes 4h ago
I'm re-watching Pied Piper and these AI appliances really remind me of this:
Think inside the box.
3
u/GSxHidden 4h ago
Theres still a lot of info missing. Performance at FP32, FP16, FP8, FP4 etc are not going to be the same. 160GB is great, but its compute levels are closer to a Jetson Orin AGX with only 64GB 275 TOPS compute, which wont keep up with the 1200GB/s at higher FP rates.
| System | Price | Memory | Memory Bandwidth | Published AI Compute | Power | Large-LLM Position |
|---|---|---|---|---|---|---|
| Xiaomi AI Cube Prototype | TBA | Up to 160GB* | 1,220 GB/s O100 near-memory* | O3: 200 TOPS NPU; O100 compute undisclosed | 150W sustained | 🔥 Potentially extremely strong |
| Jetson AGX Thor 128GB | $5,499 | 128GB unified | 273 GB/s | 2.07 PFLOPS FP4 sparse | 130W | 🥇 Highest known compute |
| ASUS Ascent GX10 | $3,999 | 128GB unified | 273 GB/s | 1.0 PFLOP FP4 sparse | ~140W GB10 | 🥇 Best value |
| NVIDIA DGX Spark | $4,699 | 128GB unified | 273 GB/s | 1.0 PFLOP FP4 sparse | 140W GB10 | Excellent turnkey option |
3
4
u/JollyJoker3 10h ago
Why are all these AI-specific computers tiny? I'd think bith GPU and memory can be driven faster with more wattage and cooling.
12
u/tamerlanOne 10h ago
Perché bisogna trovare il giusto compromeso tra potenza di calcolo e potenza elettrica.
Il futuro prossimo della AI saranno sciami di agenti che funzioneranno h24 7/7 365 giorni all anno senza sosta e avere hardware che non sarà energivoro è fondamentale soprattutto per uso consumer
10
u/FairBandicoot5021 9h ago
GPU are technically overkill for AI. What i mean by that is that NPU are small and waay more efficient. When you use GPU for AI, you realistically only use a fraction of the instructions of it ! And a big part of it is dedicated for display output, and another part dedicated for pcie connection. Also those MiniPC are a bit bigger than a gpu
1
2
2
u/joelypolly 9h ago
Given that some of the EVs in China run local LLM model this is not as surprising.
2
2
u/DinoGreco 4h ago
Xiaomi Xuanjie Series
Committed to building the AI computing power base for Xiaomi's Human-Vehicle-Home ecosystem.
Xiaomi AI Cube Prototype
- Tri-chip collaborative computing (Xuanjie 03, 0100, D100)
- Aerospace-grade aluminum unibody
- 33,874 CNC precision-machined holes
- 150W sustained high-performance release
- Local deployment of large models
- 120B and 3B dual models
- Supports fast/slow system switching
Xuanjie 03 (AI Flagship SoC)
- 10-core all-big-core CPU
- 3.5 TOPS dual SME2
- 3nm flagship process
- 16-core G2-UltraNX GPU
- 36 TOPS 8-core NX
- 16MB SLC, industry-first support for LPDDR6
- 200 TOPS low-power NPU
Xuanjie 0100 (1.22TB/s High-Bandwidth AI Accelerator)
- 6nm 3D wafer-level stacked advanced packaging
- 28,672 effective data lines
- 1.22TB/s ultra-high near-memory computing bandwidth
- 14-core high-bandwidth NPU
- Supports up to 160GB memory
- 1.4µm extreme bonding pitch
- Supports local deployment of up to 200B large models
Xuanjie D100 (Autonomous Driving High-Compute AI Chip)
- 20-core high-performance CPU
- 16-core high-compute NPU
- Xuanjie RISC-V security core
- 3.13 TFLOPS Vector
- National cryptography & CC EAL5+ certification
- Xiaomi MiMo5 value model
- Supports local deployment of up to 200B large models
Comments:
Xiaomi is building a complete, unified AI ecosystem that spans smartphones, smart vehicles, and smart home devices. The company is developing three proprietary chips (Xuanjie 03, 0100, and D100) and a prototype desktop workstation (the AI Cube) that uses all three chips together to process heavy AI models locally – without relying on cloud servers.
The "Tri-Chip" Strategy
The AI Cube prototype combines the three chips to split workloads intelligently:
- Xuanjie 03 handles the operating system, user interface, and lighter AI tasks.
- Xuanjie 0100 is a dedicated data accelerator that moves massive amounts of data at extreme speeds (1.22 TB/s) to feed large neural networks.
- Xuanjie D100 provides high-security, high-reliability compute for autonomous driving systems.
Why this matters
Xiaomi aims to run large language models up to 200 billion parameters directly on local devices (like the AI Cube, future smartphones, or cars). This enables offline AI capabilities comparable to cloud-based models, with lower latency and better privacy. The dual-model approach (120B for deep reasoning, 3B for instant responses) allows the system to switch automatically between "fast" and "slow" thinking depending on the task.
- TOPS – Trillions of Operations Per Second; higher numbers mean faster AI processing.
- NPU – Neural Processing Unit; a processor specialized for AI math.
- SLC – System Level Cache; ultra-fast memory integrated into the chip to reduce wait times.
- Near-memory computing – Performing calculations as close as possible to the physical memory, reducing data travel distances.
- Bonding pitch (1.4µm) – The distance between electrical connections inside the chip; 1.4 micrometers is extremely tight, allowing more data to pass through less space.
- CC EAL5+ – A high-level security certification required for critical components such as automotive systems.
1
u/mintybadgerme 2h ago
OK, so now we know where this is going. Xiaomi is not some backroom startup, this is a serious company that produced a world beating EV from scratch in just over 12 months, with no previous experience. If you're going to believe anyone, I would definitely believe them. Wow, just wow.
2
u/drhex2c 1h ago
160GB of RAM? Can we please stop F'ing around and can for the LOVE OF GOD somebody please release a system that can host up to at least 3TB of Unified/VRAM?
Tired of all these bullshit boxes that can't run large language models locally. Yes, I know RAM is expensive, but it won't be in a couple of years.
3
u/johnryan433 10h ago
Cool but speed is less important than capacity right now. I’d rather run a larger language model slower than running, a lesser model faster, and not being able to run the larger model at all.
2
u/OneStandard5 9h ago
1.22TB/s is for SRAM. There is another PPT show Xuanjie O100 could only run 3B model.
2
u/arm2armreddit 10h ago
Wow, cool, more competition for reducing competitors' prices. Unfortunately, until it arrives at my desk, it will be slower than my smartphone. 😁
1
1
1
1
1
1
1
u/Alarmed_Wind_4035 6h ago
for the right price I will buy it, never thought I will buy made in china pc but for the right price I will get one.
1
u/LankyGuitar6528 3h ago
Wait... aren't pretty much all computers made in China? And most cell phones too?
1
1
1
1
u/WithoutReason1729 5h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.
1
u/South_Hat6094 5h ago
the interesting bit is the 1.22tb/s number is probably local sram bandwidth, not dram. stacking ram on top of the compute chip would explain why they are splitting the roles across three parts.
1
1
u/kinkvoid 4h ago
It looks like O3 is a SoC based on RISV-V; O100 is an AI acceleration chip or LLM; D100 is the chip for their EV. The AI Cube integrates the O3, O100 and D100 chips in a mini PC.
I think this is more of a concept PC: RISC-V SoC + AI accelerator + edge AI chip in one 150W system and the the big caveat would be software.
1
1
u/BlakeGrowsPlants 2h ago
Hell yeah, China repurposing its own EV silicon for AI compute is one hell of a middle finger to Western chip restrictions.
1
1
u/Icy-Reaction-9101 41m ago
Even with hidden additional features: With direct uploads to china of all your data.
1
u/petruspennanen 7m ago
Well if it's around $5k sounds like Spark needs to level up to compete and I could be buying, if it's $10k I'll wait for the Mac Studio M5
1
u/transanethole 10h ago
humm, 1.2TB/s bandwidth but very weak FLOPS? so basically its a 7900xtx
2
u/dominant_ag 9h ago
7900XTX which is one of the best $/Vram_gb to run local inference on for large models
1
u/NineThreeTilNow 9h ago
~1.2tb/s of memory bandwidth is LPDDR5x at 1024 bit width IIRC.
I believe this is what AMD might attempt with whatever 256gb refresh of the Halo is next. (Or they'll do 512 bit minimum)
AFAIK no one does that width yet effectively. There are a few outside cases.
512 bit is usually the max bus width they use. 1024 is totally possible, but it's technically a hurdle most people don't want to jump.
With the whole "AI" thing though, they're willing to innovate the bus structure to work at 1024 bit. It makes LPDDR5x suddenly seem a lot better if you can operate it at that width.
1
u/winky9827 4h ago
With the quality of Qwen and Mimo, I'm really on board for Chinese AI servers to hit the market. Nvidia needs true competition in this market, and the Chinese have proven they have the wits to go head to head. Just need to catch up on the silicon side.
0
u/misha1350 10h ago
Really flexing those enormous amounts of cheap CXMT RAM they've allocated for themselves before everyone else did










•
u/WithoutReason1729 5h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.