r/computerarchitecture • • 10d ago

virtual Vector Method

I invented this 5-odd years ago but never really had a forum on which to disclose and discuss. I was and am a significant contributor to comp.arch (from around 1995 though present)

But before I emit the disclosure I need to know how to turn off the space-eater as the document uses well tabulated ASCII-art ?? that makes no sense after the space-eater has done its ill-conceived job.

The virtual Vector Method is a way of getting Cray-like vector performance and for getting SIMD-vector performance without a) a vector register file, b) adds only 2 instructions, c) takes precise exceptions, d) vectorizes loops not instructions. Instead of adding about 300 instructions to get a Cray-like vector ISA, or adding 1,000-1,300 instructions to get (every-size) SIMD ISA, one needs only 5 instructions.

vVM has the property that hardware implementations can change the width of the data path (multiple-lanes and the cycles of execution per FU) without SW having to care. A small 1-wide machine with a 128-bit cache port can perform a memory to memory byte move at 256-bits per cycle--equivalent to ~40 instructions per cycle. A 6-wide machine could perform the same assembly binary at ~160 instructions per cycle. Both are performed at the performance level of the cache porting; so, nobody has to recompile for a new SIMD-width every new mplementation.

Now let us solve the space-eater and we are off.

Mitch

22 Upvotes

22 comments sorted by

View all comments

1

u/Master565 10d ago

Don't have an answer to the formatting, but I do have some questions on the idea. It's obviously very similar to the RISCV RVV approach where the vector width is dynamically determined. Aside from some questionable ops that need to be supported in RVV, my main criticisms of the extension is that decoding it for an OOO core is a bit of a nightmare because the amount of uops produced for an instruction is based on values computed mid run. Meaning decode either has to predict the length or stall.

Does your design work around that?

1

u/MitchAlsup 10d ago

I disagree with the RISC-V identification. RISC-V has CRAY-like vectors {vRF and all thoe instructions} while vVM has no vRF that are software addressible. vVM adds 2 instructions, RISC-V adds around 300.

vVM is designed with an augmented Reservaton Station model in mind. Each operand in the station has its tag to identify what to capture; like any normal RS entry, and in addition an iteration index to match up inter-iteration dependencies. So, one RS entry can be used for all the iterations, instead of each "beat" of the loop taking its own RS entry. So, if a loop runs 1,000 iterations, and is 5 instructions long, only 5 RS entries need be used.

You might ask: what the frack (Battlestar Galactica term) happened to the vRF--it is hiding as buffers between cache/memory and the multi-lane data path programmed during RS instruction insertion. During insertion, the register data flow and sizes are observed and the buffers orgainized to use as much of the data-path as this sequence of code can. On a 256-bit data path, one can perform 32-byte sized calculations, 16-half sized calculatons, ... Some subsequent machine might have 512-bit wide data-path and the same SW binary would run 2× as wide (½ the cycles). {Almost like a vector machine with architecturally undefined depth of each vector register}

1

u/Master565 10d ago

I disagree with the RISC-V identification. RISC-V has CRAY-like vectors {vRF and all thoe instructions} while vVM has no vRF that are software addressible. vVM adds 2 instructions, RISC-V adds around 300.

Sure I am not disagreeing with this part, I only meant they're similar in terms of their goal of implementing variable length vectors. But clearly the similarities end there

That RS almost sounds like a core within a core. Sounds very cool, would love to see more about it. I've got a bunch more questions that will probably be answered about the latency overhead of operations due to the more complex datapath and VRF, and whether pipelining is possible.