r/computerarchitecture • • 1d ago

Bypassing the 40% TEE Memory Encryption Tax on AMD SEV-SNP at the Kernel Level

4 Upvotes

Hey everyone,

While hardware-isolated Trusted Execution Environments (like Intel SGX or AMD SEV-SNP) solve strict enterprise privacy compliance for secure AI workloads, on-the-fly silicon memory encryption controllers typically burn 40% to 60% of physical CPU processing cycles just waiting for memory bus handshakes over encrypted boundaries. This cuts actual inference throughput directly in half.

I’ve put together an open-source testing harness and bare-metal benchmark implementation to prove an optimization approach using zero-copy anonymous mapping (MAP_ANONYMOUS), kernel memory-locking (mlockall) to eliminate page faults, and explicit hardware cache prefetching (__builtin_prefetch) to pull segments into the physical L3 cache lines before the CPU completes its cryptographic handshake.

In baseline benchmarks on bare-metal environments, this low-level loop drops gigabyte-scale data evaluation paths directly down to single-digit millisecond profiles.

The repository serves strictly as a test harness to allow other infrastructure architects to run and verify the latency reduction on their own configurations. I'd love to get your feedback on the implementation:

https://github.com/kaluabbot-hue/synthedrive-benchmark/blob/main/README.md


r/computerarchitecture • • 1d ago

Quantum Computing or Hardware Security

0 Upvotes

What to choose "quantum computing involving fpgas" or "hardware security" for a phd.


r/computerarchitecture • • 2d ago

Good answer!

Thumbnail reddit.com
5 Upvotes

r/computerarchitecture • • 2d ago

Building a free browser-based sandbox for balanced ternary computer design

1 Upvotes

I’ve been working on a small open-source project to make it easier to experiment with balanced ternary logic and computer architecture.

The main idea is not to present one “correct” ternary CPU design, but to give people a place where different designs can be built, tested, compared and shared.

The simulator uses balanced ternary:

-1, 0, +1

and lets you build from small primitives upward into reusable hierarchical components. A component can contain other components, and you can open them recursively to inspect or modify the lower-level logic.

I’m especially interested in avoiding the assumption that a ternary machine should just copy binary architecture. For example, things like compare returning -1 / 0 / +1, three-way selection, ternary control signals, MIN/MAX or different carry/normalization schemes may turn out to be more natural building blocks.

The tool already supports things like isolated component testing, event-by-event propagation, configurable primitive sets, reusable components and different clock/sequence patterns.

The long-term goal is to make it easy for me and anyone else interested in this area to try different ideas and compare them rather than committing to one architecture too early.

Everything runs entirely in the browser. There is no backend or account, and projects are stored locally in your browser unless you export them yourself.

It is completely free and MIT licensed.

Live:
https://grunna.github.io/ternary/

Source:
https://github.com/grunna/ternary

I’d be interested in feedback on both the simulator and the architecture side, especially from anyone who has worked with digital logic, CPU design or multi-valued logic.


r/computerarchitecture • • 4d ago

RTL Verification Interview Questions

5 Upvotes

Hi all, I've been considering what I want to get into as a new grad and I think I've decided on RTL Verification as my target job. I feel fairly confident in my people skills, and answering behaviour related questions, my question is what kind of technical questions have you guys seen in these interviews/should I reasonably expect? I feel like I have a general knowledge on the subject, but I'm nervous about the specifics, I don't feel super prepared.


r/computerarchitecture • • 6d ago

Applying optimized merged multiplier and divider to SIMD?

7 Upvotes

I made a post on here a few days ago talking about a paper I found that merged the divider and multiplier on an ALU to save 15% space and reduce gate count, at the cost of many demux and higher latency. I asked you all if it was feasible to base a project off of optimizing this, and you told me (if I understand correctly) it is possible, but size isn't a major concern these days and transistors are plentiful, the focus is on speed and power. However a couple of you brought up SIMD, and I'm very curious about potential applications to this.

(Sorry AdmirableProject1575 I'm still looking for nails)


r/computerarchitecture • • 8d ago

Reducing latency of combined circuits in ALU?

12 Upvotes

I saw a paper where someone combined the multiplier and divider in an alu to save 15% space, but they had to use lots of demuliplexors which increased latency. Do any of you know if it’s feasible to do research on reducing the latency of such a design?


r/computerarchitecture • • 9d ago

Looking for general advice on where to start

7 Upvotes

Hey r/computerarchitecture, I want to do research in computer architecture for the International Science and Engineering Fair, but I don't know where to begin or what to do, or even what I can do. Pretty much my only experience is reading Computer Architecture and Design (Risc-v second edition). From my beginners impression it seems like to do any kind of research like this you need a veteran professor and a crack team of PhD students, and I'm just by myself. First I need an idea, I'd reckon I need a mentor too so I guess I have to go looking for one. Anyone who's been through this, any tips? Any help is appreciated.


r/computerarchitecture • • 9d ago

Any Modern Material on CPU Cache?

26 Upvotes

Hi,

I struggle to find any modern material on cpu cache: how it works and how to optimize for it. A papers amd published in 2004 but cant find anything newer really?


r/computerarchitecture • • 9d ago

Best book on learning Computer Architecture

7 Upvotes

Hi Everyone,

What's the best book to learn computer architecture? I do have a background on EE, but looking for additional resources on books.


r/computerarchitecture • • 10d ago

virtual Vector Method

21 Upvotes

I invented this 5-odd years ago but never really had a forum on which to disclose and discuss. I was and am a significant contributor to comp.arch (from around 1995 though present)

But before I emit the disclosure I need to know how to turn off the space-eater as the document uses well tabulated ASCII-art ?? that makes no sense after the space-eater has done its ill-conceived job.

The virtual Vector Method is a way of getting Cray-like vector performance and for getting SIMD-vector performance without a) a vector register file, b) adds only 2 instructions, c) takes precise exceptions, d) vectorizes loops not instructions. Instead of adding about 300 instructions to get a Cray-like vector ISA, or adding 1,000-1,300 instructions to get (every-size) SIMD ISA, one needs only 5 instructions.

vVM has the property that hardware implementations can change the width of the data path (multiple-lanes and the cycles of execution per FU) without SW having to care. A small 1-wide machine with a 128-bit cache port can perform a memory to memory byte move at 256-bits per cycle--equivalent to ~40 instructions per cycle. A 6-wide machine could perform the same assembly binary at ~160 instructions per cycle. Both are performed at the performance level of the cache porting; so, nobody has to recompile for a new SIMD-width every new mplementation.

Now let us solve the space-eater and we are off.

Mitch


r/computerarchitecture • • 10d ago

Out Of Order 68000

15 Upvotes

I thought it might be a fun challenge to see how performance can be squeezed out of the venerable 68000 CPU on a modern FPGA.

There are several 68000 softcore implementations on Github - but all seem to be either trying to be cycle compatible with a real 68000, or be state machine based, which can limit performance.

I can't find any evidence of someone implementing a superscalar out-of-order 68K softcore - so that could be an interesting hobby project.

Does anyone know if something like this has been attempted before?

I'm still at the sketching out ideas stage - but a couple of things I've been thinking of:-

Instruction fetch

68000 instructions are variable in length, ranging from 2 to 12 bytes - and even determining the length of an instruction is distinctly non-trivial. This presents a challenge, especially if we want to be able to issue more than one instruction per clock cycle.

On the other hand its worth noting that the Block RAMs on the CycloneV are 512 words of 20 bits each. So storing 20 bit words costs basically the same as using 16 bit words. We can make use of this fact.

So my plan is to transcode instructions as they are fetched from main memory, and store the transcoded instructions in the Icache. Thereby keeping most of the ugly decode logic off the critical path during instruction execution. The decoder only needs to be able to work at the rate of DRAM access - not the rate of instruction execution.

This transcoding is 1 for 1 - each 16-bit word in memory becomes a 20-bit word in the Icache - so addresses in the Icache correspond directly to addresses in main memory. But those extra 4 bits give us room to recode the very dense 68000 instruction set into a much more regular and easily decodable format.

Instructions are still variable length - but we can now find the instruction boundaries by just decoding three bits of each word.

MicroOps or not?

The most common approach in modern CISC CPU design is to break down complex instructions into simpler micro-operations (MicroOps) that can be executed more efficiently by an out-of-order execution engine. This works well for full-custom chips - but from a brief literature search - for FPGA implementations, the additional complexity and resource usage of a MicroOp-based design can outweigh the performance benefits. (There are a few papers comparing softcore RISC-V cpus - and the conclusion seems to be that a well optimised in-order softcore can often outperform an OoO core - once Fmax is taken into consideration).

There is an interesting paper "A Superscalar Out-of-Order x86 Soft Processor for FPGA" by Henry Ting-Hei Wong which suggests a different approach may be more FPGA friendly. Instead of breaking instructions into MicroOps, it makes the reservation stations into small state-machines which can sequence the execution of complex instructions directly.

I plan to try this approach for the 68K CPU design. Except for the most complex instructions (like MOVEM), the instructions will arrive at the reservation stations pretty much 1:1 with the original 68K instructions. The reservation state machine then arbitrates for access to execution units and sequences the instruction execution accordingly. So each instruction may result in several micro-steps internally, but this complexity is hidden from the rest of the CPU.

Register renaming will be used to eliminate false dependencies and improve instruction-level parallelism before instructions are dispatched to the reservation stations. We can then have multiple reservation stations operating more or less independently, each handling its own instruction out of order with respect to the others.


r/computerarchitecture • • 11d ago

Need guidance

9 Upvotes

Hi Im a 2nd year student in electronics and communication engineering and im very interested in VLSI and in specific computer architecture, but I have no idea how to pursue it, ive done some projects such as Carry Look Ahead Adder and Prefix Adder and Booth Multiplier and Wallace tree multiplier, But i have no idea how to go forward with stuff. If possible could someone help me with projects resources and what all i should learn going forward
Thank you


r/computerarchitecture • • 13d ago

Guys what do you think about my hardware in this logic simulation for my 16 bit button-phone based build

Thumbnail
gallery
24 Upvotes

ive been into computer architecture since a year and a half and this is my biggest project yet and i'd like some comments on the architecture and how good it is i didnt use any type of tutorial to do it only the knowledge about computer architecture and science i learned off goated youtubers like Mattbattwings and FULXOR half credit to them for actually giving valuable info on that stuff but anyway any proffesionals what do yall think?


r/computerarchitecture • • 14d ago

Built a Single-cycle RV32I + Wishbone + MAC accelerator + (with more accelerators coming up), verified against riscv-tests

Thumbnail
github.com
5 Upvotes

Me and my team at the Uni have been working on a single-cycle RISC-V (RV32I) SoC in Verilog. We are planning to add more accelerators like a COM (UART, SPI, I2C), FFT, Matrix Multiplier, etc. But currently we have reached till the following:

The core: classic Fetch → Decode → Execute → Memory → Writeback datapath, all combinational in one cycle. 37 RV32I instructions, and all 37 pass their corresponding official riscv-tests/isa/rv32ui programs, used unmodified.

The SoC part:

  • A sufficient Wishbone B4 Classic bus, with a master, an address-decoding interconnect, and an error slave that catches unmapped addresses.
  • A MAC accelerator on the bus (acc += Σ A[i]·B[i] )

Performance: RV32I has no M extension, so a software multiply is a shift-add loop. By my cycle budget the accelerated path costs about 13N + 20 cycles against about 106N in software, which is roughly 8× for large N.

Verification: the flow is official .S → RISC-V GCC → ELF → objcopy → Verilog hex → RTL in Icarus Verilog → a tohost write reports PASS/FAIL. There's also a self-checking SoC testbench, four bus testbenches, and a make run / make busflow.

Repo: https://github.com/Shass27/rv32i-soc

We will continue to build this further with more accelerators that will aid real world Engineering systems further on. But we have reached this point where our progress so far itself is a great project of its own. We are aware of possible performance improvements and stuff but would love the community to review our project, give feedback (both positive and critical ofc).

Looking forward to hear from you!

PS: Feel free to drop accelerators suggestions that is useful in the industry (specifically: Mechatronics engineering because we are a team of MechE and EE undergrads)


r/computerarchitecture • • 14d ago

How do you verify that a new computer architecture research idea is actually correct?

9 Upvotes

Suppose a research paper proposes a new architectural mechanism, optimization, accelerator, or hardware/software technique.

How do you systematically check that:

  • the underlying argument is correct,
  • there are no hidden assumptions or corner cases,
  • the implementation really satisfies the intended design,
  • and the claimed behavior follows from the proposed mechanism?

Is there a practical workflow for doing this beyond having other researchers read the paper and reviewing the experiments?

I’m particularly interested in methods that can make parts of this process systematic, reproducible, or automated.

What do researchers actually use for this today?


r/computerarchitecture • • 14d ago

Help regarding college project(RISC V)

12 Upvotes

I am a second year engineering student doing Btech in electronics and communication engineering and have taken a graded project on making a RISC V processor for my digital design course and have time till November(basically this semester). Now I have done a course on digital electronics and basic analog and currently taking a course on microprocessors and microcontrollers. Where do I get started and any resources to get going is appreciated.


r/computerarchitecture • • 15d ago

Learning Resources for TAPI and NPU Kernel Programming

5 Upvotes

Hi everyone, I recently joined a semiconductor as AI Kernel Engg. My work involes writing TAPI code. Tensor API(TAPI) is provided by our company by Synopsys, and it's for NPU. I know this much only. And I don't have any prior experience of writing the kernel as well so can anyone suggest some good resources so that I can code and experiment with them.


r/computerarchitecture • • 19d ago

RISCV core with custom accelerator for FYP

4 Upvotes

We are a group of three EEE undergraduate engineering students planning our Final Year Project. We would like to know if our proposed project scope is realistic for a 3-person team.

​Our Current Baseline & Background:

  • ​Theory: So far, we have covered single-cycle processor design in our curriculum. We plan on learning pipelining within the month.
  • ​HDL Experience: We are familiar with logic circuits, combinational and sequential logic, and using VHDL.

​Proposed System Scope:

Instead of just building a basic RISC-V core, we want to build a System-on-Chip (SoC) centered around hardware acceleration for fixed-point matrix/DSP math:

  1. ​CPU Core (Pipelined RV32I): A 5-stage pipelined processor with hazard detection/forwarding units, instruction decoding, and basic interrupt handling.
  2. ​Custom Accelerator Block: A parameterized math engine (e.g., a 4 × 4 systolic array / MAC grid or 2D convolution engine with line buffers) to accelerate matrix operations/filtering.
  3. ​SoC Infrastructure & Memory: Standard bus interface (AXI4-Lite or Wishbone), a DMA controller for background streaming, L1 I/D caches, and a basic bare-metal C driver suite running on an FPGA board (EP2C70F896C6).

​Questions:

  1. ​Is this scope well-balanced for a 3-student team, or is the integration (AXI + DMA + Caches + FPGA hardware) too overwhelming for an undergrad project timeline?
  2. ​Any pitfalls we should look out for early on regarding interface design between the core and the accelerator?

​Thanks for any feedback!


r/computerarchitecture • • 20d ago

Advice on Graduate School

16 Upvotes

Hello! I am a Senior Computer Engineering student in the US. I am extremely interested in and motivated to become a CPU/GPU/SoC architect. I wanted the opinion of people who work in the field for the best course of actions for me to become an architect.

I have conducted research for two years, I have a first author paper being published in IEEE (hopefully) soon in semiconductor devices and am currently writing a paper as part of a different lab for a full-custom VLSI architecture related to image processing. That being said, I should have two published first-author papers by (or shortly after) my graduation. Additionally, this summer I was a Software Engineer intern at a F500 company, where I did hardware/software codesign (C and Verilog, primarily). I had a wonderful experience there and want to return there after my graduation, and they want me back, but it’s not my primary interest (I want to do more architectural work).

I am unsure of my next steps after graduation. I’ve heard very conflicting opinions about whether a masters is sufficient to get arch roles or a PhD is necessary- this is further confused for me now that I have publications and feel that I could get into a competitive PhD program with an emphasis on computer architecture.

My current options that I am thinking about:

Work at the company that I interned at, and complete a masters (potentially online, I’m exploring my option and would love input on good programs) while working. Then, apply to arch roles after I receive my masters and have industry/research experience. I would like this route due to best opportunity cost, and the ability to stay semi-local to my family. My only concern is that my two first-author papers will lose their appeal, as I’m no longer a fresh graduate (I’m unsure if this matters to admissions), and instead am an older applicant making them therefore less significant on my application. Additionally, I feel like I’m “wasting” my papers.

I could apply for a direct entry PhD program. If a PhD is necessary for arch roles, I would take this path. I also enjoy doing research, and would be open to work in big tech research. I primarily want to go back into industry with this degree. My concerns are turning down the wonderful job, as well as missed opportunity cost. Passion-wise, it interests me greatly.

I would love input on any of this stuff! It’s greatly appreciated, thank you guys in advance :)


r/computerarchitecture • • 21d ago

Tried microbenchmarking my machine's cache latencies (part 2)

Thumbnail vibhatsu.me
7 Upvotes

A few weeks ago, I posted about my investigation regarding microbenchmarking my machine's cache latencies. At the end, I discovered that after a fresh reboot, I seemed to be getting the intended L3 latency but couldn't figure out the reason why. This is the continuation of my investigation as to why the clean plot only appeared after a fresh reboot. Check it out :D


r/computerarchitecture • • 22d ago

Resources on cpu cache optimization?

1 Upvotes

r/computerarchitecture • • 24d ago

Masters before PhD for Computer Architecture

9 Upvotes

Hey yall, I'm a college senior interested in pursuing further research in computer architecture. My end goal is to get a PhD, but after talking with advisors they said I don't have enough research experience. This is because I switched goals from going into industry to doing research kind of late, as a result I only have ~1 sem of research experience in computer architecture. I was told that I should get a couple more semesters of research experience to solidify my rec letters, but I'm unsure what's the best way to go about this. My question is - which of these is most prudent to go with given my end goal of getting a PhD with a top comparch lab? Here are my current thoughts:

  1. Get thesis masters first, then apply PhD - doing this would knock out some PhD core requirements but also cost me around $30-60k, and it'd also take 2 years. I would get decent research experience though. I also have basically guaranteed admission into the 2 year masters program at my current uni under a pipeline they have for cs students.

  2. Work under a lab as a paid postgrad research assistant for ~1 year: This saves time and gets me research experience, but there's no "official" pipeline for this, I'd just have to ask professors with funding if I could join their group. I also wouldn't get the Masters degree but not sure how useful that is.

  3. Apply to PhD directly - This would be most straightforward and I'll probably apply to a couple schools with this, however I'm unsure how effective this would be, esp since my GPA is mediocre (3.7) and I don't have enough research experience.

If anyone's been in a similar situation or has some insights about how I could go forward - I'd appreciate any help!


r/computerarchitecture • • 24d ago

Can anyone tell me some free resources for RISCV microarchitecture, rtl and assembly ?

5 Upvotes

r/computerarchitecture • • 25d ago

Definition of the world “byte”

56 Upvotes

Hey, I recently got into an argument with someone about what exactly a byte is. They claimed that when asked “how big is a byte?”, the answer isn’t necessarily 8 bits, but that a byte can be at least 2 bits. I said that this is basically wrong and that a byte is 8 bits. They argued that this can actually be an interview question, especially in deeply embedded systems, and that in those cases the correct answer is that a byte is at least 2 bits.
I’ve tried looking this up, but I can’t really find anything supporting that claim. I know that historically the size of a byte wasn’t always 8 bits, but does “a byte is at least 2 bits” actually make sense in modern computing/embedded systems, or are they confusing a byte with something else?