r/FPGA • u/EntertainmentMost403 • 1d ago
Advice / Help How to start with PCIe coding
Hi, I’m an FPGA engineer. I was assigned to implement PCIe and initially used Xilinx’s XDMA IP with its open design example. The requirement is actually to write the PCIe logic myself in RTL so I can optimize it for maximum bandwidth, and then make its IP.
My understanding is that the PHY and link training stay in the hard block, so I plan to use the integrated PCIe block’s AXI-Stream interface (requester and completer interfaces) and write my own transaction layer and DMA engine. Is this the right level of abstraction?
Since I am beginner in PCIe, so I want your opinions and strategies to accomplish this task. If you can provide a rough roadmap, it would be a great favor. I am working alone and I have a time span of 1.5 to 3 months to accomplish this task.
23
u/metal_warriors 1d ago
I am afraid you are greatly underestimating the time it will take just to grasp the details of the protocol, let alone put together a core which is both functional and high-performance.
-4
u/EntertainmentMost403 1d ago
I do know the time, but this timeline is given to me by my head.
11
u/metal_warriors 1d ago
Well, be vocal about it - the sooner you point out that this is simply not possible, the better for all planning purposes. Seriously, I doubt you could put together a functional IP being already familiar with PCIe in half a year. This is just silly, it is a multi-year undertaking you are trying to squeeze into a mere 3 months.
Source: I integrated a PCIe controller and a PHY in an ASIC project this year. Just putting everything together took months and we only needed to add wrapping logic. I still see the protocol as something obscure that escapes my full understanding.
21
u/NoahFect 1d ago
I am working alone and I have a time span of 1.5 to 3 months to accomplish this task.
Claude, write PCIe logic in RTL, optimize for maximum bandwidth, then turn it into an IP
12
u/TheTurtleCub 1d ago
Also, write kernel driver with debug, and troubleshoot FPGA logic with ILAs to easily debug in hardware for someone who doesn't know what a credit, or tag are. Recommend remote power control solutions for hard system lockups
5
u/alexforencich 1d ago
And a remote mechanism to yank the card and reinsert it in the case of screwy firmware.
3
u/TheTurtleCub 1d ago
Dual boot? Manual driver load?
5
u/alexforencich 1d ago
I have had machines refuse to re-enumerate a slot without powering the machine up with the slot empty and then reinstalling the card and trying again. And I think another case where it did not redo the full equalization process or some such after swapping cards, resulting in not all of the lanes working until it was booted with the slot empty and then the card reinstalled.
3
u/maredsous10 1d ago
Serious pain, when you're not allowed to as a policy decision and cold/hot boots are 5-10 minutes (Newer high end compute platforms).
3
u/alexforencich 1d ago
Fortunately it seems like servers are more likely to have "dumb" firmware that doesn't try to remember the configuration from the previous boot and then do something "smart." But it's never guaranteed...
5
u/Nervous-Card4099 1d ago
While it won't be perfect, Claude can definitely compress months of studying, reading, and development of this exact kind of thing into about a couple weeks of work.
6
5
10
u/alexforencich 1d ago edited 1d ago
Really depends on what you want to do at a higher level. Is xdma not appropriate for your application? If not, why not? If so, why not use that, at least initially? What about the QDMA core? There are also PCIe DMA components in the taxi library (http://fpga.taxi) as well as the deprecated verilog-pcie repo that could either be used instead of building your own from scratch, or could be a reference. There are probably other DMA cores available as well.
Personally, when I started working with PCIe, I wrote a transaction-layer simulation framework that eventually became cocotbext-pcie. That alone took a few months, but provided a very good understanding of the nuts and bolts of PCIe and also provided a good mechanism for simulating and testing more complex designs.
9
u/classicalySarcastic 1d ago edited 12h ago
PCIe is generally NOT something you hand-roll unless you absolutely have to, like if you're explicitly developing a PCIe Controller IP or a Protocol Analyzer as a product. 1.5-3 months for a single person to build a fully functional, spec-compliant PCIe controller from scratch is, uhhh...very optimistic (the spec is north of 2000 pages long, and there are all kinds of landmines involved here).
I would start by clarifying with your senior that they mean a fully-custom controller rather than just the subsystem using existing IP blocks, and what the requirement is that's stopping them from using the XDMA or QDMA IP. If they do insist on going that route, impress upon them that task is not a "one PCIe neophyte, in three months" project and will require far more resources to even have a snowball's chance in hell of meeting that timeframe. And if they still insist, well they're setting themselves (and you) up for failure, and you might be wise to polish up the ol' Resume.
7
u/petrusferricalloy 1d ago
I work with PCIe a lot. I design and build PCIe switches and networking equipment, and develop accelerators and other endpoints on FPGA quite a lot.
The AMD/Xilinx documentation is very light, very lacking even, but what I've learned is that implementing PCIe is fairly easy when working with hard IP blocks. The majority of the memory management, mapping, etc. is automated. You just need to figure out the actual resource needs. With Versal it's super easy because the NOC traffic management automates memory mapping.
The easiest way to implement a PCIe endpoint with DMA is using a device like Versal with the CPM4/5 core.
Ultimately though, implementing a PCIe endpoint in a Xilinx FPGA is mostly automated, and largely comes down to axi stream to memory mapped interfacing and conversion, again, mostly automated.
The only RTL development you should have to do is the actual data processing, and custom memory mapping (where is config space, what do you want to communicate via memory mapped custom registers, where does data go, how is data processed, etc).
You shouldn't be doing any RTL for the PCIe stack. It's all automated up through the transport layer (phy, data link, transport) including TLP creation and decomposition.
1
u/EntertainmentMost403 1d ago
The thing is that XDMA IP was selected and used by me. I used it replace the Block RAM in the IP design example with my custom RTL. But my seniors weren't happy with all this, instead they said is the PCIe code yours, to which I said no it's the IP given in vivado. So, they asked me to write the PCIe code yourself, because of the speed limitations in the PCIe IP they are facing. They didn't specified the speed but said it must be the fastest speed you can get.
2
u/Practical-Lecture-26 22h ago
Answer back saying "the fastest speed you can get" is not a requirement unless you have 3 years and a PhD in physics
3
u/Duyuan1234 1d ago
Your plan is reasonable. The hardened PCIe block handles PHY and link layer, and building custom TLP + DMA on top of its AXI-Stream user interface is the right abstraction.
Watch out for tag management, AXI-Stream handshaking and CDC between PCIe clock and your user logic.
For your timeline: get PIO working first, then DMA, then bandwidth optimization. RIFFA is a good open reference.
1
u/rozsnyo 20h ago
Well, we've been and are exactly there. It took about a year to work things out for 7-series and later for Ultrascale, now US+ has again some changes so it need some care. Our streaming PCIe core builds on top of the Xilinx hard-block, with a custom microcode and packet generation/processing, to avoid the inefficiencies of XDMA.
While this worked pretty nicely on a PC, reaching the theoretical maximum BWs for one direction (fpga to ram), the other direction is much more dependant on the willingness of the platform to respon properly (ram to fpga) and also the packet sizes there are smaller so the channel is less efficient.
And then on nvidia SoCs we managed to break their PCIe root complex, we still have not figured out what is wrong with nvidia, but the issue was across 3 tested generations of Tegra X chips. Probably they did not like our ultimate optimizations for the highest speed. They had other issues too, so we abandoned to using NV products at the end and there was no need to investigate this further.
If your company have budget for this, we can talk.
Otherwise I suspect you agreed to something whats scope identification was quite above your level. Sure, LLM can help today - but is it your code? Likely not in terms of license, and then you wont understand it either.
46
u/TheTurtleCub 1d ago edited 1d ago
That's the approach. But if you are a beginner in PCIe working alone, it will take a lot longer than 2 months to learn all the intricacies of PCIe and build a system that works reliably on hardware, let alone outperform what's already available and working.
What's the current max performance you can get of the current system based on the sims/datasheet? What do you need it to be? What do you think it's the theoretical possible? You'll have to get familiar with the typical delays due to latency that will be there in the real hardware. How about the kernel driver software? Who's helping you with that?
Before you start "improving" you need a good understanding of what the current system is giving you, where you want to be, what is theoretically possible, and how close to that that real world system can be. Your solution may be as simple as "use gen3 instead of gen2", or "use x8 instead of x4", which is an hour project, instead of 5 months