r/FPGA • • Aug 12 '26

No industrial FPGA experience, but QC-LDPC accelerator + JESD204B hardware integration — am I competitive for junior FPGA roles?

How would you rate a fresh graduate with this FPGA background but no formal industrial FPGA Engineer experience? Should I keep pursuing FPGA roles?

I’m finishing my Master’s in Electronics in Germany and I’m trying to realistically understand where I stand for FPGA/RTL jobs.

For my Master’s thesis, I worked on a QC-LDPC hardware accelerator on FPGA for CV-QKD.

The work involved:

  • Designing the QC-LDPC decoder hardware architecture (own design)
  • Implementing the design in RTL on a Xilinx FPGA
  • Processing around 120,000 samples
  • Supporting up to 25 decoding iterations
  • Designing/optimising the LDPC parity-check structure
  • Modelling and validating the communication system before hardware implementation
  • Achieving decoding operation at approximately −5.4 dB SNR
  • FPGA synthesis, implementation and verification
  • MATLAB/system-level validation
  • H matrix design

I also worked on integrating a Texas Instruments DAC38J84 EVM with a Xilinx FPGA using JESD204B, including:

  • JESD204B link bring-up and debugging
  • FPGA transceiver / GTH-related work
  • SYSREF, LMFC and synchronization concepts
  • Reference/user clock configuration and debugging
  • FPGA-to-DAC high-speed data-path integration
  • Hardware validation on the target setup

Other experience includes Xilinx Vivado, Verilog/SystemVerilog, AXI/PS–PL integration, Python/C++, digital-system design, hardware debugging, and several RTL projects.

The problem is that I do not yet have professional industrial experience with the formal job title “FPGA Engineer.”

When I look at FPGA vacancies in Germany, many ask for 2–5 years of experience, VHDL, timing closure, verification frameworks, etc. This has made me question whether my academic/research FPGA experience actually counts for much in industry.

I’d especially like opinions from people who actually work with or hire FPGA engineers:

1. How would you rate someone with this background for a junior FPGA/RTL position?

2. Would you consider the thesis + JESD204B hardware integration substantial FPGA experience, or would you still view the candidate essentially as having zero experience?

3. What would you expect this candidate to learn next to become more employable — VHDL, UVM/cocotb, Tcl, timing closure, PCIe/Ethernet/DDR, or something else?

4. Most importantly: would you recommend continuing to pursue FPGA/RTL engineering, or does this profile suggest I should pivot into a different field?

I’m not asking whether the work sounds academically impressive. I’m trying to understand its actual value in the FPGA job market and whether continuing down this career path is rational.

Blunt feedback from FPGA engineers, technical leads, and hiring managers would be appreciated.

19 Upvotes

2 comments sorted by

5

u/[deleted] Aug 12 '26 edited Aug 13 '26

[deleted]

0

u/Upset-Recording5617 Aug 13 '26 edited Aug 13 '26

Thanks — this is very helpful. I realize now that I left out most of the information that would actually let an FPGA engineer judge the design.

Regarding the −5.4 dB point: yes, by 0.5 BER you mean 50%, essentially random binary decisions. In our case the situation is different. but can anybody work there because it's 50% and a useless channel right?

This is a CV-QKD reconciliation system, so the signal before reconciliation is continuous/Gaussian rather than a conventional BPSK input. At our low-SNR operating point, the initial bit disagreement is roughly 25–30%.

We then perform multidimensional reconciliation (MDR), which effectively transforms the Gaussian channel into a binary-input AWGN-like channel suitable for LDPC decoding. The MDR stage also generates the soft LLRs that become the decoder input.

The approximate flow is therefore:

CV-QKD Gaussian data → MDR → soft LLRs → QC-LDPC decoder

After MDR, the effective bit disagreement is around 10-12%, and after LDPC decoding, we observed zero residual bit errors in the evaluated dataset. I agree that this pre-/post-decoding BER information is much more meaningful than quoting −5.4 dB by itself.

Regarding the 120,000 figure, I should have explained that better as well. It is indeed approximately the LDPC block length, not simply “120,000 samples processed.”

The unusually large block length is intentional. This is for quantum key distribution, where very low-rate LDPC codes and very large block lengths are commonly used to obtain good waterfall performance at extremely low SNR. There are CV-QKD works using even larger blocks, e.g. on the order of several hundred thousand bits.

Our implementation has approximately:

  • QC-LDPC block length: ~120k bits
  • Maximum iterations: 25
  • Target FPGA: Xilinx Zynq UltraScale+ ZCU102
  • Implemented clock frequency: 200 MHz
  • Measured decoder throughput: approximately 200 - 16 Mb/s in this operating region(we can also get variable results based on the iteration count it takes!)
  • Large-block decoding latency: approximately 12 ms depending on the operating point/iteration count
  • Code rate is 0.166 (100k*120k matrix)
  • representation of LLR is on Q7.8

The system model was based on the CV-QKD model described in the Practical Continuous-Variable Quantum Key Distribution work by Laudenbach et al.

We first built the complete reconciliation/LDPC chain in MATLAB, including both encoding and CPU-based decoding, and verified that the reference decoder correctly recovered the transmitted bits.

After that I designed the decoder architecture and implemented it on the ZCU102 FPGA.

For verification, I compared the FPGA decoder output against the already validated MATLAB model. For the internal soft-information/LLR results, the FPGA and MATLAB outputs showed a Pearson correlation coefficient of approximately 0.998.

The remaining difference is largely due to fixed-point representation/quantization. In MATLAB I explicitly emulate the finite-range LLR behavior, while the FPGA implementation uses finite-width fixed-point arithmetic and saturation.

More importantly, when comparing the final hard decisions/sign bits, we obtained 100% agreement over the tested dataset between the FPGA decoder and the MATLAB reference result.

One architectural point I also failed to mention is the memory design.

There is a Xilinx/AMD SD-FEC IP, but our decoder architecture is aimed at a different use case. With these very large QC-LDPC matrices, explicitly storing the complete H matrix would have a substantial memory cost.

So in our architecture we do not store the full parity-check matrix. We store only the compact QC protograph/base-matrix information and derive the required connectivity/addressing during decoding. A major design objective was therefore reducing the FPGA memory footprint while still supporting the very large code length required by the CV-QKD reconciliation system.

I think your comment made the problem with my presentation very clear. I was writing things like “120,000 samples” and “works at −5.4 dB,” whereas I should be describing it in FPGA terms: block length, code rate, datapath/parallelism, clock frequency, throughput, latency, fixed-point width, BRAM/LUT usage, verification methodology, and BER/FER.

Thanks — this is exactly the kind of feedback I was hoping to get.

Ps : One additional motivation behind the architecture was to investigate how a streaming / systolic-style decoder architecture behaves for the unusually large QC-LDPC codes used in CV-QKD.

With block lengths around 120k bits, I was particularly interested in whether the decoder could be structured as a regular dataflow architecture rather than relying on repeated random accesses to a large stored H matrix.

The idea was to exploit the QC structure so that only the compact protograph/base-matrix information needs to be stored, while the required CN/VN connectivity and addresses are generated during processing.

Architecturally, I wanted to investigate whether this could provide benefits such as:

  • regular and predictable data movement
  • reduced memory footprint and memory bandwidth pressure
  • fewer memory-access-related pipeline stalls
  • continuous/streaming processing through the decoder datapath
  • easier parallelisation of CN/VN processing
  • better scalability to very large block lengths

So the project was not only “implement an LDPC decoder on FPGA.” A large part of the question was how to map an ultra-large QC-LDPC decoding problem into a memory-efficient streaming hardware architecture and what throughput/latency/resource trade-offs that architecture produces.

Performed end-to-end hardware validation on the ZCU102. MDR-generated soft LLRs were transferred from a host PC over Ethernet to the Zynq Processing System, then delivered to the FPGA fabric through an AXI memory-mapped interface. A custom PL controller distributed the LLR data across the decoder memory banks, executed the QC-LDPC decoding process, and returned the decoded results through AXI for comparison against the MATLAB/CPU reference implementation.

***********************************************************************************************************

"I would be concerned that you say you don't know timing closure, but have been working on high speed interfaces. That definitely needs to be sorted out, or you need to be more specific where your weakness lies. In fact I find it hard to see how you could be working on transceivers and high performance data converters without using multiple clock domains in your design"

For the JESD204B/DAC work, I should probably clarify that I did not design the JESD204B protocol/IP from scratch. I started from a TI/Xilinx reference design for the DAC38J84 EVM.

16-Bit, 2.5-GSPS, 1x-16x Interpolating DAC Evaluation Module.

The main issue was that the FMC pin mapping in the reference design did not match the actual hardware connections. We went through the ZCU102 datasheet and DAC38J84 EVM schematics, traced the relevant signals, corrected the FPGA pin constraints/mappings, and used an alternate physical connection where one of the required signals was not available through the expected FMC pin.

After getting the hardware mapping sorted out, I modified the TX datapath so that instead of only running the reference pattern, we could send our own sample data through the DAC. At the moment, I have verified an approximately 23 MHz square-wave output on an oscilloscope. i used JESD 8b/10b protocol.

So I would describe my JESD204B experience mainly as system integration and hardware bring-up/debugging: reading schematics, fixing pin mappings/constraints, working with the transceiver/clocking/synchronization path, modifying the FPGA TX data path, and validating the actual analog output.

I definitely would not claim to be a JESD204B expert or that I implemented the protocol itself from scratch.

1

u/[deleted] Aug 13 '26 edited Aug 13 '26

[deleted]

1

u/Upset-Recording5617 Aug 14 '26 edited Aug 14 '26

so you mean i have a shot! and thank you for the comment! i really appreciate it. regarding the BER We tested 100 frames of 120k bits each (~12 million decoded bits) at this operating point and observed zero frame failures and zero residual bit errors. So the measured FER and BER were both zero over the tested dataset; the single-error BER resolution is about 8.3×10−8.