At a glance
| Metric | Value | Provenance |
|---|---|---|
| Standard / role | FIPS-197 AES-256 + NIST SP 800-38D GCM, AEAD (encrypt+decrypt) | [FIPS-197] / [SP-800-38D] |
| Cipher | AES-256 encryption only (GCM never inverts AES) | [RTL] |
| External stream widths | 8 / 32 / 128 / 512 / 1024 / 1536 / 3072-bit | [RTL] |
| Peak throughput span | 0.080 Gbps (4xs) → 1.0752 Tbps (2xl, aggregate) | post-route |
| Maturity | Technical preview; production qualification in progress | this note |
Supported profile
| Cipher / mode | AES-256-GCM AEAD |
| Operations | Authenticated encryption and decryption |
| Key | 256 bits |
| IV | 96 bits only |
| Authentication tag | 128 bits only; truncated-tag verification is not supported |
| AAD alignment | 2xs supports byte-partial final AAD for AAD-only messages; use word-aligned AAD when payload follows. Other cores support byte-granular AAD |
| Clock / reset | Single clock; synchronous active-low reset |
| Interfaces | Core-specific valid/ready streams; detailed port maps to follow |
The family
AES-256-GCM is a family of nine cores across five
datapath designs, grouped into three classes. All share the
same CTR/GCM semantics (FIPS-197 + NIST SP 800-38D) and many
common building blocks.
| Class | Datapath designs | SKUs | Optimization |
|---|---|---|---|
| Serial | 1-, 8-, and 32-bit serial AES-256 forward cipher | 4xs 3xs 2xs | ASIC area first; BRAM-free on FPGA |
| Iterative | 128-bit iterative AES, on-the-fly key expansion | xs | smallest streaming core (~16k LUT) |
| Pipelined | 1–24 parallel 128-bit lanes across one to six packet streams, fully pipelined | s m l xl 2xl | throughput first (up to a peak aggregate 1.075 Tbps) |
The three serial cores trade datapath width for area: narrower
designs target fewer logic resources but take more cycles per
block. On the U55C, fixed overhead leaves 4xs and
3xs at essentially the same FPGA footprint;
4xs is retained for its sky130 target.
Post-route results
Preliminary post-route results. All nine cores were placed and routed on the same
Alveo U55C target and meet timing at the
listed clocks
(xcu55c-fsvh2892-2L-e, Virtex UltraScale+ HBM,
1.3M LUTs).
| SKU | Bus | Clock (MHz) | LUTs | FFs | BRAM | Peak throughput (Gbps) |
|---|---|---|---|---|---|---|
4xs | 8b | 400 | 3,379 | 3,193 | 0 | 0.080 |
3xs | 8b | 400 | 3,377 | 3,194 | 0 | 0.139 |
2xs | 32b | 385 | 7,163 | 7,343 | 0 | 0.261 |
xs | 128b | 350 | 15,549 | 9,321 | 12 | 3.0 |
s | 128b | 350 | 16,562 | 9,453 | 17 | 44.8 |
m | 512b | 350 | 37,564 | 24,371 | 22.5 | 179.2 |
l | 1024b | 350 | 69,335 | 44,867 | 28.5 | 358.4 |
xl | 1536b | 350 | 100,774 | 65,025 | 34.5 | 537.6 |
2xl | 3072b | 350 | 208,411 | 177,836 | 52.5 | 1,075.2 |
- Serial cores are BRAM free and DSP free. They use pure LUT/FF fabric.
- Pipelined cores (
s–2xl) sustain one AES block per lane per cycle once full, so peak core throughput = bus width × clock. - The
l,xl, and2xlfigures aggregate two, three, and six independent 512-bit packet streams. Listed peaks require all streams active, at least 23 active contexts per packet engine (25 for2xl), and packets of at least 896 bytes. xsuses iterative AES (15 cycles/block), so its crypto throughput is 3.0 Gbps.
Power estimates
Post-route estimates use default toggle rates; no workload VCD was supplied. Treat these as comparative, medium-confidence figures. Device-static power, approximately 3.28 W in these estimates, dominates the smaller cores.
| SKU | Clock (MHz) | Total (W) | Dynamic (W) | Static (W) | Est. W/Gbps |
|---|---|---|---|---|---|
4xs | 400 | 3.385 | 0.111 | 3.275 | 42.3 |
3xs | 400 | 3.386 | 0.111 | 3.275 | 24.4 |
2xs | 385 | 3.439 | 0.163 | 3.276 | 13.2 |
xs | 350 | 3.764 | 0.482 | 3.283 | 1.3 |
s | 350 | 3.824 | 0.540 | 3.284 | 0.09 |
m | 350 | 4.641 | 1.340 | 3.301 | 0.03 |
l | 350 | 5.485 | 2.166 | 3.319 | 0.02 |
xl | 350 | 6.290 | 2.954 | 3.336 | 0.01 |
2xl | 350 | 9.037 | 5.641 | 3.396 | 0.008 |
Design space
Peak core throughput versus U55C LUT count on log axes, with each core colored by estimated total power. The pipelined designs scale throughput by adding AES lanes and independent packet engines.
Verification
Verification combines FIPS-197 and NIST AES-GCM vectors, deterministic software-reference equivalence checks, and randomized multi-stream soaks. Coverage includes encryption, authenticated decryption, and invalid-tag behavior across the core family.
Formal verification covers GF(2^128) arithmetic and selected key-management and scheduling properties. Additional functional and formal coverage will continue through production qualification.