CapabilityCognitive automation2023-08-10
On RTLLM's 30 larger design tasks, GPT-4 produced 15 functionally correct designs, counting a design as correct if one of five attempts passed its testbench
Chip design engineeroccupation page →Event date / reported
2023-08-10
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Writing RTL
Writing the hardware description code (Verilog, VHDL) that describes each block's logic.
Being augmented≈ Platform inference
Where this applies
An academic benchmark of 30 designs, built because earlier tests used designs that were relatively simple, small and proposed by their own authors. It scores syntax, functionality and design quality. GPT-4 achieved 81% correct syntax and 15 of 30 correct functionalities, where a design counts as functionally correct if any of five attempts passes the testbench. Models of 2023; it does not measure use in chip projects.
What this means
On larger designs, a model working alone gets about half right even with five tries each.
What it does not yet show
A 2023 benchmark of 30 designs; it does not measure power, performance and area in real projects.
What you can check
Open arXiv 2308.05345 (RTLLM) and find "81% correct syntax and 15/30 correct functionalities" in the PDF.
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Lu, Liu, Zhang, Xie (HKUST) — RTLLM: An Open-Source Benchmark for Design RTL Generation with Large Language Model, arXiv 2308.05345 (v1, 10 Aug 2023) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.