CapabilityCognitive automation2023-12-10
GPT-4 passed 43.5% of 156 human-written Verilog problems on the first try in VerilogEval, whose authors say it is confined to boilerplate code for small designs
Chip design engineeroccupation page →Event date / reported
2023-12-10
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Writing RTL
Writing the hardware description code (Verilog, VHDL) that describes each block's logic.
Being augmented≈ Platform inference
Where this applies
A benchmark of 156 problems from the Verilog instructional website HDLBits, checked for functional correctness by simulation against a golden solution. In the corrected v2 results GPT-4 passed 43.5% of the human-written problems on the first try and 58.9% within ten tries. The authors say their evaluations are confined to boilerplate code generation for relatively small-scale designs. Models of 2023 on instructional problems; the authors' employer designs chips.
What this means
Models can write working hardware code for a minority of textbook-style problems on the first try.
What it does not yet show
Short instructional problems with 2023 models; it does not measure real chip projects.
What you can check
Open arXiv 2309.07544 (VerilogEval) and read Table II in the v2 PDF.
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Liu et al. (NVIDIA) — VerilogEval: Evaluating Large Language Models for Verilog Code Generation, arXiv 2309.07544 (v2, 10 Dec 2023) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.