Verilog DB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation (2025)

AUTHORS:

P. Calzada, Z. Ibnat, T. Rahman, K. Kandula, D. Lu, S. K. Saha, F. Farahmandi, and Mark Tehranipoor

As Large Language Models (LLMs) become increasingly important for hardware design automation, the quality and scale of training data are critical to generating reliable RTL code. VerilogDB introduces a robust dataset and preprocessing framework built to support LLM training and fine-tuning for Verilog and RTL generation. The researchers collected and validated more than 20,000 Verilog samples through an automated process that checks syntax, runs logic synthesis, and extracts relevant module metadata to ensure higher-quality data. The work provides a scalable foundation for advancing AI-driven hardware design and enabling more capable, reliable LLM applications throughout the RTL development process.

tags
RESEARCH PAPERS
Book a Demo!

Contact us today to request a demo for one of our products!

Demo Request Form