Close

Presentation

Notsotiny: A Large, Living Benchmark for RTL Code Generation
DescriptionLLMs have shown early promise in generating RTL code, yet evaluating their capabilities in realistic setups remains a challenge. So far, RTL benchmarks have been limited in scale, skewed toward trivial designs, offering minimal verification rigor, and vulnerable to data contamination. In this paper, we introduce NotSoTiny, a benchmark that assesses LLM on structurally rich and context-aware RTL generation tasks, while being resilient to contamination. Built from hundreds of representative hardware designs produced by the TinyTapeout community, our automated pipeline removes duplicates, verifies correctness using simulation and formal equivalence checking, and continuously incorporates new designs to mitigate leakage. Evaluation results show that these tasks are significantly more challenging than prior benchmarks, emphasizing NotSoTiny's effectiveness in revealing the current limitations of LLMs applied to hardware design and in guiding the refinement of this promising technology.