BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152639Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T170300
DTEND;TZID=America/Los_Angeles:20260728T170300
UID:dac_DAC 2026_sess306_WIP3170@linklings.com
SUMMARY:NPU-Based Energy-Efficient LLM Inference for Semiconductor Equipme
 nt Control Code Generation
DESCRIPTION:Sanghyeok Park, Seunghoo Hong, and Simon Woo (Sungkyunkwan Uni
 versity)\n\nIndustrial Algorithmic Pattern Generator (ALPG) code generatio
 n\nwith large language models (LLMs) has shown strong results under\nhigh-
 cost configurations using proprietary models and GPU-based\npipelines. How
 ever, such setups limit practical deployment in on-\npremise semiconductor
  environments. In this work, we ask how\nfar a cost-constrained pipeline c
 an approach industrial viability.\nWe deploy an FP8-quantized open-source 
 Llama 3.3 model on an\nemerging low-power NPU accelerator (Renegade) and e
 valuate\ndomain-specific code generation. Compared to an NVIDIA A100\nGPU,
  the NPU achieves comparable throughput (80–110 tokens/s)\nwhile reducing 
 steady-state power by 50–70%, yielding 2–3× higher\ntokens-per-watt effici
 ency. Structural validity improves from 40% to\n60–65% with lightweight ru
 le-based correction. These results high-\nlight the trade-off between ener
 gy efficiency and domain accuracy,\nsuggesting that energy-efficient NPUs 
 can approach practical indus-\ntrial deployment without relying on high-co
 st model and hardware\nconfigurations.\n\nTrack: Student\n\n
END:VEVENT
END:VCALENDAR
