ICML 2026 · Execution-based benchmark

Rule2DRC Leaderboard

Synthesizing KLayout DRC scripts from natural-language chip design rules, graded by execution on held-out layouts. 1,000 tasks, 13,921 evaluation layouts.

Public set · 1,000 tasks Private set · 100 tasks
Docs in context No context Docs in context No context
# Model
GPT-OSS-120B paper reasoning effort medium; no-context is a single run via OpenRouter · 2026-05
GPT-OSS-20B paper reasoning effort medium; no-context is a single run via OpenRouter · 2026-05
Qwen3-30B-A3B-Instruct-2507 paper instruct (non-reasoning); no-context is a single run via OpenRouter · 2026-05
Qwen3-235B-A22B-Thinking-2507 thinking · 2026-07
Qwen3-235B-A22B-Instruct-2507 instruct (non-reasoning) · 2026-07
Qwen3-30B-A3B-Thinking-2507 thinking; with-docs output capped at 9.8K (provider limit) · 2026-07
GLM-5.2 thinking · 2026-07
MiniMax-M3 thinking · 2026-07
MiMo-v2.5 thinking · 2026-07
DeepSeek-V4-Pro thinking · 2026-07
Hy3 2026-07

Metrics

Protocol

Submit a model