ICML 2026 · Execution-based benchmark
Synthesizing KLayout DRC scripts from natural-language chip design rules, graded by execution on held-out layouts. 1,000 tasks, 13,921 evaluation layouts.
| Public set · 1,000 tasks | Private set · 100 tasks | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Docs in context | No context | Docs in context | No context | ||||||
| # | Model | ||||||||
| GPT-OSS-120B paper | |||||||||
| GPT-OSS-20B paper | |||||||||
| Qwen3-30B-A3B-Instruct-2507 paper | |||||||||
| Qwen3-235B-A22B-Thinking-2507 | |||||||||
| Qwen3-235B-A22B-Instruct-2507 | |||||||||
| Qwen3-30B-A3B-Thinking-2507 | |||||||||
| GLM-5.2 | |||||||||
| MiniMax-M3 | |||||||||
| MiMo-v2.5 | |||||||||
| DeepSeek-V4-Pro | |||||||||
| Hy3 | |||||||||
scripts/leaderboard/run_eval.sh from the Rule2DRC repository against any OpenAI-compatible endpoint, then open an issue or pull request with the printed result row.