it's great that with decbench, we have a high-quality public benchmark.
yet, decompiler evaluation remains hard, and I see two primary challenges right now:
1. high-quality, real-world-like, fully private data. Frontier LLMs have seen virtually every open source function in existence and leakage is not fully avoidable, something we also cannot rule out for our model.
2. finding a better proxy for semantic fidelity. A program can recompile, but do something entirely different. A program can also recompile, have a totally different byte-level structure from the groundtruth, and still do the same job. Dynamic-execution based attestation is tough to get correct at scale, but might be a better signal in the long run.