FrontierOR
FrontierOR tracks two evaluation modes. In one-shot evaluation, we monitor the latest frontier models as each directly generates an optimization algorithm for every task. In test-time evolution, we track latest agent frameworks and harnesses that iteratively generate, execute, and refine programs under a fixed budget, measuring whether feedback-driven search improves the quality-time frontier beyond direct model generation.
Hover over a metric column header (Exec. rate / Feasibility / Sol. quality / QTE) to see its definition and value range.
FrontierOR Full (n=—)
One-shot performance of different LLMs. Ranked by QTE.
FrontierOR Hard (n=—)
The Hard subset comprises 50 tasks whose problem class or instance structure is computationally demanding and where Gurobi fails to reach optimality within a one-hour budget — a Gurobi-saturated tail of the full benchmark.