Aria AI · Project 25 / 25

ORBench — Solver and Agent Benchmark

Compare structured CP-SAT, fixed heuristics, and text-to-schema agents on 100 small inventory, assignment, and scheduling tasks. The numeric validator is the judge. Unevaluated models stay not_evaluated.

Open live demo Dataset Organization Project 22 Dock
100 tasksInventory, assignment, scheduling
60 / 40Development / test, 20 interactive
ValidatorFeasibility and gap, not an LLM judge
CPUOR-Tools · no paid API

Organization Gradio hosting needs Team/Enterprise. This page is the AriaAICompany project card; the iframe is the live benchmark on the personal PRO account when org compute is unavailable.

اوربنچ

روش‌های ساخت تصمیم را روی مسئله‌های کوچک موجودی، تخصیص و زمان‌بندی مقایسه می‌کند. داور validator عددی است. خانه مدل اجرا‌نشده «ارزیابی نشده» است، نه صفر.