Compare structured CP-SAT, fixed heuristics, and text-to-schema agents on
100 small inventory, assignment, and scheduling tasks. The numeric
validator is the judge. Unevaluated models stay not_evaluated.
100 tasksInventory, assignment, scheduling60 / 40Development / test, 20 interactiveValidatorFeasibility and gap, not an LLM judgeCPUOR-Tools · no paid API
Organization Gradio hosting needs Team/Enterprise. This page is the AriaAICompany project card;
the iframe is the live benchmark on the personal PRO account when org compute is unavailable.
اوربنچ
روشهای ساخت تصمیم را روی مسئلههای کوچک موجودی، تخصیص و زمانبندی مقایسه میکند. داور validator عددی است. خانه مدل اجرانشده «ارزیابی نشده» است، نه صفر.