How do I know my agent got better between versions and did not just get lucky?
الفرصة
When an agent takes hundreds of sequential steps over hours, traditional A/B evaluation breaks down because an early tool call shapes every subsequent choice, making outcomes path-dependent across runs. Running enough independent trials to get reliable signal costs as much compute as training. The field defaults to proxy metrics such as step success rate and tool call accuracy that demonstrably do not correlate with end-task outcomes on real work. A 2025 audit of 445 published LLM benchmarks documented construct-validity failures at scale: vague task definitions, repurposed short-horizon datasets, and missing statistical tests, all of which become more severe as task horizon grows. Teams building agentic products ship on manual spot-checks because no principled, reproducible evaluation methodology for long-horizon agents exists.
لماذا تهم
Without a reliable way to measure whether an agent got better, you cannot systematically improve agents that are already deployed in high-stakes tasks.
كيف أقيّم الفرصة
نقاط الفرصة هي قراءتي الشخصية لا قياس دقيق: مدى تأثير المشكلة، وتكرار مواجهتها، وشُح الحلول المتاحة لها اليوم. كلما ارتفعت النقاط، كان البناء في رأيي أجدر بالاهتمام.
مقدار الألم الذي تسببه حين تظهر.
مدى تكرار مواجهة الناس لها فعلياً.
مدى شُح الأدوات الجيدة المتاحة لها اليوم.
مزيد من المشكلات التي تستحق الحل
لماذا تنساني كل تطبيقات الذكاء الاصطناعي في اللحظة التي أغلق فيها التبويب؟
AIلماذا لا يزال تعلم مجال جديد رهيناً بمعرفة الأسئلة الصحيحة؟
AIلماذا لا يستطيع غير المتخصص التحقق مما أخبره به الذكاء الاصطناعي للتو؟
AIلماذا نختبر النماذج على المعايير القياسية ثم نطلقها بناءً على الحدس؟
AIلماذا لا تملك وكلاء الذكاء الاصطناعي ذاكرة لأخطائها الخاصة؟
AIلماذا لا يمكنني مراجعة ما تدرّب عليه النموذج فعلاً؟