Why can I tell a test is flaky but never find out why?
الفرصة
Detecting that a test is flaky is mostly solved; several tools do it reliably across reruns. Diagnosing the specific root cause is not. LLMs evaluated on real-world flaky test datasets achieve F1 scores above 0.88 for detection but drop below 0.57 on the same inputs when asked to identify root cause. The few automated repair tools only handle order-dependent or implementation-dependent flakiness and fail on resource contention, async timing, and environment drift. Developers spend hours reading logs and adding print statements, then watch the failure refuse to reproduce locally.
لماذا تهم
Root cause attribution is the missing step that turns a flakiness detector from a dashboard metric into a tool that actually reduces CI time.
كيف أقيّم الفرصة
نقاط الفرصة هي قراءتي الشخصية لا قياس دقيق: مدى تأثير المشكلة، وتكرار مواجهتها، وشُح الحلول المتاحة لها اليوم. كلما ارتفعت النقاط، كان البناء في رأيي أجدر بالاهتمام.
مقدار الألم الذي تسببه حين تظهر.
مدى تكرار مواجهة الناس لها فعلياً.
مدى شُح الأدوات الجيدة المتاحة لها اليوم.
مزيد من المشكلات التي تستحق الحل
لماذا يكون البرنامج الذي نعتمد عليه أكثر من غيره هو الأصعب في الاستخدام؟
Techلماذا لا أمتلك أياً من البيانات التي أولّدها؟
Techلماذا لا أستطيع الحصول على إيصال يُثبت أن بياناتي قد حُذفت فعلاً؟
Techلماذا لا أستطيع معرفة ما إذا كان ما يعمل فعلاً يطابق ما أعلنه SBOM الخاص بي؟
Techلماذا تنهار كل سلسلة مصدر C2PA فور وصول المحتوى إلى وسائل التواصل الاجتماعي؟
Techلماذا تمر الاختبارات التي تم إنشاؤها بواسطة الذكاء الاصطناعي على الأخطاء التي كان من المفترض أن تكتشفها؟