Skip to content
Tech

Why can I tell a test is flaky but never find out why?

79

機会

Detecting that a test is flaky is mostly solved; several tools do it reliably across reruns. Diagnosing the specific root cause is not. LLMs evaluated on real-world flaky test datasets achieve F1 scores above 0.88 for detection but drop below 0.57 on the same inputs when asked to identify root cause. The few automated repair tools only handle order-dependent or implementation-dependent flakiness and fail on resource contention, async timing, and environment drift. Developers spend hours reading logs and adding print statements, then watch the failure refuse to reproduce locally.

重要な理由

Root cause attribution is the missing step that turns a flakiness detector from a dashboard metric into a tool that actually reduces CI time.

機会をどう評価するか

Opportunity Scoreは測定値ではなく、私自身の見解です。どれほど痛みを伴うか、どれほど頻繁に影響を与えるか、そして今日時点で解決策がいかに少ないか。スコアが高いほど、構築する価値が高いと私は考えています。

深刻度7/10

それが現れたときにどれほどの痛みをもたらすか。

頻度9/10

実際にどれほど頻繁に人々がそれに直面するか。

ホワイトスペース8/10

今日時点で、それに対する優れたツールがいかに少ないか。

解決する価値のある問題をもっと見る