Skip to content
AI

How do I know my agent got better between versions and did not just get lucky?

81

機会

When an agent takes hundreds of sequential steps over hours, traditional A/B evaluation breaks down because an early tool call shapes every subsequent choice, making outcomes path-dependent across runs. Running enough independent trials to get reliable signal costs as much compute as training. The field defaults to proxy metrics such as step success rate and tool call accuracy that demonstrably do not correlate with end-task outcomes on real work. A 2025 audit of 445 published LLM benchmarks documented construct-validity failures at scale: vague task definitions, repurposed short-horizon datasets, and missing statistical tests, all of which become more severe as task horizon grows. Teams building agentic products ship on manual spot-checks because no principled, reproducible evaluation methodology for long-horizon agents exists.

重要な理由

Without a reliable way to measure whether an agent got better, you cannot systematically improve agents that are already deployed in high-stakes tasks.

機会をどう評価するか

Opportunity Scoreは測定値ではなく、私自身の見解です。どれほど痛みを伴うか、どれほど頻繁に影響を与えるか、そして今日時点で解決策がいかに少ないか。スコアが高いほど、構築する価値が高いと私は考えています。

深刻度7/10

それが現れたときにどれほどの痛みをもたらすか。

頻度9/10

実際にどれほど頻繁に人々がそれに直面するか。

ホワイトスペース9/10

今日時点で、それに対する優れたツールがいかに少ないか。

解決する価値のある問題をもっと見る