Every agent researcher knows the uncomfortable gap between a leaderboard score and a system that survives contact with reality. A model can ace a static test set and still collapse the moment someone changes the tool schema, adds latency, or asks a follow-up question the rubric never anticipated. Live evaluation with an adversarial human in the loop is a different animal entirely.
That is, structurally, what Startup Battlefield 200 is. TechCrunch has revealed the next wave of judges for Disrupt 2026: Nell Daly, Michael Palank, Aditi Maliwal, Chrystal Huang, and Grace Ge. They will spend October 13-15 in San Francisco assessing founders in real time. Strip away the stage lighting and what you have is a human eval use with five graders, no answer key, and unlimited follow-up turns.
Why live judging is harder than any static test
In agent research, we distinguish between offline evaluation and interactive evaluation for a good reason. Offline benchmarks ask a fixed question and score a fixed answer. Interactive evaluation lets the evaluator probe, adapt, and attack the weakest link. The second is far more informative and far more expensive, which is why almost nobody runs enough of it.
A pitch panel is interactive evaluation applied to companies instead of models. The founder arrives with a prepared trajectory. The judges immediately go off-distribution. Where does the margin actually come from? What happens when your largest customer renegotiates? What is your defensibility when the underlying model provider ships the same feature for free?
Anyone who has watched an agent handle a well-crafted prompt and then fall apart on an ambiguous one recognizes the pattern. Performance on the happy path tells you almost nothing about performance under distribution shift. Five investors asking unscripted questions across three days generates exactly the kind of off-path pressure that reveals whether the reasoning underneath is real or memorized.
What multiple judges actually buy you
There is a technical reason panels beat single evaluators, and it is not just fairness. Single-judge evaluation inherits that judge’s biases wholesale. Anyone who has used a model-as-judge setup knows the failure mode: the evaluator develops consistent preferences for length, confidence, or a particular framing, and those preferences quietly become the optimization target.
Multiple independent evaluators reduce that variance. Different priors, different sector exposure, different scar tissue from prior investments. Agreement across a heterogeneous panel is a stronger signal than enthusiasm from any one member. Disagreement is informative too. When evaluators split sharply, you are usually looking at something genuinely novel or genuinely underspecified, and knowing which matters enormously.
For founders building agent products specifically, this matters more than it used to. The category is crowded with demos that work beautifully in a controlled environment. The questions that separate the field tend to be architectural:
- What happens to your system when a tool call fails silently rather than loudly?
- How do you measure success on tasks where correctness is subjective?
- What is your fallback when the model behind your product changes behavior in a version bump?
- Where does human review sit in the loop, and what does it cost per task?
- How do you keep evaluation from drifting as your own product shapes user behavior?
None of these are answerable with a slide. They require that the team has actually instrumented their system and looked at the failures. Investors who have seen enough agent pitches have learned to ask, and a live format gives them room to keep asking.
The transfer problem
One honest caveat. Pitch judging shares a weakness with every human evaluation protocol: performance on the eval is not the same as performance in deployment. Presenting well under pressure correlates with founder quality, but it is a proxy, and proxies degrade when people optimize for them directly. Pitch coaching is, in effect, eval-set contamination.
The counterweight is variety. A panel drawn from different backgrounds is harder to game than a single rubric, because there is no unified target to overfit to. The same principle applies when we build evaluation suites for agents. Diversity of test conditions does more for signal quality than depth on any single axis.
What to take from this
If you are building agent systems and planning to be in the room, treat the panel the way you would treat a red team. Assume the questions will target your weakest assumption. Know your failure modes before someone else finds them, because the useful outcome of a hard question is not surviving it, it is learning something from how badly you handled it.
Disrupt 2026 runs October 13-15 in San Francisco. Savings of up to $200 on tickets are available before rates rise on September 25 at 11:59 p.m. PT, and registering early is the cheaper path.
The interesting part is not who wins. It is watching which architectural claims hold up when a human evaluator refuses to stay on script.
đź•’ Published: