The answer
Evals are valuable. Offline evals catch regressions on a known test set before a deploy, and online evals can score real production traffic. Both still depend on the rubrics, evaluators, samples, and populations a team chooses to measure.
Production introduces language, intents, edge cases, and user groups that a fixed dataset may not represent. A green eval can coexist with a recurring issue in a population the evaluator does not isolate, or with gradual behavior change that needs a time-based comparison.
Cross-session behavioral analysis complements those scores. Flowlines groups what is happening across real sessions, which issues recur, who is affected, and which sessions show the behavior. The useful combination is defined evaluation plus open-ended production analysis.
Last reviewed 2026-09-04