Lesson 4 · Verification

1.4 · Evaluation, Deadlines, and Operational Success

Offline model quality is necessary but insufficient. Physical AI must be evaluated through closed-loop behavior, real-time timing, interventions, constraint violations, and recovery from faults.

Learning outcomes

  • Separate offline, closed-loop, and operational evaluation.
  • Build an end-to-end latency budget.
  • Select metrics that expose physical rather than linguistic competence.

Three evaluation levels

Offline metrics test prediction on recorded data. Closed-loop tests measure behavior when the policy changes its own future inputs. Operational tests include uptime, intervention rate, recovery, calibration, and performance under faults.

Deadline before average speed

A control rate defines a deadline. The relevant timing statistic is the synchronized end-to-end distribution—often a high percentile—not an average of isolated components. Parallel paths must be modeled by their actual critical path.

A deployment scorecard

Track success rate, completion time, intervention rate, constraint violations, calibration, out-of-distribution coverage, worst-case latency, and behavior after sensor or model faults. Stratify results by environment, object, instruction, and horizon.

Key takeaway

A model is ready only when the closed-loop system meets both behavioral and timing requirements.