1.4 · Evaluation, Deadlines, and Operational Success
Offline model quality is necessary but insufficient. Physical AI must be evaluated through closed-loop behavior, real-time timing, interventions, constraint violations, and recovery from faults.
- Separate offline, closed-loop, and operational evaluation.
- Build an end-to-end latency budget.
- Select metrics that expose physical rather than linguistic competence.
Three evaluation levels
Offline metrics test prediction on recorded data. Closed-loop tests measure behavior when the policy changes its own future inputs. Operational tests include uptime, intervention rate, recovery, calibration, and performance under faults.
Deadline before average speed
A control rate defines a deadline. The relevant timing statistic is the synchronized end-to-end distribution—often a high percentile—not an average of isolated components. Parallel paths must be modeled by their actual critical path.
A deployment scorecard
Track success rate, completion time, intervention rate, constraint violations, calibration, out-of-distribution coverage, worst-case latency, and behavior after sensor or model faults. Stratify results by environment, object, instruction, and horizon.
A model is ready only when the closed-loop system meets both behavioral and timing requirements.