This article has been provided by Gleb Tsipursky, CEO of Disaster Avoidance Experts.

Engineering teams are being asked to make consequential decisions from a deceptively simple object: a benchmark score.

A model appears more accurate, capable or efficient than its competitors, and the number begins to stand in for a much harder question—whether the system will behave dependably in the environment where it will actually be used.

The distinction matters because benchmark performance is not operational fitness.

NIST’s recent work on AI evaluation separates benchmark accuracy from generalised accuracy. The first describes performance on a fixed set of test items. The second tries to estimate performance across a broader population of similar tasks. Those measures can differ, and both depend on assumptions about the test data, the target population and uncertainty.

For engineers, this is familiar territory. A component can meet a laboratory specification and still fail once temperature, vibration...

Parents
  • I support the article prepared by Gleb.  There is emerging body of evidence that AI models provide acceptable responses 70% of the time. The remaining responses range from error or misleading to harmful. This is why the article is important and relevant when using AI models. A second influencing characteristic of AI models is that the guidelines for the AI model goals include improving the profitability of the developer of the AI model and that decisions are based on word associations and not human concepts.

Comment
  • I support the article prepared by Gleb.  There is emerging body of evidence that AI models provide acceptable responses 70% of the time. The remaining responses range from error or misleading to harmful. This is why the article is important and relevant when using AI models. A second influencing characteristic of AI models is that the guidelines for the AI model goals include improving the profitability of the developer of the AI model and that decisions are based on word associations and not human concepts.

Children
No Data