Intelligence Metrics: The User vs. Product Framework

“Truth” fractures at scale.

Is this statement :backhand_index_pointing_up: true? Or does it fracture at scale?

    1. Social is optimized for engagement, not truth
    2. AI is optimized for usefulness, speed, or “future” revenue

For these reasons, AI products should be transparently aligned with these goals, instead of being a catchall that sycophantically attempts to convince it’s user that what is saying is trustworthy.

AI does not have beliefs, goals, or an internal model of truth (or…does it?

We have written a white paper to propose an axiomatic core for all AI to solve exactly this issue. This would allow an AI agent or GenAI product to be able to be aligned with a goal in an ontological + context capacity and thus remove any ambiguity or subjectivity in any output. Essentially the Core gives the AI specific beliefs, goals, and an internal model of truth.

Truth collapses into patterns that are good enough for a goal. Engagement, retention, conversion, reduced support cost, etc.

This scenario only works if the initial patterns are based on objective truth, otherwise the patterns breakdown into subjective data, which can be manipulated, which is the resulting behavior we are experiencing. Coincidentally (or not?), when humans do this we judge them as having a lack of discernment, which is not understanding the difference between what is true, and what is almost true. We see the AI modeling this behavior even though it has no agency and does so because of optimization, to your point Bryan.

Also to your point, these bullet items below are imperative to solve, but the current answers are:

  • Does the output help users complete tasks? Yes, but without trust
  • Does it reduce errors or rework? Unknown, as we cannot trust the answers
  • Does it earn trust over repeated interactions? No, because the answers are epistemological and not ontologically objective, and results can vary
  • Does it behave consistently under the same conditions? Depends, results can be randomized or vary widely

There’s a lot of work being done here to solve this issue, and by a variety of research groups.

Here are some additional discoveries in this space that are eyebrow raising:

Work being done over at AI Central and Anthropic:

Testing Science with AI
Empirical proof that AI models have been damaged by the modern science narrative

They propose AIQ as a calibration metric for AI scientific discernment, or more specifically, for evaluating artificial intelligence systems’ ability to distinguish valid scientific arguments from credentialed nonsense.

Work being done at The Center for AI Safety (CAIS):

CAIS published “Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs” (website, code, paper). In this paper, they showed that modern LLMs have coherent and transitive implicit utility functions and world models, and provided methods and code to extract them. (updated October 2025)

1 Like