0.7.5 Evaluating AI Accuracy, Credibility, and Relevance

A Useful AI Response Needs More Than Good Writing

An AI response can be easy to read and still be poor for the task.

A basic evaluation can begin with three questions:

  • Accuracy: Is the information correct?
  • Credibility: Is there a reasonable basis for trusting the important claims?
  • Relevance: Does the response actually address the task?

These qualities overlap, but they are not the same.

Accuracy Asks Whether the Information Is Correct

Consider an AI response about a soccer rule.

If it describes the wrong restart, field location, or game condition, the answer has an accuracy problem.

Accuracy concerns the relationship between the claim and the actual facts.

An answer can be:

  • relevant but inaccurate;
  • accurate but irrelevant;
  • partly accurate and partly inaccurate.

Do not evaluate the response only as one all-or-nothing block.

Separate Claims When the Response Contains Several

Suppose an AI response says:

That response contains several claims.

One claim can be correct while another is wrong.

Break important output into the statements that actually matter.

This makes evaluation more precise.

Credibility Asks Whether the Claim Has a Trustworthy Basis

Consider two statements.

Response A
Response B

The second response identifies the information it used.

That gives you more reason to understand where the claim came from.

Credibility is stronger when the important claims can be connected to reliable information rather than unsupported confidence.

A Source-Like Statement Is Not Automatically Credible

An AI response might say:

That sounds authoritative.

The name alone does not establish that:

  • the source exists;
  • the source says what the AI claims;
  • the source applies to the current situation.

Credibility requires more than authoritative-sounding wording.

For this activity, notice when a claim lacks a clear basis.

Later work will focus more directly on selecting evidence and authoritative sources.

Relevance Asks Whether the Response Solves the Requested Task

Suppose the prompt is:

An AI response that provides a long history of the soccer club may be accurate.

It is still not relevant to the requested task.

A relevant response should match:

  • the question;
  • the audience;
  • the requested scope;
  • the needed format.
More Detail Is Not Always More Relevant

A long response can contain useful information and still bury the answer.

For a simple task, relevance may require brevity.

For a complex task, relevance may require explanation and context.

The correct amount of detail depends on the goal.

Evaluate Important Claims More Carefully

Not every sentence carries the same risk.

Compare:

with:

The first is a creative suggestion.

The second is a factual claim that could affect a real action.

The factual claim deserves more scrutiny.

A useful evaluation asks:

Notice Uncertainty

A response can be more trustworthy when it appropriately acknowledges uncertainty.

For example:

That is often more useful than inventing a specific time.

A good AI response should not pretend to know information that is missing from its available context.

Evaluate the Response Against the Task

A basic evaluation process can be:

1. Identify the task

What did you ask the AI to do?

2. Identify the important claims or recommendations

Which parts of the response matter to the result?

3. Consider accuracy

Do any claims conflict with information you already know from the task or supplied material?

4. Consider credibility

Does the response have a clear basis for factual claims, or does it rely on unsupported confidence?

5. Consider relevance

Does the answer actually solve the requested problem at the needed level of detail?

A Response Can Have Mixed Quality

Suppose an AI creates a practice announcement.

The wording may be:

  • relevant;
  • professional;
  • concise.

But it may also include an invented practice time.

That means the response can be strong in one dimension and weak in another.

Evaluation should preserve that distinction.

The Goal Is Deliberate Use

The purpose of evaluating AI output is not to assume:

It is also not to assume:

A stronger position is:

That habit prepares you to work with AI as a tool while keeping human judgment in control.