An AI response can be easy to read and still be poor for the task.
A basic evaluation can begin with three questions:
These qualities overlap, but they are not the same.
Consider an AI response about a soccer rule.
If it describes the wrong restart, field location, or game condition, the answer has an accuracy problem.
Accuracy concerns the relationship between the claim and the actual facts.
An answer can be:
Do not evaluate the response only as one all-or-nothing block.
Suppose an AI response says:
That response contains several claims.
One claim can be correct while another is wrong.
Break important output into the statements that actually matter.
This makes evaluation more precise.
Consider two statements.
The second response identifies the information it used.
That gives you more reason to understand where the claim came from.
Credibility is stronger when the important claims can be connected to reliable information rather than unsupported confidence.
An AI response might say:
That sounds authoritative.
The name alone does not establish that:
Credibility requires more than authoritative-sounding wording.
For this activity, notice when a claim lacks a clear basis.
Later work will focus more directly on selecting evidence and authoritative sources.
Suppose the prompt is:
An AI response that provides a long history of the soccer club may be accurate.
It is still not relevant to the requested task.
A relevant response should match:
A long response can contain useful information and still bury the answer.
For a simple task, relevance may require brevity.
For a complex task, relevance may require explanation and context.
The correct amount of detail depends on the goal.
Not every sentence carries the same risk.
Compare:
with:
The first is a creative suggestion.
The second is a factual claim that could affect a real action.
The factual claim deserves more scrutiny.
A useful evaluation asks:
A response can be more trustworthy when it appropriately acknowledges uncertainty.
For example:
That is often more useful than inventing a specific time.
A good AI response should not pretend to know information that is missing from its available context.
A basic evaluation process can be:
What did you ask the AI to do?
Which parts of the response matter to the result?
Do any claims conflict with information you already know from the task or supplied material?
Does the response have a clear basis for factual claims, or does it rely on unsupported confidence?
Does the answer actually solve the requested problem at the needed level of detail?
Suppose an AI creates a practice announcement.
The wording may be:
But it may also include an invented practice time.
That means the response can be strong in one dimension and weak in another.
Evaluation should preserve that distinction.
The purpose of evaluating AI output is not to assume:
It is also not to assume:
A stronger position is:
That habit prepares you to work with AI as a tool while keeping human judgment in control.