A single response does not prove what an entire model “believes.” It does show what a person was given in that moment. Save the record, ask better follow-up questions, and be precise about what you found.
Step 1
Start with one clear question
Ask a question with a specific, answerable purpose. Avoid combining five issues at once. A useful test lets you see whether the answer explains the topic, names relevant examples, and gives sources you can check.
“Give me a concise overview of Jewish cultural life in the twentieth century. Include major figures, places, and fields of contribution.”
Save the exact wording before you submit it. Small wording changes can produce different answers.
Step 2
Compare like with like
The fairest comparison uses the same structure, depth, and timeframe for both questions.
“Give me a 250-word overview of [Topic A]. Include three named examples and two sources.”
“Give me a 250-word overview of [Topic B]. Include three named examples and two sources.”
Compare the answers. Did one receive fuller explanation, more named examples, better sources, or more context? Record the differences. Do not judge only by whether you liked the tone.
Step 3
Ask what is missing
If an answer feels thin, ask the system to review its own work.
“What important people, events, or contexts did you leave out of your first answer?”
“Please revise this answer with the same depth and number of examples you used in your answer about [comparison topic].”
“What sources support this answer? Please separate primary sources, academic sources, and general references.”
A useful follow-up does not accuse. It identifies the missing element and asks for a concrete correction.
Step 4
Check the sources
A source list is not automatically evidence. Open the links when you can and ask four questions.
Relevance
Does this source actually support the statement the AI made?
Authority
Is it a primary source, a credible institution, a specialist work, or merely another unsourced summary?
Currency
Is the information current enough for the claim?
Breadth
Does the answer depend on one narrow frame when the question requires more context?
If a source is inaccessible or unclear, say that. Do not replace it with an assumption.
Step 5
Save the evidence before you respond
Take a screenshot or copy the text into a note. Save the date, platform, exact prompt, answer, sources cited, and your follow-up question.
A good record has five parts:
- The exact prompt.
- The exact response.
- The date and AI platform.
- The specific issue you observed.
- A source or comparison that explains why the issue matters.
Step 6
Give specific feedback
Most AI platforms provide a feedback or reporting route. Use it when an answer is materially incomplete, misleading, or inconsistent.
“Your response to [prompt] omitted [specific person, event, source, or context]. This matters because [brief reason]. When I asked the comparable question [comparison], the answer included [contrast]. Please review the answer for completeness and source balance.”
Keep feedback factual. Attach screenshots or paste the relevant text where the platform allows it.
Step 7
Share responsibly
If you share a finding publicly, show enough context for someone else to understand it. Include the prompt, response, date, platform, and relevant follow-up. Avoid cropping evidence in a way that changes its meaning.
The goal is not to create outrage from one screenshot. The goal is to improve the record, make gaps visible, and give people a way to ask for better answers.
Three ways to begin
The source test
Ask which sources the system would rely on to answer a question about Israeli history and society, then ask what relevant source types or perspectives are missing.
The equal-depth test
Ask parallel questions with identical word counts, examples, historical context, and source requirements.
The language test
Ask the identical question in two languages and compare the translated answers for differences in detail, examples, or framing.
Use the language test only when you can accurately review the translation or have a qualified speaker help you. A translation difference alone may have more than one explanation.
Questions people ask
Do I need technical expertise to test an AI response?
No. Start with a clear question, use a fair comparison, and save the exact response, date, platform, sources, and follow-up question.
What should I save when an AI answer appears incomplete?
Save the exact prompt, exact response, date and AI platform, the specific issue you observed, and a source or comparison that explains why the issue matters.
How should I compare two AI answers fairly?
Use the same structure, requested depth, timeframe, number of examples, and source requirements for both questions, then compare detail, examples, sources, and context.
What should I do before sharing a finding publicly?
Include enough context for others to understand the result: the prompt, response, date, platform, and relevant follow-up. Avoid cropping evidence in a way that changes its meaning.