Home | Search


2026-09-05

Memory tests might be useful for persuasion

Disclaimer

Main

gpt-5.6-sol high on memory tests (may contain hallucins)

To determine which topics, concepts, or words are easiest to retrieve from long-term semantic memory, use several complementary tests. Recall tests should be primary; recognition and association tests help diagnose why recall succeeds or fails.

Direct recall tests

| Test | Procedure | Useful measurements | Best for | |---|---|---|---| | Semantic/category fluency | Give a category such as “animals,” “tools,” or “programming languages”; produce as many examples as possible in 60–120 seconds. | Correct unique words, first-response latency, words per 15-second interval, repetitions, cluster size, switches between subcategories | Comparing broad topic accessibility | | Bounded-list recall | Ask for members of a known finite set: planets, chemical elements, state capitals, keyboard shortcuts | Percentage of set recalled, omissions, order, latency | Topics with clearly defined membership | | Fact cued recall | Ask short-answer questions: “What is the capital of Peru?” | Accuracy, response time, confidence, spelling/semantic closeness | Comparing factual topics | | Definition-to-word recall | Present a definition: “A word meaning fear of confined spaces.” Ask for the word. | Exact recall, partial recall, response time, tip-of-the-tongue reports | Vocabulary and technical terminology | | Word-to-definition production | Present a word and ask the person to define it without choices. | Number of correct semantic features, precision, misconceptions | Depth rather than mere label retrieval | | Semantic-feature generation | Give “tiger” and request category, appearance, habitat, behavior, and function | Correct features, distinctive features, feature diversity | Richness of conceptual knowledge | | Picture/object naming | Show an object or picture and ask for its name | Naming accuracy, latency, semantic substitutions | Concrete nouns and visually represented concepts | | Associate-cued recall | Present “doctor—?” when the learned or expected target is “nurse” | Recall probability under different cues, response latency | Strength and cue dependence of semantic links | | Topic knowledge dump | Give 3–5 minutes to write everything known about a topic | Number of correct propositions, breadth, depth, organization | Complex topics rather than isolated words |

Category fluency, picture naming, word–picture matching, synonym judgments, and semantic-association tasks are commonly combined because they probe different routes into semantic knowledge. (pmc.ncbi.nlm.nih.gov)

For category fluency, examine more than total word count. Average cluster size, number of subcategories reached, and number of switches provide information about the structure and searchability of a person’s semantic network. (pmc.ncbi.nlm.nih.gov)

Tests that separate recall failure from knowledge loss

1. Recall-then-recognition

For every item:

  1. Ask for an unaided answer.
  2. If recall fails, provide a hint.
  3. If that fails, provide multiple choices.
  4. Finally, show the answer and ask whether it feels familiar.

This produces a useful hierarchy:

Recognition alone should not be interpreted as easy recall.

2. Word–picture matching

Say or display a word and ask the participant to select the corresponding picture among semantically similar alternatives. Success despite naming failure suggests that conceptual knowledge remains available even when the verbal label cannot be retrieved.

3. Semantic-association judgment

Present a target and two alternatives:

The participant selects the more strongly associated item. The Pyramids and Palm Trees–style task measures access to semantic relationships without requiring the participant to produce the word. (pmc.ncbi.nlm.nih.gov)

4. Synonym or category judgment

Examples:

These primarily test meaning discrimination rather than recall.

5. Tip-of-the-tongue test

After failed recall, ask for:

This distinguishes completely unavailable items from items that are strongly activated but temporarily inaccessible.

Indirect accessibility tests

These are useful supplements but are not direct tests of voluntary recall:

Someone can perform well on these while still being unable to recall a word unaided.

Tests specifically for long-term retention

Delayed repeated recall

Administer equivalent recall tests at intervals such as:

Measure:

Avoid supplying feedback between tests if you want to measure retention without the test itself becoming additional study.

Relearning or “savings”

After a long delay, restudy forgotten items and measure:

An item may be unavailable in ordinary recall but relearn much faster than a genuinely novel item, indicating residual memory.

Cross-cue retrieval

Test the same concept using different cues:

A concept recalled through many independent cues is more robustly accessible than one recalled only under the original wording.

Recommended measurements

For every word or fact, record:

  1. Recall probability: Percentage of occasions or participants retrieving it.
  2. Correct-response latency: How quickly it is recalled.
  3. Cue requirement: Free recall, weak hint, strong hint, or recognition.
  4. Delayed retention: Whether it remains accessible over time.
  5. Retrieval consistency: Whether success repeats across sessions.
  6. Confidence: Preferably collected before feedback.
  7. Semantic precision: Exact answer versus vague or related answer.
  8. Error type: Omission, semantic substitution, phonological error, or misconception.
  9. Cue independence: Number of different cue formats that produce recall.
  10. Clustering position: Whether the word appears early, late, or only within a chain of related words.

A practical minimum battery

For each topic:

  1. Two-minute topic fluency
  2. 20–40 short-answer cued-recall questions
  3. Definition-to-word recall
  4. Recall-then-recognition for failures
  5. Repeat after one week and one month

Randomize item order and use the same scoring rules for every topic.

To rank topics, compare:

To rank individual words or facts more rigorously, use an item-response model or a mixed-effects model containing separate effects for participant and item. This prevents an easy participant or an unusually difficult question from being mistaken for a topic-level effect.

Important controls

Words and topics differ for reasons other than memory. Where possible, control or record:

Also note that semantic fluency is not a pure measure of stored semantic knowledge: processing speed, search strategy, working memory, and executive control can affect performance. (pmc.ncbi.nlm.nih.gov)

The strongest design therefore combines free recall, cued recall, recognition, response time, and delayed retesting, rather than treating any single test as definitive.

Subscribe

To subscribe, enter "Subscribe " followed by Signal username, email address or whatsapp phone number. You will receive atmost one update per month.

Alternatively, you can subscribe via RSS, or a third-party service like blogtrottr

Comment

Enter comment