2026-09-05

# Memory tests might be useful for persuasion

Disclaimer
 - **Quick Note**

Main
 - Many famous youtubers aimed at mass audience agree that you should use 5th grader (or ideally 3rd grader) level vocabulary. Imagine the dumbest person you know and target someone dumber than that.
 - I will go one step further and wonder whether it is worth measuring more explicitly, which words are easy to recall even for even people with low curiosity or IQ or knowledge or similar. Like, instead of just using the standard 3rd grader vocabulary, maybe you can more explicitly map out the vocabulary of individual people or of specific demographics.
 - About human memory
   - Working memory supposedly stores 7 bits of information for a few seconds to a few minutes at max. Most of itt gets forgetten but some of it can get moved to long-term memory.
   - Long-term memory is typically semantic and episodic. People have semantic maps inside their heads. The dumbest person you know is probably carrying a very different semantic map than you are.
 - If you knew what semantic maps people are actually carrying inside their heads, you might be more effective at making content to persuade them

# gpt-5.6-sol high on memory tests (may contain hallucins)

To determine **which topics, concepts, or words are easiest to retrieve from long-term semantic memory**, use several complementary tests. Recall tests should be primary; recognition and association tests help diagnose *why* recall succeeds or fails.

## Direct recall tests

| Test | Procedure | Useful measurements | Best for |
|---|---|---|---|
| **Semantic/category fluency** | Give a category such as “animals,” “tools,” or “programming languages”; produce as many examples as possible in 60–120 seconds. | Correct unique words, first-response latency, words per 15-second interval, repetitions, cluster size, switches between subcategories | Comparing broad topic accessibility |
| **Bounded-list recall** | Ask for members of a known finite set: planets, chemical elements, state capitals, keyboard shortcuts | Percentage of set recalled, omissions, order, latency | Topics with clearly defined membership |
| **Fact cued recall** | Ask short-answer questions: “What is the capital of Peru?” | Accuracy, response time, confidence, spelling/semantic closeness | Comparing factual topics |
| **Definition-to-word recall** | Present a definition: “A word meaning fear of confined spaces.” Ask for the word. | Exact recall, partial recall, response time, tip-of-the-tongue reports | Vocabulary and technical terminology |
| **Word-to-definition production** | Present a word and ask the person to define it without choices. | Number of correct semantic features, precision, misconceptions | Depth rather than mere label retrieval |
| **Semantic-feature generation** | Give “tiger” and request category, appearance, habitat, behavior, and function | Correct features, distinctive features, feature diversity | Richness of conceptual knowledge |
| **Picture/object naming** | Show an object or picture and ask for its name | Naming accuracy, latency, semantic substitutions | Concrete nouns and visually represented concepts |
| **Associate-cued recall** | Present “doctor—?” when the learned or expected target is “nurse” | Recall probability under different cues, response latency | Strength and cue dependence of semantic links |
| **Topic knowledge dump** | Give 3–5 minutes to write everything known about a topic | Number of correct propositions, breadth, depth, organization | Complex topics rather than isolated words |

Category fluency, picture naming, word–picture matching, synonym judgments, and semantic-association tasks are commonly combined because they probe different routes into semantic knowledge. ([pmc.ncbi.nlm.nih.gov](https://pmc.ncbi.nlm.nih.gov/articles/PMC2255584/?utm_source=openai))

For category fluency, examine more than total word count. **Average cluster size**, **number of subcategories reached**, and **number of switches** provide information about the structure and searchability of a person’s semantic network. ([pmc.ncbi.nlm.nih.gov](https://pmc.ncbi.nlm.nih.gov/articles/PMC5796672/?utm_source=openai))

## Tests that separate recall failure from knowledge loss

### 1. Recall-then-recognition

For every item:

1. Ask for an unaided answer.
2. If recall fails, provide a hint.
3. If that fails, provide multiple choices.
4. Finally, show the answer and ask whether it feels familiar.

This produces a useful hierarchy:

- **Free recall success:** highly accessible
- **Cued recall only:** stored but cue-dependent
- **Recognition only:** familiar but not independently retrievable
- **Recognition failure:** probably weak, absent, or misunderstood knowledge

Recognition alone should not be interpreted as easy recall.

### 2. Word–picture matching

Say or display a word and ask the participant to select the corresponding picture among semantically similar alternatives. Success despite naming failure suggests that conceptual knowledge remains available even when the verbal label cannot be retrieved.

### 3. Semantic-association judgment

Present a target and two alternatives:

- target: `pyramid`
- choices: `palm tree`, `pine tree`

The participant selects the more strongly associated item. The Pyramids and Palm Trees–style task measures access to semantic relationships without requiring the participant to produce the word. ([pmc.ncbi.nlm.nih.gov](https://pmc.ncbi.nlm.nih.gov/articles/PMC2805958/?utm_source=openai))

### 4. Synonym or category judgment

Examples:

- Are “rapid” and “swift” similar in meaning?
- Is a whale a fish or a mammal?
- Which item does not belong: violin, trumpet, piano, telescope?

These primarily test meaning discrimination rather than recall.

### 5. Tip-of-the-tongue test

After failed recall, ask for:

- Confidence that the answer is known
- First letter
- Number of syllables
- Rhyming words
- Related concepts
- Whether the answer would be recognized

This distinguishes completely unavailable items from items that are strongly activated but temporarily inaccessible.

## Indirect accessibility tests

These are useful supplements but are **not direct tests of voluntary recall**:

- **Lexical decision:** Decide quickly whether a letter string is a real word.
- **Semantic priming:** Measure whether seeing `doctor` speeds recognition of `nurse`.
- **Reading or naming latency:** Measure how quickly a displayed word is pronounced.
- **Stroop-like semantic interference:** Measure how strongly irrelevant meanings interfere.
- **Sorting by meaning:** Group words or pictures into categories.

Someone can perform well on these while still being unable to recall a word unaided.

## Tests specifically for long-term retention

### Delayed repeated recall

Administer equivalent recall tests at intervals such as:

- Initial baseline
- 1 day
- 1 week
- 1 month
- 3 months

Measure:

- Recall accuracy at each delay
- Response-time changes
- Probability of recovering a previously missed item
- Forgetting rate
- Consistency across occasions

Avoid supplying feedback between tests if you want to measure retention without the test itself becoming additional study.

### Relearning or “savings”

After a long delay, restudy forgotten items and measure:

- Trials needed to relearn
- Study time needed
- Accuracy after one refresher
- Speed of subsequent recall

An item may be unavailable in ordinary recall but relearn much faster than a genuinely novel item, indicating residual memory.

### Cross-cue retrieval

Test the same concept using different cues:

- Question → answer
- Definition → word
- Picture → name
- Category → example
- Related concept → target
- Real-world scenario → relevant concept

A concept recalled through many independent cues is more robustly accessible than one recalled only under the original wording.

## Recommended measurements

For every word or fact, record:

1. **Recall probability:** Percentage of occasions or participants retrieving it.
2. **Correct-response latency:** How quickly it is recalled.
3. **Cue requirement:** Free recall, weak hint, strong hint, or recognition.
4. **Delayed retention:** Whether it remains accessible over time.
5. **Retrieval consistency:** Whether success repeats across sessions.
6. **Confidence:** Preferably collected before feedback.
7. **Semantic precision:** Exact answer versus vague or related answer.
8. **Error type:** Omission, semantic substitution, phonological error, or misconception.
9. **Cue independence:** Number of different cue formats that produce recall.
10. **Clustering position:** Whether the word appears early, late, or only within a chain of related words.

## A practical minimum battery

For each topic:

1. **Two-minute topic fluency**
2. **20–40 short-answer cued-recall questions**
3. **Definition-to-word recall**
4. **Recall-then-recognition for failures**
5. **Repeat after one week and one month**

Randomize item order and use the same scoring rules for every topic.

To rank topics, compare:

- Mean free-recall accuracy
- Median correct response time
- Percentage requiring hints
- One-month retention
- Cross-cue success

To rank individual words or facts more rigorously, use an **item-response model** or a **mixed-effects model** containing separate effects for participant and item. This prevents an easy participant or an unusually difficult question from being mistaken for a topic-level effect.

## Important controls

Words and topics differ for reasons other than memory. Where possible, control or record:

- Prior familiarity and expertise
- Word frequency
- Age at which the word was learned
- Concreteness and imageability
- Word and answer length
- Emotional relevance
- Cultural and language background
- Recency of use
- Quality and number of retrieval cues
- Question difficulty and ambiguity
- Typing or speaking speed

Also note that semantic fluency is not a pure measure of stored semantic knowledge: processing speed, search strategy, working memory, and executive control can affect performance. ([pmc.ncbi.nlm.nih.gov](https://pmc.ncbi.nlm.nih.gov/articles/PMC5619366/?utm_source=openai))

The strongest design therefore combines **free recall, cued recall, recognition, response time, and delayed retesting**, rather than treating any single test as definitive.


