New paper from @marthaflindersand me: "Evaluating the Robustness of Analogical Reasoning in Large Language Models" 🧵 (1/6)

Nov 22, 2024 · 2:23 PM UTC

3
30
113
18,290
This is a much-extended follow-up on our earlier pre-print on "counterfactual tasks" in letter-string analogies. (2/6)
1
9
2,210
We revisit earlier claims of "emergent" zero-shot analogical reasoning in LLMs. We investigate robustness of previously published results in three domains: letter-string analogies, digit matrices, and story analogies. (3/6)
1
7
2,236
We evaluate humans and three GPT models on variations of tasks and also look at answer ordering effects. (4/6)
1
9
2,361
We find that on most of our experiments, while humans are able to adapt to variations in tasks and are insensitive to answer order, GPT models show brittleness to variations and can be highly sensitive to answer order. Our conclusions: (5/6)
1
8
38
2,700
Sort replies: Relevant Recent Liked
Replying to @MelMitchell1
Fwiw, when you prompt with things like “give ONLY the answer with no explanation” you are eliciting worse-than-default performance from GPT-4 class models My mantra: “analysis before answer, always”
2
4
448
Our goal was to replicate the original results (which used this prompt) and test for robustness. Other prompts including CoT might yield different results.
1
6
397