New paper from @marthaflindersand me:
"Evaluating the Robustness of Analogical Reasoning in Large Language Models"
馃У
(1/6)
3
30
113
18,290
This is a much-extended follow-up on our earlier pre-print on "counterfactual tasks" in letter-string analogies.
(2/6)
1
9
2,210
We revisit earlier claims of "emergent" zero-shot analogical reasoning in LLMs. We investigate robustness of previously published results in three domains: letter-string analogies, digit matrices, and story analogies.
(3/6)
1
7
2,236
We find that on most of our experiments, while humans are able to adapt to variations in tasks and are insensitive to answer order, GPT models show brittleness to variations and can be highly sensitive to answer order.
Our conclusions:
(5/6)
Nov 22, 2024 路 2:23 PM UTC
1
8
38
2,700

