Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm

Guardado en:
书目详细资料
发表在:arXiv.org (Dec 24, 2024), p. n/a
主要作者: Hu, Xiaoyang
其他作者: Lewis, Richard L
出版:
Cornell University Library, arXiv.org
主题:
在线阅读:Citation/Abstract
Full text outside of ProQuest
标签: 添加标签
没有标签, 成为第一个标记此记录!
实物特征
摘要:Cognitive tasks originally developed for humans are now increasingly used to study language models. While applying these tasks is often straightforward, interpreting their results can be challenging. In particular, when a model underperforms, it's often unclear whether this results from a limitation in the cognitive ability being tested or a failure to understand the task itself. A recent study argued that GPT 3.5's declining performance on 2-back and 3-back tasks reflects a working memory capacity limit similar to humans. By analyzing a range of open-source language models of varying performance levels on these tasks, we show that the poor performance instead reflects a limitation in task comprehension and task set maintenance. In addition, we push the best performing model to higher n values and experiment with alternative prompting strategies, before analyzing model attentions. Our larger aim is to contribute to the ongoing conversation around refining methodologies for the cognitive evaluation of language models.
ISSN:2331-8422
Fuente:Engineering Database