Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks

Pink, Mathis; Vo, Vy A.; Wu, Qinyuan; Mu, Jianing; Turek, Javier S.; Hasson, Uri; Norman, Kenneth A.; Michelmann, Sebastian; Huth, Alexander; Toneva, Mariya

Computer Science > Computation and Language

arXiv:2410.08133 (cs)

[Submitted on 10 Oct 2024]

Title:Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks

Authors:Mathis Pink, Vy A. Vo, Qinyuan Wu, Jianing Mu, Javier S. Turek, Uri Hasson, Kenneth A. Norman, Sebastian Michelmann, Alexander Huth, Mariya Toneva

View PDF HTML (experimental)

Abstract:Current LLM benchmarks focus on evaluating models' memory of facts and semantic relations, primarily assessing semantic aspects of long-term memory. However, in humans, long-term memory also includes episodic memory, which links memories to their contexts, such as the time and place they occurred. The ability to contextualize memories is crucial for many cognitive tasks and everyday functions. This form of memory has not been evaluated in LLMs with existing benchmarks. To address the gap in evaluating memory in LLMs, we introduce Sequence Order Recall Tasks (SORT), which we adapt from tasks used to study episodic memory in cognitive psychology. SORT requires LLMs to recall the correct order of text segments, and provides a general framework that is both easily extendable and does not require any additional annotations. We present an initial evaluation dataset, Book-SORT, comprising 36k pairs of segments extracted from 9 books recently added to the public domain. Based on a human experiment with 155 participants, we show that humans can recall sequence order based on long-term memory of a book. We find that models can perform the task with high accuracy when relevant text is given in-context during the SORT evaluation. However, when presented with the book text only during training, LLMs' performance on SORT falls short. By allowing to evaluate more aspects of memory, we believe that SORT will aid in the emerging development of memory-augmented models.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2410.08133 [cs.CL]
	(or arXiv:2410.08133v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2410.08133

Submission history

From: Mathis Pink [view email]
[v1] Thu, 10 Oct 2024 17:17:38 UTC (1,229 KB)

Computer Science > Computation and Language

Title:Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators