← All articles

RESEARCH / MULTIMODAL MEMORY

Simplex VM0First-person experience for understanding and decisions.

A memory foundation for spatial intelligence.

Evaluated on EgoLifeQA, ATM-Bench and Mem-Gallery. Three settings examine different dimensions of long-term memory.

VM0 / MULTIMODAL MEMORYCONCEPTUAL OVERVIEW

01 / EXPERIENCE

Visual contextt − 3
Conversationt − 2
Life recordt − 1

02 / MEMORY

HTDM
LDM
TEXT + MULTIMODAL

Experience across time

03 / RETRIEVAL

QA question in the present
01
02
03

Context for understanding and decisions

Experiences → memory → relevant context

Simplex VM0 explores multimodal long-term memory for robots. It combines textual and multimodal records so that past first-person experience can inform present understanding, spatial context and decisions. This research direction brings together LDM (Lifelong Developmental Memory) and HTDM (Hierarchical Temporal Developmental Memory).

01 / EVALUATION

Evidence across three memory settings.

01 / FIRST-PERSON LIFELOG

EgoLifeQA

Questions about events, people and temporal relationships in first-person daily-life records.

84.20% A1_JAKE-500/ 74.01% manual-2.9K

VM0 EgoLifeQA results on A1_JAKE-500 and manual-2.9K
Figure 1. EgoLifeQA. Accuracy on A1_JAKE-500 (500 questions) and manual-2.9K (2,905 questions, six participants). Inputs are video descriptions and lifelog text, truncated at the question timestamp.
02 / CROSS-SOURCE PERSONAL MEMORY

ATM-Bench

Retrieval and reasoning across personal memories from images, videos and email.

85.78% regular QS/ 66.91% Hard QS

VM0 ATM-Bench regular and Hard results
Figure 2. ATM-Bench. Question score (QS) from internal development evaluations. The regular and Hard sets use different pipelines.
03 / MULTIMODAL CONVERSATION

Recovering textual and visual details from long-running multimodal conversations.

91.64% normalized mean judge score

VM0 Mem-Gallery results
Figure 3. Mem-Gallery. Normalized mean judge score across 1,711 questions.