Key Features

Schemas and entities organize information about people, facts, and knowledge.
Separate nodes represent emotion, preferences, personality, and relationships.
The library includes transcription, voiceprint, scene, emotion, and embedding processing.
The streaming interface can prefetch relevant memory before a turn ends.
Routing and ranking select only the most relevant top-k memories.
The library is installable as voicemem.
The repository includes a local web application for testing voice memory.
The project releases ChatMem-400K and associated memory-aware model resources.

Its streaming dual-part architecture separates schema-and-entity factual memory from emotion and persona nodes. Audio processing includes transcription, speaker information, and other signals; retrieval routes and ranks memory before injecting only a small top-k subset into context. Streaming lookup can begin before the speaker finishes a turn.


VoiceMem is useful for developers building personalized voice assistants, digital humans, and interactive devices. The public package and repository include a web demo, model resources, and the ChatMem dataset family. Local components and external model calls have distinct setup and operating costs despite the free software release.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!