ai.memory-localmemory-ts
purpose
Conversation history management with LocalMemory, message limits, and auto-summarization.
rules
- Import
LocalMemoryfrom@microsoft/teams.ai. This is the built-in memory class that implements theIMemoryinterface for managing conversation history with automatic overflow handling. - Pass a
maxvalue to theLocalMemoryconstructor to cap the number of messages retained. When the limit is reached, the collapse strategy is triggered automatically. Choose a value that balances context quality with token budget (e.g., 20-50 messages for typical chat bots). - Set
collapse.strategyto'half'(default) to summarize and discard the oldest half of messages when the limit is hit, or'full'to summarize all messages into a single summary message. The'half'strategy preserves recent context while the'full'strategy maximizes compression. - Provide a
collapse.model-- anOpenAIChatModelinstance used to generate the summary when collapse is triggered. This can be the same model used for chat or a cheaper/faster model dedicated to summarization. - Pass the
LocalMemoryinstance as themessagesproperty of theChatPromptconstructor. The prompt reads from and writes to this memory automatically on eachprompt.send()call. - For multi-turn bots, maintain a
Map<string, LocalMemory>keyed by conversation ID. Create a newLocalMemoryper conversation to prevent history leaking across users or channels. - Use the
IMemoryinterface methods (push,pop,get,set,delete,values,length,where,collapse) for programmatic access to conversation history. Callmemory.where(predicate)to filter messages by role or content. - Seed initial context by passing a
messagesarray to theLocalMemoryconstructor. Use this for few-shot examples or system-level context that should always be present at the start of a conversation. - Call
memory.collapse()manually when you need to free token budget mid-conversation (e.g., before a large function call result). The method returns the summary message orundefinedif collapse was not needed. - For production deployments that must survive restarts, serialize
memory.values()to persistent storage (database, blob) and rehydrate by passing the stored messages array to a newLocalMemoryconstructor.
patterns
Basic LocalMemory with collapse
import { LocalMemory, ChatPrompt } from '@microsoft/teams.ai';
import { OpenAIChatModel } from '@microsoft/teams.openai';
const model = new OpenAIChatModel({
apiKey: process.env.OPENAI_API_KEY,
model: 'gpt-4o',
});
const summaryModel = new OpenAIChatModel({
apiKey: process.env.OPENAI_API_KEY,
model: 'gpt-4o-mini',
});
const memory = new LocalMemory({
max: 50, // Keep up to 50 messages
messages: [], // Optional initial messages
collapse: {
strategy: 'half', // Summarize oldest half when full
model: summaryModel, // Model used for summarization
},
});
const prompt = new ChatPrompt({
model,
instructions: 'You are a helpful assistant.',
messages: memory,
});
const result = await prompt.send('Hello!');Per-conversation memory with Map
import { LocalMemory, ChatPrompt, Message } from '@microsoft/teams.ai';
const conversationMemories = new Map<string, LocalMemory>();
app.on('message', async ({ send, activity }) => {
const convId = activity.conversation.id;
// Get or create per-conversation memory
if (!conversationMemories.has(convId)) {
conversationMemories.set(convId, new LocalMemory({
max: 30,
collapse: {
strategy: 'half',
model: summaryModel,
},
}));
}
const prompt = new ChatPrompt({
model,
instructions: 'You are a helpful assistant.',
messages: conversationMemories.get(convId)!,
});
const result = await prompt.send(activity.text);
if (result.content) {
await send(result.content);
}
});IMemory interface methods
// Push a message manually
memory.push({ role: 'user', content: 'Hello' });
// Get message count
const count = memory.length();
// Retrieve all messages
const allMessages = memory.values();
// Filter messages by role
const userMessages = memory.where((msg) => msg.role === 'user');
// Get a specific message by index
const first = memory.get(0);
// Replace a message at index
memory.set(0, { role: 'system', content: 'Updated context' });
// Remove the last message
memory.pop();
// Delete message at index
memory.delete(2);
// Manually trigger collapse/summarization
const summary = await memory.collapse();pitfalls
- Sharing a single LocalMemory across conversations: All users see each other's history. Always key memory instances by conversation ID (or user ID for 1:1 bots).
- Setting
maxtoo low: A max of 5-10 causes frequent collapse, losing important context. Start with 20-50 and tune based on your token budget and average conversation length. - Setting
maxtoo high: Exceeding the model's context window causes truncation errors or degraded response quality. Keepmax * average_message_tokenswell under the model's context limit. - Forgetting
collapse.model: If you set a collapse strategy but omit the model, summarization will fail silently and old messages will simply be dropped instead of summarized. - Memory lost on restart:
LocalMemoryis in-memory only. Bot process restarts lose all conversation history. For production, serializememory.values()to a database and rehydrate on startup. - Passing a raw
Message[]instead ofLocalMemory: Passing a plain array asmessagesworks for simple cases but you lose collapse, max limits, and theIMemoryinterface. UseLocalMemoryfor anything beyond trivial demos. - Not cleaning up stale conversations: The
Mapgrows indefinitely. Implement a TTL or LRU eviction policy to remove inactive conversation memories.
references
- Teams AI Library v2 -- GitHub
- @microsoft/teams.ai -- npm
- OpenAI Context Window Limits
- Conversation History Best Practices -- Microsoft Learn
instructions
This expert covers conversation history management with LocalMemory in Teams AI v2. Use it when you need to:
- Configure
LocalMemorywith max message limits and collapse strategies - Choose between
'half'and'full'collapse strategies for summarization - Implement per-conversation or per-user memory isolation using a
Map - Use the
IMemoryinterface methods for programmatic history access - Seed conversations with initial context messages
- Persist and rehydrate conversation history across bot restarts
Pair with ai.chatprompt-basics-ts.md for passing memory to ChatPrompt constructor, and state.storage-patterns-ts.md for persisting conversation history across restarts.
research
Deep Research prompt:
"Write a micro expert on memory in Teams AI (TypeScript). Cover LocalMemory configuration, max messages, collapse strategies (half/full), supplying a summarization model, and state scoping (per-user vs per-conversation). Include practical code patterns and warnings about memory leakage across conversations."