ai.streaming-ts
purpose
Real-time streaming of AI responses with typing indicators and progressive rendering.
rules
- Use the
onChunkcallback inprompt.send()options to receive text chunks as they arrive from the LLM. Each chunk is astringfragment of the ongoing response. - Inside
onChunk, callstream.emit(chunk)to send the accumulated text to the user with a typing indicator. Thestreamobject is available on the handler context (ctx.stream). stream.emit()accepts either a plainstringor aMessageActivityinstance. UseMessageActivitywhen you need to attach feedback buttons, AI-generated markers, or citations to the streaming message.- Call
stream.update(text)to send a status update (e.g.,"Thinking...","Searching documents..."). Status updates are separate from the accumulated content and display as informative indicators. stream.close()is called automatically when the message handler returns. It sends the final message containing all accumulated content, attachments, and entities. You do not need to call it manually in typical usage.- If you need to finalize the stream early (e.g., after an error), call
stream.close()explicitly. After close, furtheremit()calls are ignored. - Streaming works internally by batching: content is queued and flushed in batches of up to 10 items every 500ms. Text accumulates across chunks so the final message contains the complete response.
- Listen to stream events with
stream.events.on('chunk', handler)for each sent chunk andstream.events.once('close', handler)for the final message. Use these for logging, analytics, or post-processing. - When combining streaming with
MessageActivityfeatures (feedback, citations), construct a newMessageActivityin eachonChunkcall. The stream accumulates content across emissions automatically. - Do not call
await send()for the final message when streaming --stream.close()handles it. Calling bothsend()and allowing the auto-close results in duplicate messages.
patterns
Basic text streaming with onChunk
import { ChatPrompt } from '@microsoft/teams.ai';
app.on('message', async ({ send, stream, activity }) => {
const prompt = new ChatPrompt({ model, instructions: 'You are a helpful assistant.' });
// Stream chunks as they arrive
const response = await prompt.send(activity.text, {
onChunk: (chunk: string) => {
stream.emit(chunk); // Sends typing indicators with accumulated text
},
});
// stream.close() is called automatically after the handler returns,
// sending the final message with all accumulated content
});Streaming with feedback buttons and AI markers
import { MessageActivity } from '@microsoft/teams.api';
app.on('message', async ({ stream, activity }) => {
const prompt = new ChatPrompt({ model, instructions: 'You are a helpful assistant.' });
const response = await prompt.send(activity.text, {
onChunk: (chunk: string) => {
// Emit a MessageActivity with feedback buttons on each chunk
stream.emit(new MessageActivity(chunk).addFeedback());
},
});
// Final message automatically includes feedback buttons
});Stream API with status updates and event listeners
app.on('message', async ({ stream, activity }) => {
// Show a status while the LLM is thinking
stream.update('Searching documents...');
const prompt = new ChatPrompt({ model, instructions: 'You are a research assistant.' });
// Listen for stream events
stream.events.on('chunk', (sentActivity) => {
console.log('Chunk sent to user');
});
stream.events.once('close', (sentActivity) => {
console.log('Final message delivered:', sentActivity.id);
});
const response = await prompt.send(activity.text, {
onChunk: (chunk: string) => {
stream.emit(chunk);
},
});
// stream.close() sends the final message automatically
});pitfalls
- Calling
send()after streaming: If you callawait send(response.content)after streaming, the user receives a duplicate final message. The auto-close onstream.close()already sends the complete response. - Forgetting
stream.emit()insideonChunk: DefiningonChunkwithout callingstream.emit()means the user sees nothing until the final message. TheonChunkcallback alone does not send anything to the client. - Calling
stream.close()too early: Explicitly closing the stream beforeprompt.send()resolves discards remaining chunks. Only callclose()manually for error bailout scenarios. - Heavy computation in
onChunk: The callback fires on every token. Expensive operations (API calls, database writes) insideonChunkcreate backpressure and degrade streaming performance. Log or buffer instead. - Not handling errors during streaming: If the LLM request fails mid-stream, the user sees partial text with no indication of failure. Wrap
prompt.send()in try/catch and callstream.emit('An error occurred.')followed bystream.close()in the catch block. - Assuming chunk boundaries are semantic: Chunks are raw token fragments, not words or sentences. Do not parse or process individual chunks as complete text units.
- Ignoring batching behavior: The SDK batches up to 10 items every 500ms. Very rapid
emit()calls do not produce 1:1 client updates. This is normal and expected.
references
- Teams AI Library v2 -- GitHub
- Teams Streaming Protocol -- Microsoft Learn
- @microsoft/teams.ai -- npm
- OpenAI Streaming -- API Reference
instructions
This expert covers real-time streaming of AI responses in Teams AI v2. Use it when you need to:
- Stream LLM responses to the user with typing indicators using
onChunkandstream.emit() - Display status updates during long-running operations with
stream.update() - Combine streaming with
MessageActivityfor feedback buttons and AI-generated markers - Understand the internal batching mechanism (10 items / 500ms) and its effect on UX
- Handle errors gracefully during streaming
- Use stream events (
chunk,close) for logging and analytics
Pair with ai.chatprompt-basics-ts.md for prompt.send() with onChunk, ai.citations-feedback-ts.md for combining streaming with feedback buttons, and runtime.routing-handlers-ts.md for ctx.stream.
research
Deep Research prompt:
"Write a micro expert on streaming AI responses in Teams SDK v2 (TypeScript). Explain how ctx.stream works, how onChunk accumulates text, how to emit MessageActivity vs strings, and how to combine streaming with typing indicators, final messages, and error handling. Include at least two patterns: (1) plain text streaming, (2) streaming with addAiGenerated/addFeedback."