For a long-running conversation, don’t keep appending every message to the prompt. Even when the model accepts a large context window, old turns consume tokens, add noise, and can push out the information needed for the current question.
A useful design separates recent conversation from durable memory. Keep a bounded window of recent turns for conversational flow, and save only useful facts or summaries separately when the application needs them later.
Start with session-scoped history
For a straightforward LangChain chain, RunnableWithMessageHistory can load and save messages by session ID. This in-memory example shows the wiring; it is not suitable for production persistence, and it does not limit the history by itself.
from langchain_core.chat_history import InMemoryChatMessageHistory
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.runnables.history import RunnableWithMessageHistory
# Keep a separate message history for each conversation.
store = {}
def get_session_history(session_id: str) -> InMemoryChatMessageHistory:
if session_id not in store:
store[session_id] = InMemoryChatMessageHistory()
return store[session_id]
prompt = ChatPromptTemplate.from_messages([
("system", "Answer the user clearly and use the conversation when relevant."),
MessagesPlaceholder(variable_name="history"),
("human", "{input}"),
])
chain = prompt | model # Replace model with your chat model instance.
chat = RunnableWithMessageHistory(
chain,
get_session_history,
input_messages_key="input",
history_messages_key="history",
)
reply = chat.invoke(
{"input": "What did we decide about the launch date?"},
config={"configurable": {"session_id": "conversation-42"}},
)
Keep the prompt within a budget
Before each model call, trim history to a token budget that leaves room for the system instructions, the current request, and the model’s response. LangChain’s trim_messages utility can help; choose the token counter and budget for the model you are actually using. Preserve complete turns where possible rather than cutting a message in half.
When older context still matters, summarize it instead of retaining every raw message. A practical pattern is to keep a compact conversation summary plus the newest few turns, updating the summary as the conversation grows. Treat summaries as fallible: keep important user-provided details explicit, and let the user correct them.
For agents and production apps
- Use a persistent store. In-memory history disappears when the process restarts. Store conversations by user and session, with appropriate access controls and retention rules.
- Use LangGraph checkpointing for agent state. A checkpointer can persist state between turns; pass a stable
thread_id so the right conversation is resumed. This is distinct from adding all past messages to every prompt.
- Keep durable facts selective. If users need information recalled across separate conversations, retrieve relevant saved facts or documents rather than injecting an entire chat archive.
- Test retrieval and token use. Check that the system recalls relevant details, avoids unrelated history, and stays within the model’s context limit.
The right memory setup depends on what the application must remember. Recent turns support continuity; summaries and retrieved facts support longer-term recall. Keeping those jobs separate makes context easier to control and easier to debug.