How to handle conversation memory in LangChain for long-context LLM applications?

Asked 22 days ago Updated 18 hours ago 121 views

0

When building conversational assistants with Large Language Models, keeping track of conversation history without exceeding token context limits is a common challenge. Standard buffer memory simply appends every user and assistant message to the prompt, which quickly exhausts the context window or inflates API costs.

I am working on a customer support bot where conversations can last dozens of turns. I want to know how developers implement robust conversation memory in LangChain when dealing with long user sessions.

Here is my current basic setup using Python and LangChain summary buffer memory:

from langchain_openai import ChatOpenAI
from langchain.memory import ConversationSummaryBufferMemory

# Initialize the LLM with low temperature for deterministic context summarizing
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.0)

# Set up summary buffer memory with a strict token threshold
memory = ConversationSummaryBufferMemory(
    llm=llm,
    matoken_limit=300  # Summarizes older turns when context exceeds 300 tokens
)

# Simulate adding interactions to the conversation history
memory.save_context(
    {"input": "Our deployment failed on stage 2 with a database connection timeout."}, 
    {"output": "Please verify the PostgreSQL connection string and subnet security groups."}
)

# Retrieve current context formatted for prompt injection
current_history = memory.load_memory_variables({})
print(current_history["history"])

1 Answer


0

For a long-running conversation, don’t keep appending every message to the prompt. Even when the model accepts a large context window, old turns consume tokens, add noise, and can push out the information needed for the current question.

A useful design separates recent conversation from durable memory. Keep a bounded window of recent turns for conversational flow, and save only useful facts or summaries separately when the application needs them later.

Start with session-scoped history

For a straightforward LangChain chain, RunnableWithMessageHistory can load and save messages by session ID. This in-memory example shows the wiring; it is not suitable for production persistence, and it does not limit the history by itself.

from langchain_core.chat_history import InMemoryChatMessageHistory
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.runnables.history import RunnableWithMessageHistory

# Keep a separate message history for each conversation.
store = {}

def get_session_history(session_id: str) -> InMemoryChatMessageHistory:
    if session_id not in store:
        store[session_id] = InMemoryChatMessageHistory()
    return store[session_id]

prompt = ChatPromptTemplate.from_messages([
    ("system", "Answer the user clearly and use the conversation when relevant."),
    MessagesPlaceholder(variable_name="history"),
    ("human", "{input}"),
])

chain = prompt | model  # Replace model with your chat model instance.

chat = RunnableWithMessageHistory(
    chain,
    get_session_history,
    input_messages_key="input",
    history_messages_key="history",
)

reply = chat.invoke(
    {"input": "What did we decide about the launch date?"},
    config={"configurable": {"session_id": "conversation-42"}},
)

Keep the prompt within a budget

Before each model call, trim history to a token budget that leaves room for the system instructions, the current request, and the model’s response. LangChain’s trim_messages utility can help; choose the token counter and budget for the model you are actually using. Preserve complete turns where possible rather than cutting a message in half.

When older context still matters, summarize it instead of retaining every raw message. A practical pattern is to keep a compact conversation summary plus the newest few turns, updating the summary as the conversation grows. Treat summaries as fallible: keep important user-provided details explicit, and let the user correct them.

For agents and production apps

  • Use a persistent store. In-memory history disappears when the process restarts. Store conversations by user and session, with appropriate access controls and retention rules.
  • Use LangGraph checkpointing for agent state. A checkpointer can persist state between turns; pass a stable thread_id so the right conversation is resumed. This is distinct from adding all past messages to every prompt.
  • Keep durable facts selective. If users need information recalled across separate conversations, retrieve relevant saved facts or documents rather than injecting an entire chat archive.
  • Test retrieval and token use. Check that the system recalls relevant details, avoids unrelated history, and stays within the model’s context limit.

The right memory setup depends on what the application must remember. Recent turns support continuity; summaries and retrieved facts support longer-term recall. Keeping those jobs separate makes context easier to control and easier to debug.

Write Your Answer