OpenAI added a guide on Context Engineering for the Agents SDK to their cookbook, and this is probably the most competent approach to memory management.
Instead of rummaging through thousands of old messages, the agent maintains a structured user profile and a "notebook".
🟢 How it works? benifits & Pitfalls
☞ State Object: a centralized information hub in the form of a JSON object, which is stored locally. It includes a profile (hard facts: name, ID, loyalty status) and notes (unstructured notes: "likes hotels in the center").
☞ Injection: before each launch, this state is fed into the system prompt in YAML format: for the profile and Markdown for the notes. Not everything at once, of course, but only what is needed at the moment.
☞ Distillation: the most interesting part. The agent doesn't just chat, it has a tool save_memory_note. If you said in the conversation: "I don't eat meat", the agent calls this tool and saves the Session Note (temporary note) in real time.
☞ Consolidation: garbage collection for memory. After the session ends, a separate process is launched, which takes the temporary notes, compares them with the global ones, removes duplicates and resolves conflicts according to the principle "the newer overrides the older".
BENEFITS:
☞ agent starts behaving like a personal assistant without retraining.
☞ There are clear rules: what the user said now > session notes > global settings.
☞ We don't mix everything together, but separate hard data (for example, from CRM) and soft data (preferences from the chat).
OpenAI's approach of dividing into Session Memory and Global Memory looks reliable, but requires direct hands in writing the consolidation logic. Without this, your agent will quickly turn into a demented grandfather who remembers what never happened.
PITFALLS:
You need to make a separate call to the LLM after each dialogue to tidy up the memory. If the model glitches at this stage, it may write a hallucination into the "long memory" or delete something important. Here, strict frameworks are needed.
The context window is not elastic. Although models have a huge context, dragging "War and Peace" from the user's notes is costly in terms of money and timing. You will have to periodically trim the history, leaving only the essence.
Guide • #AI #ML #LLM #Guide #OpenAI
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
