🔴 What was problem with ReAct?
Agents in ReAct keep a detailed “diary”: they think, take an action (search, click), record the result, and repeat the cycle.
This makes the process transparent, but in long tasks the history quickly grows → context limit → loss of details.
🟢 What is ReSum’s solution?
☞ When the context nears the limit, the agent stops and writes a summary: verified facts + still open questions.
+4.5% quality improvement compared to ReAct
Pass@1: 33.3% and 18.3% on the challenging BrowseComp tests
☞ Then it continues from this summary instead of a long conversation.
☞ A separate 30B model for summaries that better handles “noisy” pages and highlights important information.
☞ Reinforced learning ReSum-GRPO: up to +8.2% with ReSum-GRPO : the agent receives a reward only for the final answer, which is distributed across all intermediate steps. This teaches it to gather correct facts and create concise, useful summaries.
Agents stay within token budget → solve complex web search → analysis tasks better than classic ReAct.
🤖 Data Science, ML & Big Data with @DataXplore
