Back to News
RSS feedhuggingface.co

IBM Research Finds AI Agent Memory Must Be Calibrated to Model Capability

Summary

IBM Research evaluates ALTK-Evolve, which distills an agent's past trajectories into inference-time guidelines. Across eight models and 585 AppWorld tasks, strong models benefited from full guideline sets, while weaker models performed best with compact cores and task-specific retrieval. gpt-oss-120b gained 16.1 percentage points with only 5% more tokens, whereas GLM-5 showed no measurable gain.