RSS feedhuggingface.co
IBM Research Finds AI Agent Memory Must Be Calibrated to Model Capability
Summary
IBM Research evaluates ALTK-Evolve, which distills an agent's past trajectories into inference-time guidelines. Across eight models and 585 AppWorld tasks, strong models benefited from full guideline sets, while weaker models performed best with compact cores and task-specific retrieval. gpt-oss-120b gained 16.1 percentage points with only 5% more tokens, whereas GLM-5 showed no measurable gain.