Same evidence, different memory: an agent experiment on a MacBook
Two local models, 270 runs, and what changed when identical reports arrived in a different order. Full methods, results, and limitations.
Notes, reflections, and questions on life, technology, and learning.
Same evidence, different memory: an agent experiment on a MacBook
Two local models, 270 runs, and what changed when identical reports arrived in a different order. Full methods, results, and limitations.
Terminal-Bench 4.0: Why Agent Benchmarks Need Maintenance, Not Hype
A beginner-friendly breakdown of resource calibration, task fixes, saturated tasks, and why benchmarks should behave more like software.
Prime Agent and RLMs: The Complete Beginner's Guide
Understanding why the next jump in AI agents may come from the harness around the model, not only the model itself.
Yusuf (12:86): I Complain of My Anguish and Sorrow Only to Allah
A reflection on grief, sabr, tawakkul, and the kind of hope that still gets up and searches.