07/06 01:55 The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems 
06/29 07:56 Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement 
06/23 12:51 HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions 
06/15 14:09 Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning 