07/06 01:55 The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems 
06/29 07:56 Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement 
06/23 12:51 HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions 