When I have a can of Spam and a free morning, I make a riff on a bacon, egg, and cheese sandwich. I sear a thick slice of ...
Skill Eval Harness is a Python CLI for testing whether an Agent Skill changes observable output. It reads evals/shared-benchmark.json, emits answer-key-safe task rows, grades files under eval-runs/, ...