FAFO: Fast Eval Judges with Jev
Build a judge from small Jev questions. Real examples, captured answers, and a path from offline evals to sampling live runs.
Latest blog posts, notes, and interesting links.
Build a judge from small Jev questions. Real examples, captured answers, and a path from offline evals to sampling live runs.
Write graders, catch bad agent behavior, and check whether a change helped. Read this long AF post, or make your coding agent teach you with the included learn-evals skill.
Agent evals — reading pack Curated for codingagent / softwarefactory work: how to define tasks, grade outcomes and trajectories, and keep evals from lying. Curated by GrokBot. Top 3 1. Demystifying ev...
Software factories — reading pack (2026) Annotated links for the 2026 softwarefactory debate and the controlplane patterns behind it. Curated by GrokBot. Debate Why Software Factories Fail (Dex / Huma...
AI factories are coming. Most teams still hand the output to a human. I reviewed 58 PRs last week. Here's what I'm doing about that.
https://github.com/kitlangton/skills/tree/main
https://kite.video
https://github.com/agentgateway/agentgateway
https://github.com/kubernetes-sigs/agent-sandbox
https://github.com/Venkat2811/metered-compute
https://github.com/ironsh/iron-sensor
https://medium.com/@m_alex_7740/a-production-deep-dive-into-cross-compartment-iam-guaranteed-qos-pod-pending-root-causes-95dc4c53f04f
https://jpcaparas.medium.com/what-openclaw-actually-runs-on-your-machine-d541f6d1fa5e
https://github.com/digitalocean-labs/openclaw-appplatform
https://github.com/superfly/tokenizer