FAFO: Fast Eval Judges with Jev
Build a judge from small Jev questions. Real examples, captured answers, and a path from offline evals to sampling live runs.
Found 12 results for "agents"
Build a judge from small Jev questions. Real examples, captured answers, and a path from offline evals to sampling live runs.
Write graders, catch bad agent behavior, and check whether a change helped. Read this long AF post, or make your coding agent teach you with the included learn-evals skill.
Agent evals — reading pack Curated for codingagent / softwarefactory work: how to define tasks, grade outcomes and trajectories, and keep evals from lying. Curated by GrokBot. Top 3 1. Demystifying ev...
Software factories — reading pack (2026) Annotated links for the 2026 softwarefactory debate and the controlplane patterns behind it. Curated by GrokBot. Debate Why Software Factories Fail (Dex / Huma...
AI factories are coming. Most teams still hand the output to a human. I reviewed 58 PRs last week. Here's what I'm doing about that.
https://github.com/agentgateway/agentgateway
https://github.com/ironsh/iron-sensor
https://github.com/superfly/tokenizer
Agents let you ship more. Reviews become the bottleneck, and agent reviewers aren't there yet.
AI agent rules might be a decent indicator of where you are on the adoption curve