Dramaturgy

Research without a fake bibliography.

We have not published papers. This page is the working stance: agents should be judged against stateful twins, with every side effect in the show report.

Twins, not mocks that forget

A twin keeps service state across the run. The agent writes; the next call reads what it wrote. That is the difference between a stub and a rehearsal.

Traces before scores

Define judging criteria after you can see provider calls, latency, and forbidden effects. A pass with a blocked GitHub check-run is still a story in the prompt book.

One line of code

Reconnect by changing base URLs. If the agent cannot switch houses that cheaply, the eval is testing the wiring, not the performance.