One agent builds it. A second, much tougher one compares it to something real you picked — without being told which is which. It keeps going, try after try, until yours is the one that wins.
Sign in with Google · runs in your CI · results arrive as pull requests
Judged blind against a page you admire — the critic is shown both and asked which one it would actually finish signing up to.
Judged blind against the manufacturer’s photograph — two cameras on the same model, so a generation can’t win by getting one angle right.
What improves is the harness — the context boundary, the rules inside it, the tools wired to the model and the retry path when one fails. Scored on cases the builder never saw.
Judged blind against a shipped racing game, on the things a still frame cannot fake — wheel blur, weight on the outside tyre, the dust a slide throws.
Anything can generate a first draft. What makes work good is being measured against something real, repeatedly, by something that isn't trying to please you.
Named, fetchable, comparable — a competitor's live page, an actual photograph, a published score. Given something vague, a critic invents a comparison and approves everything.
Separate context from the builder, labels stripped, shown both. If it can pick out yours, the generation lost — and it says exactly what gave it away.
Each round feeds the last verdict back to the builder. It runs for hours without you, and you read the result rather than supervising the process.
Where it ships, what varies, what budget. Three questions that stop a loop winning the comparison while shipping something you can't use.
Your code never leaves your repository. MacroEvo sends the brief and receives the results — the work happens in your CI, on your key or your Claude subscription.
Google sign-in. No password to manage, nothing to download.
Install the GitHub App on the repos you choose. It sees only what you select, and you can revoke it from GitHub at any time.
The reference worth beating, and the three trap questions. This is the part that decides whether any of it works.
Watch generations converge on the dashboard. Take the pull request when it wins.
The loop doesn't stop at good enough. It stops when the critic can't tell yours from the real thing.