CodeEvolver proposes changes to your prompts and code, tests each one against your data, and keeps only what wins — then proves it on data the engine never saw, and hands you a PR to review.
A reflection agent analyzes failures in your system. A coding agent proposes candidate mutations to prompts, code, and model configuration. An evolutionary optimizer scores every candidate against your evaluation dataset and keeps only the winners.
The output is a pull request with the full experiment history and associated scores attached.
A real optimization run on the PhantomWiki benchmark — three rounds, nine candidates evaluated, a 59% relative lift over baseline, validated on a holdout test set.