CodeEvolver builds self-improving AI systems for industries where mistakes are costly, powered by our proprietary Artificial Selection Algorithm.

optimizedbaseline
scoreRelative Gain80%38.4%69.0%baselineoptimized
80%improvement on PhantomWiki benchmark
Synthetic knowledge-base QA · 0.38 → 0.69
The task: QA over a generated corpus — fully reproducible, contamination-free evals.
What CodeEvolver changed: direct passage extraction, document reranking, and an added reflection step.
+80%
PhantomWiki
0.38 → 0.69
+85%
HoVer
0.47 → 0.88
+64%
RAG QA Arena - Tech
failures 7.2% → 2.6%

CodeEvolver powers improvements over many iterations.

Our Artificial Selection Algorithm runs deliberate experiments and tests each one against your data, refining your code over many cycles before handing you a PR to review. Learn about the algorithm here.

0.380.69accuracyaccepted changes →optimized
codeevolver — phantomwiki run
BASELINEEvaluating the starting program
score 0.38
PROMPTRewrote the agent's reasoning instructions
evaluate → score 0.46 ▲ +0.08 · kept
PROMPTTightened how it extracts the final answer
evaluate → score 0.54 ▲ +0.08 · kept
CODEAdded a second pass to catch answers it missed
evaluate → score 0.61 ▲ +0.07 · kept
PROMPTSharpened the follow-up search strategy
evaluate → score 0.66 ▲ +0.05 · kept
CODEAdded answer-normalization post-processing
evaluate → score 0.69 ▲ +0.03 · kept
optimization complete · best 0.69

We cut a LegalTech company's AI failures by 69%.

In legal, accuracy isn't optional — a wrong answer can cost a case and a client.

The client was managing a complex rules engine for parsing legal documents.

CodeEvolver autonomously created an updated rules engine that reduced failures by 69% on a challenging set of test cases. See case study.

Before
59% accuracy on difficult, failure-prone cases
codeevolver
Changes made
+ 685 new rules
+ 21 topics
+ 12 prompts optimized
✓ 18,000 tests run
After
88% accuracy on the same evaluation set
Incorrect output Correct output
Before CodeEvolver, the AI was accurate 59% of the time. CodeEvolver autonomously made changes spanning 685 new rules, 21 topics, and 12 optimized prompts, validated against 18,000 test cases. After, the AI scores 88% on the same evaluation set, a 29-point gain in accuracy.

We run the optimization. You approve the PR.

Most teams engage us end-to-end: we build the evals, run the optimization, and review the code before handing off a PR for your approval. We offer self-service for teams that want to run the optimization themselves.

S1
Data Curation
Evaluation datasets curated from your observability logs, simulations, and domain experts — with ground-truth labels where your reward function needs them.
S2
Eval-Driven Engineering
We define the reward functions and constraints that make optimization trustworthy — objective metrics your team agrees measure what “better” actually means.
S3
The Optimization Engine
CodeEvolver runs thousands of experiments across your prompts, code, and configuration — and delivers the winning system as a pull request with the full experiment history.
S4
Private Deployment
For teams in regulated industries: run everything inside your own walls. Your code and data never leave your environment.

Built for teams where
AI errors are expensive.

The answers worth knowing

Short, direct answers to the things engineering leaders want to know. If yours isn't here, ask us.

No. The engine works against your evaluation datasets in an isolated environment. Changes only reach your codebase as a pull request your team reviews and merges like any other PR.
A reflection agent studies your system's failures; a coding agent drafts the fixes. Every change is a deliberate experiment, not a guess.
Your data and any constraints you set. Every candidate is scored against your evals — winners survive, losers are discarded.
A pull request with the full diff and every experiment's score. Your team reviews and merges.
No — building it is part of the engagement. We curate evaluation datasets from a variety of sources, including your observability logs, public sources, and synthetic data. It's also the most valuable artifact you keep, besides the optimization output itself.
We split eval data into training, validation, and holdout splits, so we can measure any overfitting. Final numbers are reported on a holdout set the engine has never seen — the same discipline we use in our published benchmark results.
Prompts (instructions, few-shot examples), code (context pipelines, reasoning steps, tools, module architecture), and model configuration — jointly. Our research shows joint optimization beats prompt-only optimization by 5–15 points on public benchmarks. You can also set any constraint to limit the kind of changes the engine can make.
Typically 4-6 weeks end-to-end: ~1-2 weeks to set up data and evals (if not already available), ~1 week to run the optimization, and then 2-4 weeks to review, integrate, and test.
Three ways. We optimize the whole system, not just prompts — that's worth 5–15 extra points in our experiments. We operate as a research-backed service, not a library you have to integrate and learn. And we validate every result on held-out data before you see it.
We charge based on outcomes. Projects are scoped per AI workflow. Talk to us — pricing depends on factors like the optimization complexity and whether you need a private deployment.
Yes. Private deployment options keeps your code and data entirely inside your walls. Additionally, we provide a data processing agreement that outlines all that we do to keep your data safe and confidential. Nothing from your system is used to train our models.
Start With a Single Use Case

Make your software
measurably more reliable.

Get a Baseline AssessmentSee the Research