Define a score
Provide evaluate(candidate_path). Test correctness, measure a benchmark, or combine the criteria that matter to you. Higher is better.
Give us an evaluator.
We’ll evolve the code.
From a measurable objective to a better Python program.
No GPU. No model API keys. No infrastructure to manage.
Python SDK · OpenEvolve engine · A budget you control
from easyevolve import EasyEvolve
result = EasyEvolve(
evaluator="eval.py"
).run(iterations=100, budget=2.0)
print(result.best_code)
print(result.best_score)def solve(values):
return valuesSelect the next step to inspect a mutation.
Write down what good looks like.
Let a population of programs work toward it.
Provide evaluate(candidate_path). Test correctness, measure a benchmark, or combine the criteria that matter to you. Higher is better.
Choose a model level and a maximum budget. The service generates an initial candidate, then selects, mutates, and evaluates programs in isolated sandboxes.
Get the best program, its score, and an iteration history. Reconnect with a job ID, or stop the search whenever you have enough.
Each external operation reserves its price cap before it starts. When the next operation won’t fit, the search stops.
Candidate programs and your evaluator run together in a fresh remote VM with outbound internet disabled.
Stripe Checkout funds prepaid credits. A saved device token lets later runs use your available balance.
The job keeps running when your Python process disconnects. Use its ID to reconnect and retrieve the result.