EXAMPLES / START WITH A SCOREMake the objective
Run your first search ↗
Make the objective
executable.
Two standard-library evaluators you can adapt.
Keep correctness in the score you optimize.
01 / CORRECTNESS
A sorting function
Ask for solve(values). Score the proportion of tests passed. A docstring makes the contract legible even without a task description.
Useful for learning the API. Five test cases are a demonstration, not a comprehensive correctness proof.
"""Implement solve(values: list[int]) -> sorted list[int]."""
import importlib.util
def evaluate(candidate_path: str) -> dict:
spec = importlib.util.spec_from_file_location("candidate", candidate_path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
cases = [[], [1], [3, 1, 2], [5, -1, 5, 0], list(range(30, 0, -1))]
passed = sum(module.solve(x.copy()) == sorted(x) for x in cases)
return {"score": passed / len(cases)}02 / PERFORMANCE
Faster Fibonacci
Require correct answers first. Then maximize the inverse execution time across repeated calls. Include a broader validation suite before using the result in production.
Timing varies across sandboxes. Use multiple samples and stable inputs for meaningful measurements.
"""Implement fibonacci(n: int) -> int for n >= 0."""
import importlib.util
import time
def evaluate(candidate_path: str) -> dict:
spec = importlib.util.spec_from_file_location("candidate", candidate_path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
for n, expected in [(0, 0), (1, 1), (10, 55), (20, 6765)]:
if module.fibonacci(n) != expected:
return {"score": 0.0}
start = time.perf_counter()
for _ in range(200):
if module.fibonacci(20) != 6765:
return {"score": 0.0}
elapsed = time.perf_counter() - start
return {"score": 1.0 / (elapsed + 0.001)}