EXAMPLES / START WITH A SCORE

Make the objective
executable.

Two standard-library evaluators you can adapt.
Keep correctness in the score you optimize.

01 / CORRECTNESS

A sorting function

Ask for solve(values). Score the proportion of tests passed. A docstring makes the contract legible even without a task description.

Useful for learning the API. Five test cases are a demonstration, not a comprehensive correctness proof.

eval.py
"""Implement solve(values: list[int]) -> sorted list[int]."""
import importlib.util

def evaluate(candidate_path: str) -> dict:
    spec = importlib.util.spec_from_file_location("candidate", candidate_path)
    module = importlib.util.module_from_spec(spec)
    spec.loader.exec_module(module)
    cases = [[], [1], [3, 1, 2], [5, -1, 5, 0], list(range(30, 0, -1))]
    passed = sum(module.solve(x.copy()) == sorted(x) for x in cases)
    return {"score": passed / len(cases)}
02 / PERFORMANCE

Faster Fibonacci

Require correct answers first. Then maximize the inverse execution time across repeated calls. Include a broader validation suite before using the result in production.

Timing varies across sandboxes. Use multiple samples and stable inputs for meaningful measurements.

eval.py
"""Implement fibonacci(n: int) -> int for n >= 0."""
import importlib.util
import time

def evaluate(candidate_path: str) -> dict:
    spec = importlib.util.spec_from_file_location("candidate", candidate_path)
    module = importlib.util.module_from_spec(spec)
    spec.loader.exec_module(module)
    for n, expected in [(0, 0), (1, 1), (10, 55), (20, 6765)]:
        if module.fibonacci(n) != expected:
            return {"score": 0.0}
    start = time.perf_counter()
    for _ in range(200):
        if module.fibonacci(20) != 6765:
            return {"score": 0.0}
    elapsed = time.perf_counter() - start
    return {"score": 1.0 / (elapsed + 0.001)}
Run your first search ↗