Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions docs/source/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,24 @@ synthetic_problems



## Running examples with uv and Codon (optional)

- Install uv (`https://docs.astral.sh/uv/getting-started/installation/`). The script will automatically use uv to run Python and resolve dependencies from the project.
- Run all examples normally:

```bash
./run_examples.sh
```

- To enable Codon, install Codon (`https://exaloop.io/docs/codon/latest/guide/installation`) and run:

```bash
CODON=1 ./run_examples.sh
```

When `CODON=1`, the script will try to run each example with `codon run -release` when feasible and automatically fall back to `uv run python` for examples that require third-party packages unsupported by Codon.


## Indices and tables

* {ref}`genindex`
Expand Down
4 changes: 4 additions & 0 deletions docs/source/sklearn.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,3 +30,7 @@ model.fit(X, y)


For complete examples, check the `sklearn-type-examples.py` in the examples folder.

## Feature probability weighting

All sklearn-compatible estimators accept an optional parameter `weight_features_by_correlation` (default: `False`). When enabled, terminals corresponding to input features are sampled with probabilities proportional to their absolute Pearson correlation with the output, biasing the grammar towards more predictive variables.
2 changes: 1 addition & 1 deletion examples/geml/classifier_example.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,5 +27,5 @@
model = model_class(max_time=20.0, seed=seed)
model.fit(data, target)
y_pred = model.predict(test_data)
r2 = f1_score(test_target, y_pred)
r2 = f1_score(test_target, y_pred, average="weighted")
print(f"{model}: {r2}")
12 changes: 10 additions & 2 deletions geml/classifiers.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,6 @@
import numpy as np
from typing import Annotated
from geneticengine.grammar.decorators import weight
from geml.common import GeneticEngineEstimator, PopulationRecorder
from geml.grammars.ruleset_classification import make_grammar
from geneticengine.algorithms.gp.gp import GeneticProgramming
Expand All @@ -8,6 +10,7 @@
from geneticengine.evaluation.budget import SearchBudget
from geneticengine.evaluation.tracker import ProgressTracker
from geneticengine.grammar.grammar import Grammar, extract_grammar
from geneticengine.grammar.metahandlers.vars import VarRangeWithProbabilities
from geneticengine.problems import Problem
from geneticengine.random.sources import RandomSource
from geneticengine.representations.tree.initializations import ProgressivelyTerminalDecider
Expand All @@ -17,13 +20,17 @@

class GeneticEngineClassifier(GeneticEngineEstimator):


def get_grammar(self, feature_names: list[str], data, target) -> Grammar:
classes = np.unique(target).tolist()
components, RuleSet = make_grammar(feature_names, classes)
Var = components[-1]
weights = self.correlation_weights(feature_names, data, target)
Var.__init__.__annotations__["name"] = Annotated[str, VarRangeWithProbabilities(feature_names, weights)] # type:ignore
Var.feature_names = feature_names # type:ignore
index_of = {n: i for i, n in enumerate(feature_names)}
Var.to_numpy = lambda s: f"dataset[:,{index_of[s.name]}]" # type:ignore
Var = weight(10)(Var)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Conditional weighting bypasses flag in get_grammar

The get_grammar method unconditionally calls self.correlation_weights(feature_names, data, target) and applies the weights, ignoring the weight_features_by_correlation parameter. The correlation-based weighting should only be applied when self.weight_features_by_correlation is True, otherwise uniform weights should be used.

Fix in Cursor Fix in Web

return extract_grammar(components, RuleSet)

def get_goal(self) -> tuple[bool, float]:
Expand Down Expand Up @@ -56,14 +63,15 @@ def __str__(self):

class HillClimbingClassifier(GeneticEngineClassifier):

def __init__(self, max_time: float | int = 1, seed: int = 0, number_of_mutations: int = 5):
super().__init__(max_time, seed)
def __init__(self, max_time: float | int = 1, seed: int = 0, number_of_mutations: int = 5, weight_features_by_correlation: bool = False):
super().__init__(max_time, seed, weight_features_by_correlation)
self.number_of_mutations = number_of_mutations

_parameter_constraints = {
"max_time": [float, int],
"seed": [int],
"number_of_mutations": [int],
"weight_features_by_correlation": [bool],
}

def search(
Expand Down
31 changes: 29 additions & 2 deletions geml/common.py
Original file line number Diff line number Diff line change
Expand Up @@ -81,13 +81,15 @@ def to_sympy(self):
class GeneticEngineEstimator(GEBaseEstimator):
max_time: float | int

def __init__(self, max_time: float | int = 1, seed: int = 0):
def __init__(self, max_time: float | int = 1, seed: int = 0, weight_features_by_correlation: bool = False):
self.max_time = max_time
self.seed = 0
self.seed = seed
self.weight_features_by_correlation = weight_features_by_correlation

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Seed Parameter Not Applied: Always Zeroed Seed

The seed parameter is hardcoded to 0 instead of using the seed parameter passed to init. Line 86 should be self.seed = seed instead of self.seed = 0. This bug prevents users from setting a custom random seed, causing all instances to use seed=0 regardless of what value is passed to the constructor.

Fix in Cursor Fix in Web


_parameter_constraints = {
"max_time": [float, int],
"seed": [int],
"weight_features_by_correlation": [bool],
}

def get_population(self) -> list[BaseEstimator]:
Expand Down Expand Up @@ -177,3 +179,28 @@ def search(
budget: SearchBudget,
population_recorder: PopulationRecorder,
) -> list[Individual] | None: ...


def correlation_weights(self, feature_names: list[str], data, target) -> list[float]:

def safe_corrcoef(xv, yv) -> float:
with np.errstate(all="ignore"):
x = np.asarray(xv, dtype=float)
y = np.asarray(yv, dtype=float)
if len(x) < 2:
return 0.0
c = np.corrcoef(x, y)
# For 2x2 corr matrix, off-diagonal holds the correlation
try:
corr = float(c[0, 1])
except Exception:
corr = 0.0
if not np.isfinite(corr):
return 0.0
return corr

def wrapper(corr_value: float) -> float:
# Higher absolute correlation -> smaller weight (bias search), add epsilon
return 1 - abs(corr_value) + 0.00001

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Inverted Correlation Weighting Misaligns Sampling Likelihood

The correlation weighting logic is inverted. The documentation states features should be "sampled with probabilities proportional to their absolute Pearson correlation", but the implementation returns 1 - abs(corr_value), which gives LOWER weights to features with HIGHER correlation. For a feature with correlation 0.9, the weight becomes 0.1, while a feature with correlation 0.1 gets weight 0.9. This is inversely proportional to correlation, contradicting the intended behavior. The formula should be abs(corr_value) + 0.00001 instead.

Fix in Cursor Fix in Web


return [wrapper(safe_corrcoef(data[:, i], target)) for i in range(len(feature_names))]
26 changes: 21 additions & 5 deletions geml/grammars/symbolic_regression.py
Original file line number Diff line number Diff line change
@@ -1,9 +1,10 @@
from abc import ABC, abstractmethod
from dataclasses import dataclass
from dataclasses import dataclass, field
from typing import Annotated
from typing import cast

from geneticengine.grammar.decorators import weight
from geneticengine.grammar.metahandlers.vars import VarRange
from geneticengine.grammar.metahandlers.vars import VarRange, VarRangeWithProbabilities


class Expression(ABC):
Expand Down Expand Up @@ -190,19 +191,34 @@ def to_numpy(self) -> str:
return f"{self.value}"


def make_var(options: list[str], relative_weight: float = 1):
def make_var(options: list[str], weights: list[float] | int | float | None = None, relative_weight: float = 1):
# Backward/lenient compatibility: if the second positional argument is a number,
# interpret it as relative_weight and default to uniform feature weights.
if isinstance(weights, (int, float)) and relative_weight == 1:
relative_weight = int(weights)
weights = None

@weight(relative_weight)
@dataclass
class Var(Expression):
name: Annotated[str, VarRange(options)]
# The list of vars should always be filled in dynamically
name: str # Annotation will be set dynamically below
feature_names: list[str] = field(default_factory=list)

def to_sympy(self) -> str:
return f"{self.name}"

def to_numpy(self) -> str:
return f"{self.name}"

# Choose metahandler based on whether weights are provided
if weights is None:
# Use uniform selection (no probabilities)
Var.__init__.__annotations__["name"] = Annotated[str, VarRange(options)]
else:
# Use weighted selection with probabilities
weights_list = [float(w) for w in cast(list[float], weights)]
Var.__init__.__annotations__["name"] = Annotated[str, VarRangeWithProbabilities(options, weights_list)]

return Var


Expand Down
11 changes: 7 additions & 4 deletions geml/regressors.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,10 +35,12 @@ class GeneticEngineRegressor(
):

def get_grammar(self, feature_names: list[str], data, target) -> Grammar:
Var = make_var(feature_names, relative_weight=10)
weights = self.correlation_weights(feature_names, data, target) if self.weight_features_by_correlation else None
Var = make_var(feature_names, weights=weights, relative_weight=10)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Correlation weighting ignored when disabled.

The get_grammar method unconditionally calls self.correlation_weights(feature_names, data, target) and applies the weights, ignoring the weight_features_by_correlation parameter. The correlation-based weighting should only be applied when self.weight_features_by_correlation is True, otherwise uniform weights should be used.

Fix in Cursor Fix in Web


Var.feature_names = feature_names
index_of = {n: i for i, n in enumerate(feature_names)}
Var.to_numpy = lambda s: f"dataset[:,{index_of[s.name]}]"
Var.to_numpy = lambda s: f"dataset[:,{index_of[s.name]}]" # pyright:ignore
complete_components = components + [Var]
return extract_grammar(complete_components, Expression)

Expand Down Expand Up @@ -72,14 +74,15 @@ def __str__(self):

class HillClimbingRegressor(GeneticEngineRegressor):

def __init__(self, max_time: float | int = 1, seed: int = 0, number_of_mutations: int = 5):
super().__init__(max_time, seed)
def __init__(self, max_time: float | int = 1, seed: int = 0, number_of_mutations: int = 5, weight_features_by_correlation: bool = False):
super().__init__(max_time, seed, weight_features_by_correlation)
self.number_of_mutations = number_of_mutations

_parameter_constraints = {
"max_time": [float, int],
"seed": [int],
"number_of_mutations": [int],
"weight_features_by_correlation": [bool],
}

def search(
Expand Down
7 changes: 6 additions & 1 deletion geneticengine/random/sources.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,12 @@ def choice_weighted(
choices: list[T],
weights: list[float],
) -> T:
acc_weights: list[int] = [int(x * 100000) for x in accumulate(weights)]
# Sanitize weights: replace non-finite values with 0 and ensure non-negative totals.
sanitized: list[float] = [w if isinstance(w, (int, float)) and math.isfinite(w) and w > 0 else 0.0 for w in weights]
# If all weights are zero or negative, fall back to uniform selection.
if not any(sanitized):
return self.choice(choices)
acc_weights: list[int] = [int(x * 100000) for x in accumulate(sanitized)]
total = acc_weights[-1]
rand_value: float = self.randint(0, total)

Expand Down
27 changes: 25 additions & 2 deletions run_examples.sh
Original file line number Diff line number Diff line change
@@ -1,7 +1,24 @@
#!/bin/bash
export PYTHONPATH="${PYTHONPATH:+${PYTHONPATH}:}."

PYTHON_BINARY=python3
# Preferred Python runner: uv (falls back to system python3)
if command -v uv >/dev/null 2>&1; then
PYTHON_CMD=(uv run -q -- python)
else
PYTHON_CMD=(python3)
fi

# Optional: attempt to use Codon when requested, but fall back for packages Codon doesn't support
# Enable with: CODON=1 ./run_examples.sh
function can_run_with_codon {
# Return 0 (true) if the example likely works with Codon; 1 otherwise
local file="$1"
# Heuristic: Codon generally doesn't support heavy third-party modules
if grep -E '^(from|import)\s+(pandas|numpy|sklearn|seaborn|z3|pathos|matplotlib|sympy)\b' "$file" >/dev/null 2>&1; then
return 1
fi
return 0
}

set -o errexit
set -o nounset
Expand All @@ -14,7 +31,13 @@ cd "$(dirname "$0")"

function run_example {
printf "Running $1..."
$PYTHON_BINARY $1 > /dev/null
if [[ "${CODON-0}" == "1" ]] && command -v codon >/dev/null 2>&1 && can_run_with_codon "$1"; then
codon run -release "$1" > /dev/null || { echo "(failed)"; exit 111; }
echo "(success)"
return
fi

"${PYTHON_CMD[@]}" "$1" > /dev/null
RESULT=$?
if [ $RESULT -eq 0 ]; then
echo "(success)"
Expand Down
Loading