Skip to content

Update sklearn models for feature probabilities - #365

Merged
alcides merged 7 commits into
mainfrom
2025-11-04-qbdl-5RYRJ
Nov 5, 2025
Merged

alcides merged 7 commits into
mainfrom
2025-11-04-qbdl-5RYRJ

Conversation

@alcides

@alcides alcides commented Nov 4, 2025 •

Copy link
Copy Markdown
Owner

Note

Adds optional correlation-weighted feature sampling across sklearn estimators, improves robustness of weighted choice, updates docs, and enhances the examples runner with uv and optional Codon.

  • Estimators (sklearn-compatible):
    • Add weight_features_by_correlation option in base GeneticEngineEstimator; compute feature probabilities via correlation_weights.
    • Regressors: use make_var(..., weights=...) to bias terminal sampling; keep relative weight 10.
    • Classifiers: apply VarRangeWithProbabilities with computed weights and weight(10); update HillClimbingClassifier init and _parameter_constraints.
  • Grammar (symbolic regression):
    • make_var now supports optional probability weights via VarRangeWithProbabilities, retains backward compatibility with numeric relative_weight.
  • Random:
    • choice_weighted sanitizes non-finite/non-positive weights and falls back to uniform when needed.
  • Examples:
    • Classifier example uses f1_score(..., average="weighted").
    • run_examples.sh: run with uv by default; optional Codon execution with heuristic fallback.
  • Docs:
    • Add section on running examples with uv/Codon.
    • Document feature probability weighting for sklearn estimators.

Written by Cursor Bugbot for commit 82473f0. This will update automatically on new commits. Configure here.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR is being reviewed by Cursor Bugbot

Details

You are on the Bugbot Free tier. On this plan, Bugbot will review limited PRs each billing cycle.

To receive Bugbot reviews on all of your PRs, visit the Cursor dashboard to activate Pro and start your 14-day free trial.

Comment thread geml/common.py
def __init__(self, max_time: float | int = 1, seed: int = 0, weight_features_by_correlation: bool = False):
self.max_time = max_time
self.seed = 0
self.weight_features_by_correlation = weight_features_by_correlation

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Seed Parameter Not Applied: Always Zeroed Seed

The seed parameter is hardcoded to 0 instead of using the seed parameter passed to init. Line 86 should be self.seed = seed instead of self.seed = 0. This bug prevents users from setting a custom random seed, causing all instances to use seed=0 regardless of what value is passed to the constructor.

Fix in Cursor Fix in Web

@alcides
alcides force-pushed the 2025-11-04-qbdl-5RYRJ branch from 716e930 to 2ac3c6b Compare November 5, 2025 08:56
if total <= 0:
return random.choice(self.options)
normalized = [w / total if w >= 0 else 0.0 for w in self.weights]
return random.choice_weighted(self.options, normalized)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Incorrect normalization with negative weights

Incorrect normalization when weights contain both negative and non-negative values. The code computes total = sum(self.weights) including negative weights, then normalizes by replacing negative weights with 0.0 and dividing non-negative weights by total. This results in normalized weights that don't sum to 1.0. For example, with weights [2.0, -1.0, 2.0], total=3.0, but normalized=[0.667, 0.0, 0.667] sums to 1.334. The correct approach is to filter negative weights first, compute the sum of remaining positive weights, then normalize.

Fix in Cursor Fix in Web

Comment thread geml/common.py Outdated
def wrapper(v:float) -> float:
return 1 - abs(float(np.average(v))) + 0.00001

return [wrapper(np.corrcoef(data[:, i], target)) for i in range(len(feature_names))]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Incorrect correlation weight computation using full matrix

The correlation_weights method incorrectly uses np.corrcoef(data[:, i], target) which returns a 2x2 correlation matrix, not a scalar value. The wrapper function then calls np.average() on this matrix, computing the average of all 4 elements (which includes two 1.0 values on the diagonal and two correlation coefficients), resulting in an incorrect weight calculation. The correct approach would be to extract the correlation coefficient using indexing: np.corrcoef(data[:, i], target)[0, 1].

Fix in Cursor Fix in Web

Comment thread geml/classifiers.py
Var.feature_names = feature_names # type:ignore
index_of = {n: i for i, n in enumerate(feature_names)}
Var.to_numpy = lambda s: f"dataset[:,{index_of[s.name]}]" # type:ignore
Var = weight(10)(Var)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Conditional weighting bypasses flag in get_grammar

The get_grammar method unconditionally calls self.correlation_weights(feature_names, data, target) and applies the weights, ignoring the weight_features_by_correlation parameter. The correlation-based weighting should only be applied when self.weight_features_by_correlation is True, otherwise uniform weights should be used.

Fix in Cursor Fix in Web

Comment thread geml/regressors.py
index_of = {n: i for i, n in enumerate(feature_names)}
Var.to_numpy = lambda s: f"dataset[:,{index_of[s.name]}]"
weights = self.correlation_weights(feature_names, data, target)
Var = make_var(feature_names, weights=weights, relative_weight=10)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Correlation weighting ignored when disabled.

The get_grammar method unconditionally calls self.correlation_weights(feature_names, data, target) and applies the weights, ignoring the weight_features_by_correlation parameter. The correlation-based weighting should only be applied when self.weight_features_by_correlation is True, otherwise uniform weights should be used.

Fix in Cursor Fix in Web

@alcides
alcides force-pushed the 2025-11-04-qbdl-5RYRJ branch from 1bbdf8e to b3e5e92 Compare November 5, 2025 10:00
Comment thread geml/common.py

def wrapper(corr_value: float) -> float:
# Higher absolute correlation -> smaller weight (bias search), add epsilon
return 1 - abs(corr_value) + 0.00001

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: Inverted Correlation Weighting Misaligns Sampling Likelihood

The correlation weighting logic is inverted. The documentation states features should be "sampled with probabilities proportional to their absolute Pearson correlation", but the implementation returns 1 - abs(corr_value), which gives LOWER weights to features with HIGHER correlation. For a feature with correlation 0.9, the weight becomes 0.1, while a feature with correlation 0.1 gets weight 0.9. This is inversely proportional to correlation, contradicting the intended behavior. The formula should be abs(corr_value) + 0.00001 instead.

Fix in Cursor Fix in Web

@alcides
alcides merged commit 4b1c92c into main Nov 5, 2025
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant