Testland
Browse all skills & agents

fairlearn-fairness

Compute group fairness metrics (selection rate, demographic parity, equalized odds) per sensitive feature with `MetricFrame`, then mitigate disparities using Reductions algorithms (`ExponentiatedGradient` with constraint = `DemographicParity`/`EqualizedOdds`). Wire group-disaggregated assertions into the model-evaluation gate. Use when a model's decisions affect people and a stakeholder, auditor, or regulation (ECOA, GDPR Art. 22, EU AI Act high-risk) requires evidence of per-group outcomes, or when someone reports the model treats a specific group worse.

Install with skills.sh (any agent)

npx skills add testland/qa --skill fairlearn-fairness
View source

fairlearn-fairness

Fairlearn provides "Metrics - Tools to assess which groups are negatively impacted and compare models across fairness and accuracy dimensions" and "Algorithms - Techniques to mitigate unfairness" per the Fairlearn quickstart (opens in new window). Two primitives: MetricFrame (group disaggregation) + Reductions (ExponentiatedGradient, ThresholdOptimizer).

When to use

  • Pre-deployment: assert per-group accuracy / selection rate disparities are within budget.
  • Bias incident triage: a stakeholder reports the model is unfair to group X; produce evidence + a mitigated comparison.
  • Compliance evidence (ECOA, GDPR Art. 22, EU AI Act high-risk systems): group-disaggregated metrics + mitigation provenance.

Step 1 - Install

pip install fairlearn
# OR
conda install -c conda-forge fairlearn

Per the Fairlearn quickstart (opens in new window).

Step 2 - Compute disaggregated accuracy

from fairlearn.metrics import MetricFrame
from sklearn.metrics import accuracy_score
from sklearn.tree import DecisionTreeClassifier

classifier = DecisionTreeClassifier(min_samples_leaf=10, max_depth=4)
classifier.fit(X, y_true)
y_pred = classifier.predict(X)

mf = MetricFrame(
    metrics=accuracy_score,
    y_true=y_true,
    y_pred=y_pred,
    sensitive_features=sex,
)
print(mf.by_group)
print(f"Disparity (max-min): {mf.difference()}")

Per the Fairlearn quickstart (opens in new window). sensitive_features can be a Series or a 2-D array for intersectional analysis (sex × race).

Step 3 - Compute selection-rate disparity

from fairlearn.metrics import selection_rate

sr = MetricFrame(
    metrics=selection_rate,
    y_true=y_true,
    y_pred=y_pred,
    sensitive_features=sex,
)
print(sr.by_group)
# Demographic Parity Difference (DPD)
print(f"DPD: {sr.difference()}")

DPD = max group selection rate − min group selection rate. Industry guidance often cites the 80% rule (selection rate ratio ≥ 0.8 between groups) as a soft threshold; consult legal counsel for binding thresholds in your jurisdiction.

Step 4 - Equalized odds (TPR + FPR per group)

from fairlearn.metrics import (
    true_positive_rate,
    false_positive_rate,
    MetricFrame,
)

mf = MetricFrame(
    metrics={
        "TPR": true_positive_rate,
        "FPR": false_positive_rate,
        "selection_rate": selection_rate,
    },
    y_true=y_true,
    y_pred=y_pred,
    sensitive_features=sex,
)
print(mf.by_group)

Equalized Odds requires both TPR and FPR to be equal across groups - stricter than Demographic Parity.

Step 5 - Mitigation via Reductions

from fairlearn.reductions import DemographicParity, ExponentiatedGradient

constraint = DemographicParity()
mitigator = ExponentiatedGradient(classifier, constraint)
mitigator.fit(X, y_true, sensitive_features=sex)
y_pred_mitigated = mitigator.predict(X)

Per the Fairlearn quickstart (opens in new window): this approach significantly reduces selection-rate differences while maintaining accuracy. Other constraints: EqualizedOdds, TruePositiveRateParity, FalsePositiveRateParity.

Step 6 - Threshold post-processing

from fairlearn.postprocessing import ThresholdOptimizer

postprocess = ThresholdOptimizer(
    estimator=classifier,
    constraints="demographic_parity",
    prefit=True,
)
postprocess.fit(X, y_true, sensitive_features=sex)
y_pred_pp = postprocess.predict(X, sensitive_features=sex)

Per the Fairlearn postprocessing (opens in new window) guide, Fairlearn currently supports one postprocessing technique, ThresholdOptimizer. Cheaper than retraining; trades model output for per-group threshold adjustment.

Step 7 - CI assertion

def assert_fairness(y_true, y_pred, sensitive, max_dpd=0.10):
    sr = MetricFrame(
        metrics=selection_rate,
        y_true=y_true,
        y_pred=y_pred,
        sensitive_features=sensitive,
    )
    dpd = sr.difference()
    if dpd > max_dpd:
        raise AssertionError(
            f"Demographic Parity Difference {dpd:.3f} exceeds budget {max_dpd}"
        )

assert_fairness(y_true, y_pred, sex, max_dpd=0.10)

Anti-patterns

Anti-patternWhy it failsFix
Compute aggregate accuracy onlyHides group disparitiesAlways use MetricFrame (Step 2)
Choose Demographic Parity for all problemsDP can be inappropriate when base rates legitimately differ across groupsMatch constraint to legal/ethical context: DP, EO, EOD, EOP
Mitigate via training data resampling aloneDoesn't generalize to new data; brittleUse Reductions (Step 5) or post-processing (Step 6)
Single sensitive attribute (e.g., sex only)Misses intersectional disparities (Black women)Pass 2-D sensitive_features for intersection (Step 2)
Hard-code 80% rule globallyNot legally binding everywhere; not appropriate for all metricsTune max_dpd per use case + legal counsel; use waiver template if scope-exclusion needed

Limitations

  • Fairlearn does not detect proxy discrimination (zip code as proxy for race). Pair with feature-correlation analysis and domain expertise.
  • Mitigation often comes at accuracy cost; document the tradeoff with stakeholders, not unilaterally.

References

Related skills

alibi-explainability

Generates model explanations with Alibi Explain - Anchors, Integrated Gradients, Kernel/Tree SHAP, ALE, Counterfactual Instances. Wires explainer.fit + explainer.explain into model-evaluation pipelines so that every flagged prediction ships with a "why" record auditors can reason about. Use when a model decision must be explainable to an auditor, regulator, or affected user, or when a support team cannot answer why a specific prediction was made.

deepchecks-tests

Run Deepchecks suites (data integrity, train-test validation, model evaluation) on tabular / NLP / vision data + models. Pass `result.passed_conditions()` to CI to gate on regressions; the same checks run during research, CI, and production monitoring per the Deepchecks lifecycle posture. Use before training to catch train-test leakage and data-integrity defects in a tabular, NLP, or vision dataset, and to re-run the same suite on production samples to detect drift.

evidently-monitoring

Use Evidently OSS (100+ evaluation metrics, declarative testing API) to detect data drift, target drift, and model-performance regression, wired into CI as a gate (a Report run with include_tests) and into production monitoring as a continuous check; reports as HTML + JSON for both human review and pipeline assertions. Use when you need a drift or quality gate, or a scheduled monitoring job, for a tabular ML model. Built on the Evidently API specifically: for DeepChecks-based validation suites use deepchecks-tests instead.

giskard-tests

Test ML models with Giskard's scan() vulnerability detector + test catalog (performance, robustness, fairness, data leakage, ethical issues) for tabular and NLP models. Wrap a prediction function in giskard.Model + a DataFrame in giskard.Dataset; emit test suites that pass/fail in CI. Use when a trained tabular or NLP model is about to ship with no test suite of its own, or when a feature-engineering or hyperparameter change needs a pre-merge scan for newly introduced vulnerabilities.

model-performance-regression-gate

Computes held-out metrics (accuracy, F1, AUC, RMSE) for a retrained model and compares them against the current production model, failing promotion when any metric regresses beyond a configured tolerance. Adds per-segment checks via Deepchecks WeakSegmentsPerformance so a model that improves globally but regresses on a key slice is still blocked. Use when a retrained model is a candidate for promotion and the CI pipeline must enforce a per-metric pass/fail gate before the artifact is pushed to the model registry.

model-risk-evidence-matrix

Assigns a machine learning model to a low, medium, or high risk tier from what its predictions decide about people, then derives the fairness and explainability evidence that tier must produce: group metrics per declared sensitive feature, intersectional breakdowns with per-cell counts, vulnerability scan categories, a drift monitoring plan, and per-prediction explanation logs. Supplies conventional demographic parity difference bands, a per-vulnerability-category blocking table, and evidence rules that mark a bundle incomplete or self-contradicting. Use when a model release candidate is up for promotion and someone must decide which fairness artifacts are mandatory rather than nice to have, or when a model card declares a risk tier and the attached evidence bundle has to be checked against what that tier demands.