Testland
Browse all skills & agents

AI & ML

LLM evaluation, ML model testing, AI-assisted test generation, data notebooks, search relevance.

4 plugins · 21 skills · 5 agents

qa-ai-assisted

v1.4.0

AI-assisted test generation + curation: 3 skills (ai-spec-coverage-mapper, ai-test-generator, model-based-test-graph-author) and 3 agents (ai-test-curator, ai-test-shallow-coverage-critic, mbt-suite-builder).

3 skills + 2 agents

qa-llm-evaluation

v1.4.0

LLM and prompt evaluation: 7 skills (deepeval-evaluation, giskard-llm, langfuse-tracing, llm-regression-suite-author, openai-evals, promptfoo-evaluation, ragas-evaluation) and 2 agents (llm-red-team-planner, prompt-eval-reviewer). Covers the mainstream OSS LLM-eval ecosystem: Promptfoo + OpenAI Evals + DeepEval + Ragas for functional eval, Giskard for adversarial scan, Langfuse for production observability.

7 skills + 2 agents

qa-ml-models

v1.4.0

ML model testing: 7 skills (giskard-tests, deepchecks-tests, evidently-monitoring, fairlearn-fairness, model-performance-regression-gate, model-risk-evidence-matrix, notebook-ci-pipeline-author). Covers tabular + NLP + vision validation, drift monitoring with alert triage, group fairness with a risk-tiered promotion gate, per-prediction explainability, and the Jupyter notebook CI pipeline (papermill + nbval + testbook + nbstripout).

7 skills

qa-search-relevance

v1.3.0

Search relevance testing: 6 skills (elasticsearch-relevance-tests, hybrid-search-eval-author, judgment-list-author, opensearch-relevance-tests, solr-relevance-tests, vector-search-recall-tests) and 1 agent (relevance-regression-reviewer). IR-metrics-driven NDCG / MRR / Recall@k regression detection.

4 skills + 1 agent