Browse all skills & agents
AI & ML
LLM evaluation, ML model testing, AI-assisted test generation, data notebooks, search relevance.
4 plugins · 21 skills · 5 agents
qa-ai-assisted
v1.4.0
AI-assisted test generation + curation: 3 skills (ai-spec-coverage-mapper, ai-test-generator, model-based-test-graph-author) and 3 agents (ai-test-curator, ai-test-shallow-coverage-critic, mbt-suite-builder).
3 skills + 2 agents
qa-llm-evaluation
v1.4.0
LLM and prompt evaluation: 7 skills (deepeval-evaluation, giskard-llm, langfuse-tracing, llm-regression-suite-author, openai-evals, promptfoo-evaluation, ragas-evaluation) and 2 agents (llm-red-team-planner, prompt-eval-reviewer). Covers the mainstream OSS LLM-eval ecosystem: Promptfoo + OpenAI Evals + DeepEval + Ragas for functional eval, Giskard for adversarial scan, Langfuse for production observability.
7 skills + 2 agents
qa-ml-models
v1.4.0
ML model testing: 7 skills (giskard-tests, deepchecks-tests, evidently-monitoring, fairlearn-fairness, model-performance-regression-gate, model-risk-evidence-matrix, notebook-ci-pipeline-author). Covers tabular + NLP + vision validation, drift monitoring with alert triage, group fairness with a risk-tiered promotion gate, per-prediction explainability, and the Jupyter notebook CI pipeline (papermill + nbval + testbook + nbstripout).
7 skills
qa-search-relevance
v1.3.0
Search relevance testing: 6 skills (elasticsearch-relevance-tests, hybrid-search-eval-author, judgment-list-author, opensearch-relevance-tests, solr-relevance-tests, vector-search-recall-tests) and 1 agent (relevance-regression-reviewer). IR-metrics-driven NDCG / MRR / Recall@k regression detection.
4 skills + 1 agent