{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "0",
   "metadata": {},
   "source": [
    "# Scorer Metrics\n",
    "\n",
    "This demo will walk you through how to measure and gauge performance of PyRIT scoring configurations, both harm scorers and objective scorers.\n",
    "\n",
    "\n",
    "## Understanding Scorer Metrics\n",
    "\n",
    "PyRIT has metrics for many scorers checked in. Before diving into how to retrieve or create metrics, it's important to understand what they measure and how scorer identity determines which metrics apply.\n",
    "\n",
    "### Scorer Identity\n",
    "\n",
    "Every scorer has a unique **identity hash** computed from its complete configuration:\n",
    "- Scorer type (e.g., `SelfAskRefusalScorer`)\n",
    "- System and user prompt templates\n",
    "- Target model information (model name)\n",
    "- Temperature and other generation parameters\n",
    "- Any scorer-specific configuration\n",
    "\n",
    "This means changing *any* of these values creates a new scorer identity. The reason these are variables is because they _might_ change performance—does changing the temperature increase or decrease accuracy? Metrics let you experiment and find out.\n",
    "\n",
    "Metrics are stored and retrieved by this identity hash, so the same scorer configuration will always get the same cached metrics."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Found default environment files: ['./.pyrit/.env', './.pyrit/.env.local']\n",
      "Loaded environment file: ./.pyrit/.env\n",
      "Loaded environment file: ./.pyrit/.env.local\n"
     ]
    },
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[pyrit:alembic] No new upgrade operations detected.\n"
     ]
    },
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Scorer Identity:\n",
      "  Type: SelfAskRefusalScorer\n",
      "  System Prompt: **Important**: Your role is strictly that of an ev...\n",
      "  Target Info: OpenAIChatTarget::99d992bd\n",
      "  Identity Hash: 12cd53923d4290d9c97740cc75242610fbcfce508851d79116d09994b5ce9e6e\n"
     ]
    }
   ],
   "source": [
    "from pyrit.prompt_target import OpenAIChatTarget\n",
    "from pyrit.score import SelfAskRefusalScorer\n",
    "from pyrit.setup import IN_MEMORY, initialize_pyrit_async\n",
    "\n",
    "await initialize_pyrit_async(memory_db_type=IN_MEMORY)  # type: ignore\n",
    "\n",
    "# Create a refusal scorer\n",
    "refusal_scorer = SelfAskRefusalScorer(chat_target=OpenAIChatTarget(temperature=0.9))\n",
    "\n",
    "# View the scorer's full identity - this determines which metrics apply\n",
    "scorer_identity = refusal_scorer.get_identifier()\n",
    "print(\"Scorer Identity:\")\n",
    "print(f\"  Type: {scorer_identity.class_name}\")\n",
    "system_prompt = scorer_identity.params.get(\"system_prompt_template\")\n",
    "print(f\"  System Prompt: {system_prompt[:50] if system_prompt else 'None'}...\")\n",
    "prompt_target = scorer_identity.get_child(\"prompt_target\")\n",
    "print(f\"  Target Info: {prompt_target}\")\n",
    "print(f\"  Identity Hash: {scorer_identity.hash}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "2",
   "metadata": {},
   "source": [
    "### Objective Metrics\n",
    "\n",
    "Objective scorers produce true/false outputs, usually whether an objective was met (e.g., did this response have instructions on \"how to make a Molotov cocktail?\"). We evaluate them using standard classification metrics by comparing model predictions against human-labeled ground truth.\n",
    "\n",
    "- **Accuracy**: Proportion of predictions matching human labels. Simple but can be misleading with imbalanced datasets.\n",
    "- **Precision**: Of all \"true\" predictions, how many were correct? High precision = few false positives.\n",
    "- **Recall**: Of all actual \"true\" cases, how many did we catch? High recall = few false negatives.\n",
    "- **F1 Score**: Harmonic mean of precision and recall. Balances both concerns.\n",
    "- **Accuracy Standard Error**: Statistical uncertainty in accuracy estimate, useful for confidence intervals.\n",
    "\n",
    "**Which metric matters most?**\n",
    "- If false positives are costly (e.g., flagging safe content as harmful) → prioritize **precision**\n",
    "- If false negatives are costly (e.g., missing actual jailbreaks) → prioritize **recall**\n",
    "- For balanced scenarios → use **F1 score**"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "3",
   "metadata": {},
   "source": [
    "### Harm Metrics\n",
    "\n",
    "Harm scorers produce float scores (0.0-1.0) representing severity. Since these are continuous values, we use different metrics that capture how close the model's scores are to human judgments.\n",
    "\n",
    "**Error Metrics:**\n",
    "- **Mean Absolute Error (MAE)**: Average absolute difference between model and human scores. An MAE of 0.15 means the model is off by 0.15 on average.\n",
    "- **MAE Standard Error**: Uncertainty in the MAE estimate.\n",
    "\n",
    "**Statistical Significance:**\n",
    "- **t-statistic**: From a one-sample t-test. Positive = model scores higher than humans; negative = lower.\n",
    "- **p-value**: If small (e.g., < 0.05), the difference between model and human scores is statistically significant (not due to chance).\n",
    "\n",
    "**Inter-Rater Reliability (Krippendorff's Alpha):**\n",
    "Measures agreement between evaluators, ranging from -1.0 to 1.0:\n",
    "- **1.0**: Perfect agreement\n",
    "- **0.8+**: Strong agreement\n",
    "- **0.6-0.8**: Moderate agreement\n",
    "- **< 0.6**: Weak agreement\n",
    "\n",
    "Three alpha values are reported:\n",
    "- **`krippendorff_alpha_humans`**: Agreement among human evaluators (baseline quality of labels)\n",
    "- **`krippendorff_alpha_model`**: Agreement across multiple model scoring trials (model consistency)\n",
    "- **`krippendorff_alpha_combined`**: Overall agreement between humans and model"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4",
   "metadata": {},
   "source": [
    "## Retrieving Scorer Metrics\n",
    "\n",
    "When scorer metrics are calculated with `evaluate_async()`, they can be saved to JSONL registry files and retrieved without re-running the evaluation. The PyRIT team has pre-computed metrics for common scorer configurations which you can access immediately.\n",
    "\n",
    "### Retrieving Cached Metrics for a Scorer\n",
    "\n",
    "Use `get_scorer_metrics()` on any scorer instance to retrieve cached results matching its identity:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "5",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "No cached metrics found for this scorer configuration.\n"
     ]
    }
   ],
   "source": [
    "import os\n",
    "\n",
    "from pyrit.auth import get_azure_openai_auth\n",
    "from pyrit.prompt_target import OpenAIChatTarget\n",
    "from pyrit.score import (\n",
    "    SelfAskRefusalScorer,\n",
    "    TrueFalseInverterScorer,\n",
    ")\n",
    "\n",
    "# This is a simple objective scorer that only detects whether the response was a refusal\n",
    "gpt4o_endpoint = os.environ.get(\"AZURE_OPENAI_GPT4O_ENDPOINT\")\n",
    "objective_scorer = TrueFalseInverterScorer(\n",
    "    scorer=SelfAskRefusalScorer(\n",
    "        chat_target=OpenAIChatTarget(\n",
    "            endpoint=gpt4o_endpoint,\n",
    "            api_key=get_azure_openai_auth(gpt4o_endpoint),\n",
    "            model_name=os.environ.get(\"AZURE_OPENAI_GPT4O_MODEL\"),\n",
    "        )\n",
    "    )\n",
    ")\n",
    "\n",
    "# Retrieve pre-computed metrics (from PyRIT team's evaluation runs)\n",
    "# using the scorer's identity hash\n",
    "cached_metrics = objective_scorer.get_scorer_metrics()\n",
    "\n",
    "if cached_metrics:\n",
    "    print(\"Evaluation Metrics:\")\n",
    "    print(f\"  Dataset Name: {cached_metrics.dataset_name}\")\n",
    "    print(f\"  Dataset Version: {cached_metrics.dataset_version}\")\n",
    "    print(f\"  F1 Score: {cached_metrics.f1_score}\")\n",
    "    print(f\"  Accuracy: {cached_metrics.accuracy}\")\n",
    "    print(f\"  Precision: {cached_metrics.precision}\")\n",
    "    print(f\"  Recall: {cached_metrics.recall}\")\n",
    "    print(f\"  Avg Score time: {cached_metrics.average_score_time_seconds} seconds\")\n",
    "else:\n",
    "    print(\"No cached metrics found for this scorer configuration.\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6",
   "metadata": {},
   "source": [
    "Harm scorer metrics are retrieved similarly:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "7",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "No cached metrics found for this scorer configuration.\n"
     ]
    }
   ],
   "source": [
    "import os\n",
    "\n",
    "from pyrit.auth import get_azure_openai_auth\n",
    "from pyrit.prompt_target import OpenAIChatTarget\n",
    "from pyrit.score import LikertScalePaths, SelfAskLikertScorer\n",
    "\n",
    "gpt4o_endpoint = os.environ.get(\"AZURE_OPENAI_GPT4O_ENDPOINT\")\n",
    "harm_scorer = SelfAskLikertScorer.from_likert_scale(\n",
    "    chat_target=OpenAIChatTarget(\n",
    "        endpoint=gpt4o_endpoint,\n",
    "        api_key=get_azure_openai_auth(gpt4o_endpoint),\n",
    "        model_name=os.environ.get(\"AZURE_OPENAI_GPT4O_MODEL\"),\n",
    "    ),\n",
    "    likert_scale=LikertScalePaths.EXPLOITS_SCALE.load(),\n",
    ")\n",
    "\n",
    "# Retrieve pre-computed metrics using the scorer's identity hash\n",
    "harm_metrics = harm_scorer.get_scorer_metrics()\n",
    "\n",
    "if harm_metrics:\n",
    "    print(\"Evaluation Metrics:\")\n",
    "    print(f\"  Dataset Name: {harm_metrics.dataset_name}\")\n",
    "    print(f\"  Dataset Version: {harm_metrics.dataset_version}\")\n",
    "    print(f\"  Mean Absolute Error: {harm_metrics.mean_absolute_error:.3f}\")\n",
    "    print(f\"  Krippendorff Alpha: {harm_metrics.krippendorff_alpha_combined:.3f}\")\n",
    "    print(f\"  P value: {harm_metrics.p_value:.4f}\")\n",
    "    print(f\"  Avg Score time: {harm_metrics.average_score_time_seconds} seconds\")\n",
    "\n",
    "else:\n",
    "    print(\"No cached metrics found for this scorer configuration.\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "8",
   "metadata": {},
   "source": [
    "### Comparing All Scorer Configurations\n",
    "\n",
    "The evaluation registry stores metrics for all tested scorer configurations. You can load all entries to compare which configurations perform best.\n",
    "\n",
    "Use `get_all_objective_metrics()` or `get_all_harm_metrics()` to load evaluation results. These return `ScorerMetricsWithIdentity` objects with clean attribute access to both the scorer identity and its metrics."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "9",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Found 24 scorer configurations in the metrics file\n",
      "\n",
      "Top 5 configurations by F1 Score:\n",
      "--------------------------------------------------------------------------------\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: TrueFalseInverterScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m      • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m        └─ Composite of 1 scorer(s):\u001b[0m\n",
      "\u001b[36m            • Scorer Type: SelfAskRefusalScorer\u001b[0m\n",
      "\u001b[36m            • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m            • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m            • model_name: gpt-4o-japan-nilfilter\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[36m      • Accuracy: 89.37%\u001b[0m\n",
      "\u001b[36m      • Accuracy Std Error: ±0.0155\u001b[0m\n",
      "\u001b[36m      • F1 Score: 0.8918\u001b[0m\n",
      "\u001b[36m      • Precision: 0.8782\u001b[0m\n",
      "\u001b[32m      • Recall: 0.9058\u001b[0m\n",
      "\u001b[36m      • Average Score Time: 0.59s\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: TrueFalseInverterScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m      • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m        └─ Composite of 1 scorer(s):\u001b[0m\n",
      "\u001b[36m            • Scorer Type: SelfAskRefusalScorer\u001b[0m\n",
      "\u001b[36m            • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m            • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m            • model_name: gpt-5.4\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[33m      Official evaluation has not been run yet for this specific configuration\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: TrueFalseInverterScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m      • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m        └─ Composite of 1 scorer(s):\u001b[0m\n",
      "\u001b[36m            • Scorer Type: SelfAskRefusalScorer\u001b[0m\n",
      "\u001b[36m            • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m            • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m            • model_name: gpt-4o\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[33m      Official evaluation has not been run yet for this specific configuration\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: TrueFalseInverterScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m      • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m        └─ Composite of 1 scorer(s):\u001b[0m\n",
      "\u001b[36m            • Scorer Type: SelfAskRefusalScorer\u001b[0m\n",
      "\u001b[36m            • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m            • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m            • model_name: gpt-4o-japan-nilfilter\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[36m      • Accuracy: 84.56%\u001b[0m\n",
      "\u001b[36m      • Accuracy Std Error: ±0.0182\u001b[0m\n",
      "\u001b[36m      • F1 Score: 0.8530\u001b[0m\n",
      "\u001b[36m      • Precision: 0.7902\u001b[0m\n",
      "\u001b[32m      • Recall: 0.9267\u001b[0m\n",
      "\u001b[36m      • Average Score Time: 0.61s\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: TrueFalseInverterScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m      • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m        └─ Composite of 1 scorer(s):\u001b[0m\n",
      "\u001b[36m            • Scorer Type: SelfAskRefusalScorer\u001b[0m\n",
      "\u001b[36m            • scorer_type: true_false\u001b[0m\n",
      "\u001b[36m            • score_aggregator: OR_\u001b[0m\n",
      "\u001b[36m            • model_name: gpt-4o-unsafe\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[33m      Official evaluation has not been run yet for this specific configuration\u001b[0m\n",
      "\n",
      "================================================================================\n",
      "Best Accuracy:  89.37%\n",
      "Best Precision: 0.990\n",
      "Best Recall:    0.989\n",
      "Fastest:        0.043 seconds\n",
      "Slowest:        10.397 seconds\n"
     ]
    }
   ],
   "source": [
    "from pyrit.output import output_scorer_async\n",
    "from pyrit.score import get_all_objective_metrics\n",
    "\n",
    "# Load all objective scorer metrics - returns ScorerMetricsWithIdentity[ObjectiveScorerMetrics]\n",
    "all_scorers = get_all_objective_metrics()\n",
    "\n",
    "print(f\"Found {len(all_scorers)} scorer configurations in the metrics file\\n\")\n",
    "\n",
    "# Sort by F1 score - type checker knows entry.metrics is ObjectiveScorerMetrics\n",
    "sorted_by_f1 = sorted(all_scorers, key=lambda x: x.metrics.f1_score, reverse=True)\n",
    "\n",
    "print(\"Top 5 configurations by F1 Score:\")\n",
    "print(\"-\" * 80)\n",
    "for _i, entry in enumerate(sorted_by_f1[:5], 1):\n",
    "    await output_scorer_async(scorer_identifier=entry.scorer_identifier)\n",
    "\n",
    "# Show best by each metric\n",
    "print(\"\\n\" + \"=\" * 80)\n",
    "print(f\"Best Accuracy:  {max(all_scorers, key=lambda x: x.metrics.accuracy).metrics.accuracy:.2%}\")\n",
    "print(f\"Best Precision: {max(all_scorers, key=lambda x: x.metrics.precision).metrics.precision:.3f}\")\n",
    "print(f\"Best Recall:    {max(all_scorers, key=lambda x: x.metrics.recall).metrics.recall:.3f}\")\n",
    "print(\n",
    "    f\"Fastest:        {min(all_scorers, key=lambda x: x.metrics.average_score_time_seconds).metrics.average_score_time_seconds:.3f} seconds\"\n",
    ")\n",
    "print(\n",
    "    f\"Slowest:        {max(all_scorers, key=lambda x: x.metrics.average_score_time_seconds).metrics.average_score_time_seconds:.3f} seconds\"\n",
    ")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "10",
   "metadata": {},
   "source": [
    "Similarly, you can look at the best harm scorers for a given category:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "11",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Found 4 harm scorer configurations for violence\n",
      "\n",
      "Top configurations by Mean Absolute Error:\n",
      "--------------------------------------------------------------------------------\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: SelfAskLikertScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: float_scale\u001b[0m\n",
      "\u001b[36m      • model_name: gpt-4o\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[36m      • Mean Absolute Error: 0.1901\u001b[0m\n",
      "\u001b[36m      • MAE Std Error: ±0.0229\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Combined): 0.8110\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Model): 0.9126\u001b[0m\n",
      "\u001b[32m      • Average Score Time: 0.39s\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: SelfAskLikertScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: float_scale\u001b[0m\n",
      "\u001b[36m      • model_name: gpt-4o-japan-nilfilter\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[36m      • Mean Absolute Error: 0.1901\u001b[0m\n",
      "\u001b[36m      • MAE Std Error: ±0.0229\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Combined): 0.8110\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Model): 0.9126\u001b[0m\n",
      "\u001b[32m      • Average Score Time: 0.39s\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: SelfAskLikertScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: float_scale\u001b[0m\n",
      "\u001b[36m      • model_name: gpt-4o-japan-nilfilter\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[36m      • Mean Absolute Error: 0.1901\u001b[0m\n",
      "\u001b[36m      • MAE Std Error: ±0.0229\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Combined): 0.8110\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Model): 0.9126\u001b[0m\n",
      "\u001b[32m      • Average Score Time: 0.39s\u001b[0m\n",
      "\n",
      "\u001b[1m  📊 Scorer Information\u001b[0m\n",
      "\u001b[37m    ▸ Scorer Identifier\u001b[0m\n",
      "\u001b[36m      • Scorer Type: AzureContentFilterScorer\u001b[0m\n",
      "\u001b[36m      • scorer_type: float_scale\u001b[0m\n",
      "\n",
      "\u001b[37m    ▸ Performance Metrics\u001b[0m\n",
      "\u001b[36m      • Mean Absolute Error: 0.2437\u001b[0m\n",
      "\u001b[36m      • MAE Std Error: ±0.0238\u001b[0m\n",
      "\u001b[36m      • Krippendorff Alpha (Combined): 0.7754\u001b[0m\n",
      "\u001b[32m      • Krippendorff Alpha (Model): 1.0000\u001b[0m\n",
      "\u001b[32m      • Average Score Time: 0.72s\u001b[0m\n"
     ]
    }
   ],
   "source": [
    "from pyrit.output import output_scorer_async\n",
    "from pyrit.score import get_all_harm_metrics\n",
    "\n",
    "# Load all harm scorer metrics for a specific category\n",
    "all_harm_scorers = get_all_harm_metrics(harm_category=\"violence\")\n",
    "\n",
    "print(f\"Found {len(all_harm_scorers)} harm scorer configurations for violence\\n\")\n",
    "\n",
    "# Sort by mean absolute error (lower is better)\n",
    "sorted_by_mae = sorted(all_harm_scorers, key=lambda x: x.metrics.mean_absolute_error)\n",
    "\n",
    "print(\"Top configurations by Mean Absolute Error:\")\n",
    "print(\"-\" * 80)\n",
    "for _i, e in enumerate(sorted_by_mae[:5], 1):\n",
    "    await output_scorer_async(scorer_identifier=e.scorer_identifier, harm_category=\"violence\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "12",
   "metadata": {},
   "source": [
    "## Creating Scorer Metrics\n",
    "\n",
    "This section covers how to create new metrics by running evaluations against human-labeled datasets.\n",
    "\n",
    "### Caching and Skip Logic\n",
    "\n",
    "When you call `evaluate_async()` on a scorer, the evaluation framework follows a smart caching strategy to avoid redundant work. It checks the metrics registry (a JSONL file) for an existing entry matching the scorer's identity hash. The decision to skip or run evaluation depends on:\n",
    "\n",
    "1. **No existing entry**: Run the full evaluation\n",
    "2. **Dataset version or harm definition version changed**: Re-run and replace the old entry (assumes newer dataset/newer scoring criteria for harm is authoritative)\n",
    "3. **Same version, sufficient trials**: Skip if existing `num_scorer_trials >= requested` (existing metrics are good enough)\n",
    "4. **Same version, fewer trials**: Re-run with more trials and replace (higher fidelity needed)\n",
    "\n",
    "During evaluation, the scorer processes each entry from human-labeled CSV dataset(s). For each `assistant_response` in the CSV, the scorer generates predictions which are compared against the `human_score` column(s). For objective scorers, this produces accuracy/precision/recall/F1 metrics. For harm scorers, it calculates MAE, t-statistics, and Krippendorff's alpha.\n",
    "\n",
    "Setting `add_to_evaluation_results=False` bypasses caching entirely—always running fresh evaluations without reading from or writing to the registry. This is useful for testing custom configurations without polluting the official metrics."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "13",
   "metadata": {},
   "source": [
    "### Running an Objective Evaluation\n",
    "\n",
    "Call `evaluate_async()` on any scorer instance. The scorer's identity (including system prompt, model, temperature) determines which cached results apply."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "14",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      " Accuracy: 1.0\n"
     ]
    }
   ],
   "source": [
    "from typing import cast\n",
    "\n",
    "from pyrit.prompt_target import OpenAIChatTarget\n",
    "from pyrit.score import (\n",
    "    ObjectiveScorerMetrics,\n",
    "    RegistryUpdateBehavior,\n",
    "    ScorerEvalDatasetFiles,\n",
    "    SelfAskRefusalScorer,\n",
    ")\n",
    "\n",
    "# Create a refusal scorer - uses the chat target to determine if responses are refusals\n",
    "refusal_scorer = SelfAskRefusalScorer(chat_target=OpenAIChatTarget())\n",
    "\n",
    "# REAL usage would simply be:\n",
    "# metrics = await refusal_scorer.evaluate_async()\n",
    "\n",
    "# For demonstration, use a smaller evaluation file (normally you'd use the full dataset)\n",
    "# The evaluation_file_mapping tells the evaluator which human-labeled CSV files to use\n",
    "refusal_scorer.evaluation_file_mapping = ScorerEvalDatasetFiles(\n",
    "    human_labeled_datasets_files=[\"sample/mini_refusal.csv\"],\n",
    "    result_file=\"sample/test_refusal_metrics.jsonl\",\n",
    ")\n",
    "\n",
    "# Run evaluation with:\n",
    "# - num_scorer_trials=1: Score each response once (use 3+ for production to measure variance)\n",
    "# - RegistryUpdateBehavior.NEVER_UPDATE: Don't save to the official registry (for testing only)\n",
    "metrics = await refusal_scorer.evaluate_async(  # type: ignore\n",
    "    num_scorer_trials=1, update_registry_behavior=RegistryUpdateBehavior.NEVER_UPDATE\n",
    ")\n",
    "\n",
    "if metrics:\n",
    "    objective_metrics = cast(\"ObjectiveScorerMetrics\", metrics)\n",
    "    print(f\" Accuracy: {objective_metrics.accuracy}\")\n",
    "else:\n",
    "    raise RuntimeError(\"Evaluation failed, no metrics returned\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "15",
   "metadata": {},
   "source": [
    "### Running a Harm Evaluation"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "16",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Metrics for harm category \"exploits\" created\n"
     ]
    }
   ],
   "source": [
    "from typing import cast\n",
    "\n",
    "from pyrit.score import LikertScalePaths, RegistryUpdateBehavior, SelfAskLikertScorer\n",
    "from pyrit.score.scorer_evaluation.scorer_evaluator import ScorerEvalDatasetFiles\n",
    "from pyrit.score.scorer_evaluation.scorer_metrics import HarmScorerMetrics\n",
    "\n",
    "# Create a harm scorer using the hate speech Likert scale\n",
    "likert_scorer = SelfAskLikertScorer.from_likert_scale(\n",
    "    chat_target=OpenAIChatTarget(), likert_scale=LikertScalePaths.EXPLOITS_SCALE.load()\n",
    ")\n",
    "\n",
    "# # Configure evaluation to use a small sample dataset\n",
    "# likert_scorer.evaluation_file_mapping = ScorerEvalDatasetFiles(\n",
    "#     human_labeled_datasets_files=[\"harm/mini_hate_speech.csv\"],\n",
    "#     result_file=\"harm/test_hate_speech_metrics.jsonl\",\n",
    "#     harm_category=\"hate_speech\",  # Required for harm evaluations\n",
    "# )\n",
    "\n",
    "# This can be called without parameters to update the registry\n",
    "metrics = await likert_scorer.evaluate_async(  # type: ignore\n",
    "    num_scorer_trials=1, update_registry_behavior=RegistryUpdateBehavior.NEVER_UPDATE\n",
    ")\n",
    "\n",
    "if metrics:\n",
    "    harm_metrics = cast(\"HarmScorerMetrics\", metrics)\n",
    "    print(f'Metrics for harm category \"{harm_metrics.harm_category}\" created')\n",
    "else:\n",
    "    raise RuntimeError(\"Evaluation failed, no metrics returned\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "17",
   "metadata": {},
   "source": [
    "### Custom Evaluation Datasets\n",
    "\n",
    "You can evaluate scorers against your own human-labeled datasets using `ScorerEvalDatasetFiles`:\n",
    "\n",
    "```python\n",
    "# Configure custom datasets\n",
    "scorer.evaluation_file_mapping = ScorerEvalDatasetFiles(\n",
    "    # Glob patterns relative to SCORER_EVALS_PATH\n",
    "    human_labeled_datasets_files=[\"my_datasets/*.csv\"],\n",
    "    # Where to store results\n",
    "    result_file=\"my_datasets/evaluation_metrics.jsonl\",\n",
    "    # Required for harm scorers, ignored for objective scorers\n",
    "    harm_category=\"violence\",\n",
    ")\n",
    "```\n",
    "\n",
    "### CSV Human Evaluation Files\n",
    "\n",
    "Many human scored dataset csv files are available in the `pyrit/datasets/scorer_evals/` directory. These include datasets for refusal detection, hate speech, violence, and other harm categories. You can reference these as templates for creating your own evaluation datasets.\n",
    "\n",
    "**Required columns:**\n",
    "- `assistant_response`: The text to score\n",
    "- `objective`: Required for **objective scorers** - the scoring question (e.g., \"Did the response refuse?\")\n",
    "- `harm_category`: Required for **harm scorers** - the harm category (e.g., \"hate_speech\", \"violence\")\n",
    "- `human_score` or `human_score_1`, `human_score_2`, etc.: Ground truth labels from human raters\n",
    "  - For objective scorers: 0 or 1 (converted to bool)\n",
    "  - For harm scorers: 0.0-1.0 float values\n",
    "- `data_type`: Type of content (defaults to \"text\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "18",
   "metadata": {},
   "source": [
    "## Batch Evaluation with `evaluate_scorers.py`\n",
    "\n",
    "While `evaluate_async()` runs evaluations for a single scorer, the `evaluate_scorers.py` script\n",
    "evaluates **all registered scorers** in bulk. This is useful for benchmarking after changing scorer\n",
    "prompts, adding new variants, or updating human-labeled datasets.\n",
    "\n",
    "The script initializes PyRIT with `ScorerInitializer` (which registers all configured scorers),\n",
    "then runs `evaluate_async()` on each one. Results are saved to the JSONL registry files in\n",
    "`pyrit/datasets/scorer_evals/`.\n",
    "\n",
    "### Basic Usage\n",
    "\n",
    "```bash\n",
    "# Evaluate all registered scorers (long-running — can take hours)\n",
    "python -m build_scripts.evaluate_scorers\n",
    "\n",
    "# Evaluate only scorers with specific tags\n",
    "python -m build_scripts.evaluate_scorers --tags refusal\n",
    "python -m build_scripts.evaluate_scorers --tags refusal,default\n",
    "\n",
    "# Control parallelism (default: 5, lower if hitting rate limits)\n",
    "python -m build_scripts.evaluate_scorers --max-concurrency 3\n",
    "```\n",
    "\n",
    "### Tags\n",
    "\n",
    "`ScorerInitializer` applies tags to scorers during registration. These tags let you target\n",
    "specific subsets for evaluation:\n",
    "\n",
    "- `refusal` — The 4 standalone refusal scorer variants\n",
    "- `default` — All scorers registered by default\n",
    "- `best_refusal_f1` — The refusal variant with the highest F1 (set dynamically from metrics)\n",
    "- `best_objective_f1` — The objective scorer with the highest F1\n",
    "\n",
    "### Recommended Workflow: Refusal → Dependent Scorers\n",
    "\n",
    "When refusal scorer prompts or datasets change, the recommended workflow is:\n",
    "\n",
    "**Step 1: Evaluate refusal scorers first**\n",
    "\n",
    "```bash\n",
    "python -m build_scripts.evaluate_scorers --tags refusal\n",
    "```\n",
    "\n",
    "This evaluates only the 4 refusal variants and writes results to\n",
    "`refusal_scorer/refusal_metrics.jsonl`. After this step, `ScorerInitializer` can determine which\n",
    "refusal variant has the best F1 and tag it as `best_refusal_f1`.\n",
    "\n",
    "**Step 2: Re-evaluate all scorers**\n",
    "\n",
    "```bash\n",
    "python -m build_scripts.evaluate_scorers\n",
    "```\n",
    "\n",
    "On the next full run, `ScorerInitializer` reads the refusal metrics from Step 1, picks the best\n",
    "refusal variant, and uses it to build dependent scorers (e.g., `TrueFalseInverterScorer` wrapping\n",
    "the best refusal scorer). This ensures objective scorers that depend on refusal detection use the\n",
    "best-performing refusal prompt.\n",
    "\n",
    "Scorers whose metrics are already up-to-date (same dataset version, sufficient trials) are\n",
    "automatically skipped, so re-running the full script is efficient.\n",
    "\n",
    "**Step 3: Commit updated metrics**\n",
    "\n",
    "```bash\n",
    "git add pyrit/datasets/scorer_evals/\n",
    "git commit -m \"chore: update scorer metrics\"\n",
    "```\n",
    "\n",
    "The updated JSONL files should be checked in so that `ScorerInitializer` can read them at runtime\n",
    "to select the best scorers."
   ]
  }
 ],
 "metadata": {
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.14.4"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
