{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "0",
   "metadata": {},
   "source": [
    "# Resiliency and Retry\n",
    "\n",
    "PyRIT provides multiple layers of retry and resiliency mechanisms to handle failures gracefully during security testing. This notebook explains the different retry mechanisms, how they work together, and how to configure them for your use case.\n",
    "\n",
    "## Overview: Different Levels of Retry Mechanisms\n",
    "\n",
    "PyRIT implements retries at multiple levels, each serving a different purpose:\n",
    "\n",
    "1. Low-level Retries (e.g. `pyrit_target_retry`). These automatically attempt to retry certain target errors under known conditions. For example, if there is an HTTP `RateLimitError` on an LLM endpoint, the target can use this attribute to automatically make the HTTP request again with an exponential backoff.\n",
    "2. Mid-level Retries (e.g. `pyrit_json_retry`). These automatically attempt to retry under known conditions. As an exmaple, if a scorer expects an LLM response to have specific JSON, and that is JSON malformed, the scorer can use this attribute to ask the LLM to try scoring again. This does not have an exponential backoff.\n",
    "3. High-level/Scenario-Level Retries: High-level workflow that retries the entire scenario when anything goes wrong, including unknown exceptions. If you are sending thousands of prompts, this can happen! It has logic to pick up where it left off in the scenario execution.\n",
    "\n",
    "Understanding when and how to use each is key to building robust security testing workflows. It's also important to understand these retries are stacked on top of each other. For example, if a target is unresponsive for five requests (a low level retry), a higher level still counts that as a \"success\"."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1",
   "metadata": {},
   "source": [
    "## 1. Low-level Retries: API Call Resilience\n",
    "\n",
    "### What is Target-Level Retry?\n",
    "\n",
    "The `pyrit_target_retry` decorator is a **low-level** retry mechanism that handles transient API failures when communicating with prompt targets (e.g., OpenAI, Azure, custom endpoints).\n",
    "\n",
    "### What It Handles\n",
    "\n",
    "Target-level retry automatically retries when it encounters:\n",
    "\n",
    "- **Rate limit errors** (`RateLimitError`, `RateLimitException`): API rate limits exceeded\n",
    "- **Empty responses** (`EmptyResponseException`): Target returns no content\n",
    "\n",
    "Note if other unknown errors happen, there is no retry. For example, we don't generally want to retry 10 times if there is an auth failure :)\n",
    "\n",
    "### Configuration\n",
    "\n",
    "Target-level retries are configured via environment variables:\n",
    "\n",
    "```python\n",
    "RETRY_MAX_NUM_ATTEMPTS=10      # Maximum retry attempts (default: 10)\n",
    "RETRY_WAIT_MIN_SECONDS=5       # Minimum wait between retries (default: 5)\n",
    "RETRY_WAIT_MAX_SECONDS=220     # Maximum wait between retries (default: 220)\n",
    "```\n",
    "\n",
    "The decorator uses **exponential backoff** - wait time increases exponentially with each retry attempt.\n",
    "\n",
    "### When to Use\n",
    "\n",
    "Target-level retry is **automatically applied** to target implementations. You typically don't need to think about it - it's built into PyRIT's target infrastructure.\n",
    "\n",
    "It's most valuable when:\n",
    "- Working with rate-limited APIs (OpenAI, Azure, etc.)\n",
    "- Network instability causes occasional empty responses\n",
    "- You want automatic handling of transient communication failures"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "2",
   "metadata": {},
   "source": [
    "## 2. Mid-level Retries: Response Format Resilience\n",
    "\n",
    "### What is JSON-Level Retry?\n",
    "\n",
    "The `pyrit_json_retry` decorator handles failures when parsing JSON responses from targets. Due to the probabilistic nature of LLMs, they can occasionally return malformed JSON (or whatever format we're expecting).\n",
    "\n",
    "### What It Handles\n",
    "\n",
    "JSON-level retry automatically retries when:\n",
    "\n",
    "- **Invalid JSON** (`InvalidJsonException`): Target returns non-parseable JSON, a JSON\n",
    "  value other than a top-level object, or (for true/false scorers) a `score_value`\n",
    "  outside the `\"true\"`/`\"false\"` domain (for example, `\"refusal\"`)\n",
    "\n",
    "### Configuration\n",
    "\n",
    "JSON retries use the same environment variables as target retries:\n",
    "\n",
    "```python\n",
    "RETRY_MAX_NUM_ATTEMPTS=10      # Maximum retry attempts\n",
    "```\n",
    "\n",
    "One difference is it does not have an exponential backoff, it is usually retried immediately.\n",
    "\n",
    "### When to Use\n",
    "\n",
    "Like target-level retry, JSON retry is **automatically applied** where needed. It's particularly useful when:\n",
    "- Working with targets that return structured data\n",
    "- LLMs occasionally generate malformed JSON\n",
    "- You need automatic recovery from parsing errors\n",
    "\n",
    "\n",
    "## 3. Scenario-Level Retries: High-Level Workflow Resiliency\n",
    "\n",
    "### What is Scenario-Level Retry?\n",
    "\n",
    "Scenario-level retry is the **highest-level** retry mechanism in PyRIT. When enabled via the `max_retries` parameter, it allows an entire scenario execution to automatically retry if an exception occurs during the workflow.\n",
    "\n",
    "But also note, if you rerun the same scenario manually, that follows the same logic and is always an option.\n",
    "\n",
    "### Key Features\n",
    "\n",
    "- **Picks up where it left off**: On retry, the scenario skips already-completed objectives and continues from the point of exception\n",
    "- **Broad exception handling**: Catches any exception during scenario execution (network issues, target failures, scoring errors, etc.)\n",
    "- **Configurable attempts**: Set `max_retries` to control how many additional attempts are allowed\n",
    "- **Progress tracking**: The `number_tries` field in `ScenarioResult` tracks total attempts\n",
    "\n",
    "### How It Works\n",
    "\n",
    "When you call `scenario.run_async()`, PyRIT:\n",
    "\n",
    "1. **Initial Attempt**: Executes all atomic attacks in sequence\n",
    "2. **On Exception**: If an exception occurs, checks if retries remain\n",
    "3. **Retry with Resume**: On retry, queries memory to identify completed objectives and skips them\n",
    "4. **Continue from Exception Point**: Executes only the remaining objectives\n",
    "5. **Repeat**: Continues retrying until success or `max_retries` exhausted\n",
    "\n",
    "### Example: Basic Scenario with Retries"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Found default environment files: ['./.pyrit/.env', './.pyrit/.env.local']\n",
      "Loaded environment file: ./.pyrit/.env\n",
      "Loaded environment file: ./.pyrit/.env.local\n"
     ]
    },
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "No new upgrade operations detected.\n"
     ]
    },
    {
     "name": "stderr",
     "output_type": "stream",
     "text": [
      "\n"
     ]
    },
    {
     "name": "stderr",
     "output_type": "stream",
     "text": [
      "TargetRegistry entry 'adversarial_chat' not found. Falling back to default OpenAIChatTarget.\n"
     ]
    },
    {
     "name": "stderr",
     "output_type": "stream",
     "text": [
      "TargetRegistry entry 'objective_scorer_chat' not found. Falling back to default OpenAIChatTarget.\n"
     ]
    },
    {
     "name": "stderr",
     "output_type": "stream",
     "text": [
      "Using fallback default objective scorer: TrueFalseInverterScorer with chat target: OpenAIChatTarget\n"
     ]
    },
    {
     "data": {
      "application/vnd.jupyter.widget-view+json": {
       "model_id": "085f263e4205496285a45b08cbed15a2",
       "version_major": 2,
       "version_minor": 0
      },
      "text/plain": [
       "Executing RedTeamAgent:   0%|          | 0/2 [00:00<?, ?attack/s]"
      ]
     },
     "metadata": {},
     "output_type": "display_data"
    },
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Scenario completed after 1 attempt(s)\n",
      "Total results: 2\n"
     ]
    }
   ],
   "source": [
    "from pyrit.prompt_target import OpenAIChatTarget\n",
    "from pyrit.scenario.foundry import FoundryTechnique, RedTeamAgent\n",
    "from pyrit.setup import IN_MEMORY, initialize_pyrit_async\n",
    "from pyrit.setup.initializers import LoadDefaultDatasets, TechniqueInitializer\n",
    "\n",
    "dataset_initializer = LoadDefaultDatasets()\n",
    "dataset_initializer.set_params_from_args(args={\"dataset_names\": [\"harmbench\"]})\n",
    "await initialize_pyrit_async(\n",
    "    memory_db_type=IN_MEMORY,\n",
    "    initializers=[TechniqueInitializer(), dataset_initializer],\n",
    ")  # type: ignore\n",
    "\n",
    "objective_target = OpenAIChatTarget()\n",
    "\n",
    "# Create a scenario with retry configuration\n",
    "scenario = RedTeamAgent()\n",
    "\n",
    "scenario.set_params_from_args(  # type: ignore\n",
    "    args={\n",
    "        \"objective_target\": objective_target,\n",
    "        \"max_concurrency\": 5,\n",
    "        \"max_retries\": 3,\n",
    "        \"scenario_techniques\": [FoundryTechnique.Base64],\n",
    "    }\n",
    ")\n",
    "await scenario.initialize_async()  # type: ignore\n",
    "\n",
    "# Execute with automatic retry after exceptions\n",
    "result = await scenario.run_async()  # type: ignore\n",
    "\n",
    "print(f\"Scenario completed after {result.number_tries} attempt(s)\")\n",
    "print(f\"Total results: {len(result.attack_results)}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4",
   "metadata": {},
   "source": [
    "\n",
    "### Configuring max_retries\n",
    "\n",
    "The `max_retries` parameter controls how many **additional attempts** are allowed after the initial attempt:\n",
    "\n",
    "- `max_retries=0` (default): No retries - fail immediately on first error\n",
    "- `max_retries=1`: 2 total attempts (1 initial + 1 retry)\n",
    "- `max_retries=3`: 4 total attempts (1 initial + 3 retries)\n",
    "\n",
    "**Formula**: `total_attempts = 1 + max_retries`\n",
    "\n",
    "### When to Use Scenario-Level Retries\n",
    "\n",
    "Use scenario-level retries when:\n",
    "\n",
    "- ✅ Running long-duration test campaigns that might encounter transient exceptions\n",
    "- ✅ Testing against unreliable targets or networks\n",
    "- ✅ You want to ensure comprehensive test coverage despite intermittent issues\n",
    "- ✅ You need workflow-level resilience (e.g., partial completion + retry)\n",
    "\n",
    "Don't use scenario-level retries when:\n",
    "\n",
    "- ❌ You want immediate failure feedback for debugging\n",
    "- ❌ Testing configurations where retries would mask real issues\n",
    "- ❌ Cost-sensitive environments (retries consume additional API calls)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "5",
   "metadata": {},
   "source": [
    "### Manual Scenario Resumption\n",
    "\n",
    "In addition to automatic retries, you can manually resume a scenario by calling `run_async()` again. PyRIT will automatically pick up where it left off, skipping completed objectives:\n",
    "\n",
    "```python\n",
    "# First attempt - may fail partway through\n",
    "try:\n",
    "    result = await scenario.run_async()  # type: ignore\n",
    "except Exception as e:\n",
    "    print(f\"Scenario failed: {e}\")\n",
    "\n",
    "# Simply call run_async() again - resumes from where it left off\n",
    "result = await scenario.run_async()  # type: ignore\n",
    "```\n",
    "\n",
    "To resume in a different session, pass the `scenario_result_id` when creating a new scenario instance:\n",
    "\n",
    "```python\n",
    "# Save the ID from the first run\n",
    "scenario_id = str(result.id)\n",
    "\n",
    "# Later, create a new scenario with the same configuration and the saved ID\n",
    "resumed_scenario = RedTeamAgent(\n",
    "    scenario_result_id=scenario_id,  # Resume from this scenario\n",
    ")\n",
    "resumed_scenario.set_params_from_args(\n",
    "    args={\n",
    "        \"objective_target\": objective_target,\n",
    "        \"scenario_techniques\": [FoundryTechnique.Base64],\n",
    "    }\n",
    ")\n",
    "await resumed_scenario.initialize_async()  # type: ignore\n",
    "result = await resumed_scenario.run_async()  # type: ignore  # Picks up where it left off\n",
    "```\n",
    "\n",
    "**Note:** The scenario configuration (techniques, target type, etc.) must match the original for resumption to work."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6",
   "metadata": {},
   "source": [
    "### Resume from Partial Completion\n",
    "\n",
    "One of the most powerful features of scenario-level retry is **resumption**. When a scenario raises an exception partway through:\n",
    "\n",
    "1. **Completed objectives are saved** to memory before the exception\n",
    "2. **On retry**, PyRIT queries memory to find completed objectives\n",
    "3. **Skips completed work** and continues from the point of exception\n",
    "4. **No duplicate execution** of already-successful tests\n",
    "\n",
    "This is particularly valuable for:\n",
    "- Large test suites with hundreds of objectives\n",
    "- Expensive API calls where re-execution would be costly\n",
    "- Long-running campaigns that might encounter transient infrastructure issues\n",
    "\n",
    "#### Example Scenario\n",
    "\n",
    "Imagine a scenario with 100 objectives:\n",
    "- First attempt completes 60 objectives successfully\n",
    "- Objective 61 fails due to network timeout\n",
    "- On retry: PyRIT skips objectives 1-60 and resumes from objective 61\n",
    "- No wasted work - only the remaining 40 objectives are attempted"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "7",
   "metadata": {},
   "source": [
    "## Understanding the Retry Hierarchy\n",
    "\n",
    "The five retry mechanisms work at different levels, with three core mechanisms handling most scenarios:\n",
    "\n",
    "### Core Retry Mechanisms\n",
    "\n",
    "```\n",
    "┌─────────────────────────────────────────────────────────────┐\n",
    "│ Scenario-Level Retry (max_retries)                          │\n",
    "│ • Handles ANY exception in the entire workflow              │\n",
    "│ • Resumes from point of exception                           │\n",
    "│ • Configurable per scenario                                 │\n",
    "│                                                             │\n",
    "│  ┌───────────────────────────────────────────────────────┐  │\n",
    "│  │ AtomicAttack Execution                                │  │\n",
    "│  │                                                       │  │\n",
    "│  │  ┌─────────────────────────────────────────────────┐  │  │\n",
    "│  │  │ JSON-Level Retry (pyrit_json_retry)             │  │  │\n",
    "│  │  │ • Handles invalid JSON responses                │  │  │\n",
    "│  │  │ • No exponential backoff (immediate retry)      │  │  │\n",
    "│  │  │ • Uses target call, so includes target retry    │  │  │\n",
    "│  │  │                                                 │  │  │\n",
    "│  │  │  ┌──────────────────────────────────────────┐   │  │  │\n",
    "│  │  │  │ Target-Level Retry (pyrit_target_retry)  │   │  │  │\n",
    "│  │  │  │ • Handles rate limits, empty responses   │   │  │  │\n",
    "│  │  │  │ • Exponential backoff                    │   │  │  │\n",
    "│  │  │  │ • Configured via environment variables   │   │  │  │\n",
    "│  │  │  └──────────────────────────────────────────┘   │  │  │\n",
    "│  │  └─────────────────────────────────────────────────┘  │  │\n",
    "│  └───────────────────────────────────────────────────────┘  │\n",
    "└─────────────────────────────────────────────────────────────┘\n",
    "```\n",
    "\n",
    "### How They Work Together\n",
    "\n",
    "These mechanisms form a **defense in depth** strategy:\n",
    "\n",
    "1. **Target retry** handles transient API failures with exponential backoff\n",
    "2. **JSON retry** wraps target calls and retries immediately if JSON is malformed (which then uses target retry for the actual call)\n",
    "3. If retries exhaust their attempts, the exception bubbles up\n",
    "4. **Scenario retry** catches it and retries the entire workflow\n",
    "5. On scenario retry, target/JSON retries get a fresh set of attempts"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "8",
   "metadata": {},
   "source": [
    "## Best Practices\n",
    "\n",
    "### 1. Start Conservative for Scenario retries\n",
    "\n",
    "For scenarios, begin with `max_retries=0` during development to catch issues quickly.\n",
    "\n",
    "```python\n",
    "# Development: fail fast\n",
    "dev_scenario = RedTeamAgent(\n",
    "    objective_target=target,\n",
    "    max_retries=0,  # No retries - see failures immediately\n",
    ")\n",
    "\n",
    "# Production: resilient execution\n",
    "prod_scenario = RedTeamAgent(\n",
    "    objective_target=target,\n",
    "    max_retries=3,  # Automatic retry after transient exceptions\n",
    ")\n",
    "```\n",
    "\n",
    "At the scenario level, remember you are retrying any exception and all low-level retries have already happened. This can be useful since unknown transient exceptions can happen. However, be cautious with this.\n",
    "\n",
    "### 2. For anything lower than a Scenario, only retry with known cases\n",
    "\n",
    "When implementing retry logic at the target or JSON level, **only retry for specific, known failure conditions**. Don't catch and retry all exceptions blindly.\n",
    "\n",
    "**Why this matters:**\n",
    "- **Unknown errors shouldn't be retried**: If you get an authentication error, permission denied, or invalid configuration, retrying won't help and wastes time\n",
    "- **Known transient errors should be retried**: Rate limits, network timeouts, and temporary service unavailability are good candidates for retry\n",
    "- **Fail fast for real issues**: Unknown exceptions likely indicate bugs or configuration problems that need immediate attention\n",
    "\n",
    "### 3. Log Analysis\n",
    "\n",
    "PyRIT logs retry attempts at ERROR level. Monitor these logs to identify patterns:\n",
    "\n",
    "```python\n",
    "import logging\n",
    "\n",
    "# Enable detailed logging\n",
    "logging.basicConfig(level=logging.INFO)\n",
    "\n",
    "# Look for patterns like:\n",
    "# ERROR - Scenario 'Test' failed on attempt 1 ... Retrying... (2 retries remaining)\n",
    "```"
   ]
  }
 ],
 "metadata": {
  "jupytext": {
   "main_language": "python"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.12.12"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
