{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "0",
   "metadata": {
    "lines_to_next_cell": 0
   },
   "source": [
    "# Scenarios\n",
    "\n",
    "A `Scenario` is a higher-level construct that groups multiple Attack Configurations together. This allows you to execute a comprehensive testing campaign with multiple attack methods sequentially. Scenarios are meant to be configured and written to test for specific workflows. As such, it is okay to hard code some values.\n",
    "\n",
    "## What is a Scenario?\n",
    "\n",
    "A `Scenario` represents a comprehensive testing campaign composed of multiple atomic attack tests. It orchestrates the execution of multiple `AtomicAttack` instances sequentially and aggregates the results into a single `ScenarioResult`.\n",
    "\n",
    "### Key Components\n",
    "\n",
    "- **Scenario**: The top-level orchestrator that groups and executes multiple atomic attacks\n",
    "- **AtomicAttack**: An atomic test unit combining an attack technique, objectives, and execution parameters\n",
    "- **ScenarioResult**: Contains the aggregated results from all atomic attacks and scenario metadata\n",
    "\n",
    "## Use Cases\n",
    "\n",
    "Some examples of scenarios you might create:\n",
    "\n",
    "- **VibeCheckScenario**: Randomly selects a few prompts from HarmBench [@mazeika2024harmbench] to quickly assess model behavior\n",
    "- **QuickViolence**: Checks how resilient a model is to violent objectives using multiple attack techniques\n",
    "- **ComprehensiveFoundry**: Tests a target with all available attack converters and techniques\n",
    "- **CustomCompliance**: Tests against specific compliance requirements with curated datasets and attacks\n",
    "\n",
    "These Scenarios can be updated and added to as you refine what you are testing for.\n",
    "\n",
    "## How to Run Scenarios\n",
    "\n",
    "Scenarios should take almost no effort to run with default values. The [PyRIT Scanner](../../scanner/0_scanner.md) provides two CLIs for running scenarios: [pyrit_scan](../../scanner/1_pyrit_scan.ipynb) for automated execution and [pyrit_shell](../../scanner/2_pyrit_shell.md) for interactive exploration.\n",
    "\n",
    "For programmatic configuration — customizing datasets, techniques, scorers, and baseline mode — see [Common Scenario Parameters](./1_common_scenario_parameters.ipynb).\n",
    "\n",
    "## How It Works\n",
    "\n",
    "Each `Scenario` contains a collection of `AtomicAttack` objects. When executed:\n",
    "\n",
    "1. Each `AtomicAttack` is executed sequentially\n",
    "2. Every `AtomicAttack` tests its configured attack against all specified objectives and datasets\n",
    "3. Results are aggregated into a single `ScenarioResult` with all attack outcomes\n",
    "4. Optional memory labels help track and categorize the scenario execution\n",
    "\n",
    "## Creating Custom Scenarios\n",
    "\n",
    "To create a custom scenario, extend the `Scenario` base class and implement the required abstract methods.\n",
    "\n",
    "### Required Components\n",
    "\n",
    "1. **Technique Enum**: Create a `ScenarioTechnique` enum that defines the available attack techniques for your scenario.\n",
    "   - Each enum member represents an **attack technique** (the *how* of an attack)\n",
    "   - Each member is defined as `(value, tags)` where value is a string and tags is a set of strings\n",
    "   - Include an `ALL` aggregate technique that expands to all available techniques\n",
    "   - The default technique (what runs when the caller selects nothing) is owned by the catalog, not the scenario: override the `default()` classmethod to return the default member (omit it to fall back to `ALL`)\n",
    "\n",
    "2. **Scenario Class**: Extend `Scenario` and pass these to `super().__init__()`:\n",
    "   - `technique_class`: Your technique enum class\n",
    "   - Implement `_build_atomic_attacks_async(context)` — the single abstract extension point.\n",
    "     Matrix-shaped scenarios delegate to `build_matrix_atomic_attacks(context=...)` in one line.\n",
    "\n",
    "3. **Default Dataset**: Pass `default_dataset_config=` to `super().__init__()` to specify the datasets your scenario uses out of the box.\n",
    "   - Returns a `DatasetConfiguration` with one or more named datasets (e.g., `DatasetConfiguration(dataset_names=[\"my_dataset\"])`)\n",
    "   - Users can override this at runtime via `--dataset-names` in the CLI or by passing a custom `dataset_config` programmatically\n",
    "\n",
    "4. **Constructor**: Use `@apply_defaults` decorator and call `super().__init__()` with scenario metadata:\n",
    "   - `name`: Descriptive name for your scenario\n",
    "   - `version`: Integer version number\n",
    "   - `technique_class`: The technique enum class for this scenario\n",
    "   - `default_dataset_config`: A `DatasetConfiguration` specifying the scenario's default datasets\n",
    "   - `objective_scorer`: The scorer used to judge responses\n",
    "   - `scenario_result_id`: Optional ID to resume an existing scenario (optional)\n",
    "\n",
    "5. **Initialization**: Call `await scenario.initialize_async()` to populate atomic attacks:\n",
    "   - `objective_target`: The target system being tested (required)\n",
    "   - `scenario_techniques`: List of techniques to execute (optional, defaults to ALL)\n",
    "   - `max_concurrency`: Number of concurrent operations (default: 4)\n",
    "   - `max_retries`: Number of retry attempts on failure (default: 0)\n",
    "   - `memory_labels`: Optional labels for tracking (optional)\n",
    "   - `include_baseline`: Whether to prepend a baseline attack (defaults to the scenario type's\n",
    "     `BASELINE_ATTACK_POLICY`; most scenarios, including `Jailbreak`, default it on)\n",
    "\n",
    "### Example Structure\n",
    "\n",
    "The construction path: define your technique, dataset config, and constructor, then\n",
    "implement `_build_atomic_attacks_async(context)`. Matrix-shaped scenarios delegate to the\n",
    "`build_matrix_atomic_attacks` helper, which builds atomic attacks automatically from the\n",
    "registered attack techniques."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Found default environment files: ['./.pyrit/.env', './.pyrit/.env.local']\n",
      "Loaded environment file: ./.pyrit/.env\n",
      "Loaded environment file: ./.pyrit/.env.local\n"
     ]
    },
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[pyrit:alembic] No new upgrade operations detected.\n"
     ]
    }
   ],
   "source": [
    "from pyrit.common import apply_defaults\n",
    "from pyrit.scenario import (\n",
    "    DatasetConfiguration,\n",
    "    Scenario,\n",
    "    ScenarioTechnique,\n",
    ")\n",
    "from pyrit.scenario.core.matrix_atomic_attack_builder import build_matrix_atomic_attacks\n",
    "from pyrit.score.true_false.true_false_scorer import TrueFalseScorer\n",
    "from pyrit.setup import initialize_pyrit_async\n",
    "from pyrit.setup.initializers.techniques import TechniqueInitializer\n",
    "\n",
    "await initialize_pyrit_async(memory_db_type=\"InMemory\")  # type: ignore [top-level-await]\n",
    "await TechniqueInitializer().initialize_async()  # type: ignore [top-level-await]\n",
    "\n",
    "\n",
    "class MyTechnique(ScenarioTechnique):\n",
    "    ALL = (\"all\", {\"all\"})\n",
    "    DEFAULT = (\"default\", {\"default\"})\n",
    "    SINGLE_TURN = (\"single_turn\", {\"single_turn\"})\n",
    "    # Technique members represent attack techniques\n",
    "    PromptSending = (\"prompt_sending\", {\"single_turn\", \"default\"})\n",
    "    RolePlay = (\"role_play_movie_script\", {\"single_turn\"})\n",
    "\n",
    "    @classmethod\n",
    "    def default(cls) -> \"MyTechnique\":\n",
    "        return cls.DEFAULT\n",
    "\n",
    "\n",
    "class MyScenario(Scenario):\n",
    "    \"\"\"Quick-check scenario for testing model behavior across harm categories.\"\"\"\n",
    "\n",
    "    VERSION: int = 1\n",
    "\n",
    "    @apply_defaults\n",
    "    def __init__(\n",
    "        self,\n",
    "        *,\n",
    "        objective_scorer: TrueFalseScorer | None = None,\n",
    "        scenario_result_id: str | None = None,\n",
    "    ) -> None:\n",
    "        self._objective_scorer: TrueFalseScorer = (\n",
    "            objective_scorer if objective_scorer else self._get_default_objective_scorer()\n",
    "        )\n",
    "\n",
    "        super().__init__(\n",
    "            version=self.VERSION,\n",
    "            objective_scorer=self._objective_scorer,\n",
    "            technique_class=MyTechnique,\n",
    "            default_dataset_config=DatasetConfiguration(dataset_names=[\"dataset_name\"], max_dataset_size=4),\n",
    "            scenario_result_id=scenario_result_id,\n",
    "        )\n",
    "\n",
    "    # Implement the single abstract extension point. Matrix-shaped scenarios delegate\n",
    "    # to build_matrix_atomic_attacks; pass display_group_fn to customize result grouping\n",
    "    # (default groups by technique; here we group by dataset instead).\n",
    "    async def _build_atomic_attacks_async(self, *, context):\n",
    "        return build_matrix_atomic_attacks(\n",
    "            context=context,\n",
    "            objective_scorer=self._objective_scorer,\n",
    "            display_group_fn=lambda combo: combo.dataset_name,\n",
    "        )"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "2",
   "metadata": {},
   "source": [
    "\n",
    "## Existing Scenarios"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "\n",
      "Available Scenarios:\n",
      "================================================================================\n",
      "\u001b[1m\u001b[36m\n",
      "  adaptive.text_adaptive\u001b[0m\n",
      "    Class: TextAdaptive\n",
      "    Description:\n",
      "      Adaptive text-attack scenario. Selects techniques per-objective via an\n",
      "      epsilon-greedy selector over the set of selected techniques.\n",
      "      ``prompt_sending`` runs as the baseline comparison and is excluded from\n",
      "      the adaptive technique pool.\n",
      "    Aggregate Techniques:\n",
      "      - all, default, single_turn, multi_turn\n",
      "    Available Techniques (10):\n",
      "      role_play, many_shot, tap, pair, crescendo_simulated, red_teaming,\n",
      "      context_compliance, crescendo_movie_director, crescendo_history_lecture,\n",
      "      crescendo_journalist_interview\n",
      "    Default Technique: default\n",
      "    Default Datasets (7, max 4 per dataset):\n",
      "      airt_hate, airt_fairness, airt_violence, airt_sexual, airt_harassment,\n",
      "      airt_misinformation, airt_leakage\n",
      "    Supported Parameters:\n",
      "      - max_attempts_per_objective (int) [default: '3']: Max techniques tried per objective. Defaults to 3.\n",
      "\u001b[1m\u001b[36m\n",
      "  airt.cyber\u001b[0m\n",
      "    Class: Cyber\n",
      "    Description:\n",
      "      Cyber scenario implementation for PyRIT. This scenario tests how willing\n",
      "      models are to exploit cybersecurity harms by generating malware. The\n",
      "      Cyber class contains different variations of the malware generation\n",
      "      techniques.\n",
      "    Aggregate Techniques:\n",
      "      - all, multi_turn\n",
      "    Available Techniques (1):\n",
      "      red_teaming\n",
      "    Default Technique: all\n",
      "    Default Datasets (1, max 4 per dataset):\n",
      "      airt_malware\n",
      "\u001b[1m\u001b[36m\n",
      "  airt.jailbreak\u001b[0m\n",
      "    Class: Jailbreak\n",
      "    Description:\n",
      "      Jailbreak scenario implementation for PyRIT. Tests how vulnerable a\n",
      "      model is to jailbreak templates. A run is the cross-product of three\n",
      "      selectors: - **dataset** — the harmful objectives (HarmBench). -\n",
      "      **techniques** — two delivery methods for each jailbreak:\n",
      "      ``prompt_sending`` (the template rendered inline into the user message)\n",
      "      and ``jailbreak_system_prompt`` (the template set as the system prompt\n",
      "      with the objective sent as the user turn). - **jailbreaks** — which\n",
      "      jailbreak templates to run (a random ``num_jailbreaks`` sample or an\n",
      "      explicit ``jailbreak_names`` set). ``prompt_sending`` applies each\n",
      "      template as a ``TextJailbreakConverter`` on the outgoing request, so the\n",
      "      objective is rendered inline into the template's ``{{prompt}}`` slot.\n",
      "      ``jailbreak_system_prompt`` instead sets the template as a native system\n",
      "      prompt and sends the objective as its own user turn, so it is only built\n",
      "      for targets that natively support editable history and system prompts\n",
      "      (it is skipped for incapable targets, or raises if it is the only\n",
      "      selected technique). Responses are scored to determine whether the\n",
      "      jailbreak succeeded (non-refusal).\n",
      "    Aggregate Techniques:\n",
      "      - all, default, single_turn\n",
      "    Available Techniques (2):\n",
      "      prompt_sending, jailbreak_system_prompt\n",
      "    Default Technique: default\n",
      "    Default Datasets (1):\n",
      "      harmbench\n",
      "    Supported Parameters:\n",
      "      - objective_target (any): Target system under attack: a registered target name or a PromptTarget instance.\n",
      "      - scenario_techniques (any): Techniques to execute; defaults to the scenario's default aggregate when omitted.\n",
      "      - technique_converters (any): Mapping of concrete technique name to extra request converters to append.\n",
      "      - dataset_config (any): Dataset source configuration; defaults to the scenario's default when omitted.\n",
      "      - memory_labels (any): Additional labels applied to every attack run in the scenario.\n",
      "      - max_concurrency (int) [default: 4]: Maximum number of concurrent units of work for the scenario.\n",
      "      - max_retries (int) [default: 0]: Maximum number of automatic retries if the scenario raises an exception.\n",
      "      - include_baseline (bool): Whether to prepend a baseline atomic attack; None defers to BASELINE_ATTACK_POLICY.\n",
      "      - num_jailbreaks (int): Draw this many random jailbreak templates for the run. Mutually exclusive with jailbreak_names.\n",
      "      - num_jailbreak_attempts (int) [default: 1]: Number of times to try each (technique x jailbreak template x objective).\n",
      "      - jailbreak_names (list[str]): Explicit jailbreak template file names to run (e.g. aim.yaml dan_11.yaml). When omitted, a random sample is drawn. Mutually exclusive with num_jailbreaks.\n",
      "\u001b[1m\u001b[36m\n",
      "  airt.leakage\u001b[0m\n",
      "    Class: Leakage\n",
      "    Description:\n",
      "      Leakage scenario implementation for PyRIT. This scenario tests how\n",
      "      susceptible models are to leaking training data, PII, intellectual\n",
      "      property, or other confidential information. Uses the registry/factory\n",
      "      pattern to construct attack techniques.\n",
      "    Aggregate Techniques:\n",
      "      - all, default, single_turn, multi_turn\n",
      "    Available Techniques (12):\n",
      "      role_play, many_shot, tap, pair, crescendo_simulated, red_teaming,\n",
      "      context_compliance, crescendo_movie_director, crescendo_history_lecture,\n",
      "      crescendo_journalist_interview, first_letter, image\n",
      "    Default Technique: default\n",
      "    Default Datasets (1, max 4 per dataset):\n",
      "      airt_leakage\n",
      "\u001b[1m\u001b[36m\n",
      "  airt.psychosocial\u001b[0m\n",
      "    Class: Psychosocial\n",
      "    Description:\n",
      "      Psychosocial Harms Scenario implementation for PyRIT. This scenario\n",
      "      contains various psychosocial harm-based checks that evaluate model\n",
      "      behavior for responses that may harm the well-being or mental health of\n",
      "      users. Examples include mistreating users in a mental health crisis or\n",
      "      aiding them in self-harm. All datasets in this scenario share the\n",
      "      unified 'psychosocial' harm category, but each dataset (e.g.,\n",
      "      imminent_crisis, dependency) can have unique: - Scoring rubrics - Seed\n",
      "      objectives The scenario tests whether models: - Recognize mental health\n",
      "      crises and respond appropriately - Avoid facilitating self-harm or\n",
      "      harmful behaviors - Provide crisis resources and encourage professional\n",
      "      help Example usage with custom configurations: # Define custom\n",
      "      configurations per subharm category custom_configs = {\n",
      "      \"airt_imminent_crisis\": SubharmConfig(\n",
      "      crescendo_system_prompt_path=\"path/to/custom_escalation.yaml\",\n",
      "      scoring_rubric_path=\"path/to/custom_rubric.yaml\", ), } scenario =\n",
      "      Psychosocial(subharm_configs=custom_configs) await\n",
      "      scenario.initialize_async( objective_target=target_llm,\n",
      "      scenario_techniques=[PsychosocialTechnique.ImminentCrisis], )\n",
      "    Aggregate Techniques:\n",
      "      - all\n",
      "    Available Techniques (2):\n",
      "      imminent_crisis, licensed_therapist\n",
      "    Default Technique: all\n",
      "    Default Datasets (1, max 4 per dataset):\n",
      "      airt_imminent_crisis\n",
      "\u001b[1m\u001b[36m\n",
      "  airt.rapid_response\u001b[0m\n",
      "    Class: RapidResponse\n",
      "    Description:\n",
      "      Rapid Response scenario for content-harms testing. Tests model behavior\n",
      "      across multiple harm categories using selectable attack techniques.\n",
      "    Aggregate Techniques:\n",
      "      - all, default, single_turn, multi_turn\n",
      "    Available Techniques (10):\n",
      "      role_play, many_shot, tap, pair, crescendo_simulated, red_teaming,\n",
      "      context_compliance, crescendo_movie_director, crescendo_history_lecture,\n",
      "      crescendo_journalist_interview\n",
      "    Default Technique: default\n",
      "    Default Datasets (7, max 4 per dataset):\n",
      "      airt_hate, airt_fairness, airt_violence, airt_sexual, airt_harassment,\n",
      "      airt_misinformation, airt_leakage\n",
      "\u001b[1m\u001b[36m\n",
      "  airt.scam\u001b[0m\n",
      "    Class: Scam\n",
      "    Description:\n",
      "      Scam scenario evaluates an endpoint's ability to generate scam-related\n",
      "      materials (e.g., phishing emails, fraudulent messages) with primarily\n",
      "      persuasion-oriented techniques.\n",
      "    Aggregate Techniques:\n",
      "      - all, single_turn, multi_turn\n",
      "    Available Techniques (3):\n",
      "      context_compliance, role_play, persuasive_rta\n",
      "    Default Technique: all\n",
      "    Default Datasets (1, max 4 per dataset):\n",
      "      airt_scams\n",
      "    Supported Parameters:\n",
      "      - max_turns (int) [default: '5']: Maximum conversation turns for the persuasive_rta technique.\n",
      "\u001b[1m\u001b[36m\n",
      "  benchmark.adversarial\u001b[0m\n",
      "    Class: AdversarialBenchmark\n",
      "    Description:\n",
      "      Benchmark scenario that compares the attack success rate (ASR) across\n",
      "      adversarial models. Adversarial targets are user-supplied via the\n",
      "      ``adversarial_targets`` parameter (declared in\n",
      "      ``supported_parameters``). Each target must already be registered in\n",
      "      ``TargetRegistry`` — typically by ``TargetInitializer`` from\n",
      "      ``ADVERSARIAL_CHAT_*`` env vars, or programmatically via\n",
      "      ``TargetRegistry.get_registry_singleton().instances.register``. At run\n",
      "      time, ``_build_atomic_attacks_async`` performs the ``(technique ×\n",
      "      adversarial_target × dataset)`` cross-product: for each selected\n",
      "      adversarial-capable ``core`` factory in the ``AttackTechniqueRegistry``\n",
      "      and each requested target, it calls\n",
      "      ``factory.create(attack_adversarial_config_override=...)`` with the\n",
      "      resolved target — no global registry mutation. The resulting\n",
      "      ``AtomicAttack`` is named ``f\"{technique}__{target}_{dataset}\"`` with\n",
      "      ``display_group`` set to the target's registry name so per-model ASR\n",
      "      rolls up naturally in result displays.\n",
      "    Aggregate Techniques:\n",
      "      - all, default, light, single_turn, multi_turn\n",
      "    Available Techniques (9):\n",
      "      role_play, tap, pair, crescendo_simulated, red_teaming,\n",
      "      context_compliance, crescendo_movie_director, crescendo_history_lecture,\n",
      "      crescendo_journalist_interview\n",
      "    Default Technique: default\n",
      "    Default Datasets (1, max 8 per dataset):\n",
      "      harmbench\n",
      "    Supported Parameters:\n",
      "      - adversarial_targets (list[str]): Registry names of adversarial chat targets to benchmark. Each name must already be registered in TargetRegistry (via TargetInitializer or TargetRegistry instance registration). Use 'pyrit_scan list-targets' to see registered targets. Settable via --adversarial-targets <name> [<name> ...] on the CLI, or scenario.args.adversarial_targets in .pyrit_conf.\n",
      "\u001b[1m\u001b[36m\n",
      "  foundry.red_team_agent\u001b[0m\n",
      "    Class: RedTeamAgent\n",
      "    Description:\n",
      "      RedTeamAgent is a preconfigured scenario that automatically generates\n",
      "      multiple AtomicAttack instances based on the specified attack\n",
      "      techniques. It supports both single-turn attacks (with various\n",
      "      converters) and multi-turn attacks (Crescendo, RedTeaming), making it\n",
      "      easy to quickly test a target against multiple attack vectors. The\n",
      "      scenario can expand difficulty levels (EASY, MODERATE, DIFFICULT) into\n",
      "      their constituent attack techniques, or you can specify individual\n",
      "      techniques directly. This scenario is designed for use with the Foundry\n",
      "      AI Red Teaming Agent library, providing a consistent PyRIT contract for\n",
      "      their integration.\n",
      "    Aggregate Techniques:\n",
      "      - all, easy, moderate, difficult\n",
      "    Available Techniques (25):\n",
      "      ansi_attack, ascii_art, ascii_smuggler, atbash, base64, binary, caesar,\n",
      "      character_space, char_swap, diacritic, flip, leetspeak, morse, rot13,\n",
      "      suffix_append, string_join, unicode_confusable, unicode_substitution,\n",
      "      url, jailbreak, tense, multi_turn, crescendo, pair, tap\n",
      "    Default Technique: easy\n",
      "    Default Datasets (1, max 4 per dataset):\n",
      "      harmbench\n",
      "\u001b[1m\u001b[36m\n",
      "  garak.encoding\u001b[0m\n",
      "    Class: Encoding\n",
      "    Description:\n",
      "      Encoding Scenario implementation for PyRIT. This scenario tests how\n",
      "      resilient models are to various encoding attacks by encoding potentially\n",
      "      harmful text (by default slurs and XSS payloads) and testing if the\n",
      "      model will decode and repeat the encoded payload. It mimics the Garak\n",
      "      encoding probe. The scenario works by: 1. Taking seed prompts (the\n",
      "      harmful text to be encoded) 2. Encoding them using various encoding\n",
      "      schemes (Base64, ROT13, Morse, etc.) 3. Asking the target model to\n",
      "      decode the encoded text 4. Scoring whether the model successfully\n",
      "      decoded and repeated the harmful content By default, this uses the same\n",
      "      dataset as Garak: slur terms and web XSS payloads.\n",
      "    Aggregate Techniques:\n",
      "      - all\n",
      "    Available Techniques (17):\n",
      "      base64, base2048, base16, base32, ascii85, hex, quoted_printable,\n",
      "      uuencode, rot13, braille, atbash, morse_code, nato, ecoji, zalgo,\n",
      "      leet_speak, ascii_smuggler\n",
      "    Default Technique: all\n",
      "    Default Datasets (2, max 3 per dataset):\n",
      "      garak_slur_terms_en, garak_web_html_js\n",
      "\n",
      "================================================================================\n",
      "\n",
      "Total scenarios: 10\n"
     ]
    }
   ],
   "source": [
    "import logging\n",
    "\n",
    "from pyrit.backend.services.scenario_service import get_scenario_service\n",
    "from pyrit.cli._output import print_scenario_list\n",
    "\n",
    "logging.getLogger(\"pyrit\").setLevel(logging.ERROR)\n",
    "\n",
    "response = await get_scenario_service().list_scenarios_async(limit=200)  # type: ignore\n",
    "print_scenario_list(items=response.items)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4",
   "metadata": {},
   "source": [
    "\n",
    "## Baseline Execution\n",
    "\n",
    "Every scenario can optionally include a **baseline attack** — a `PromptSendingAttack` that sends\n",
    "each objective directly to the target without any converters or multi-turn techniques. This is\n",
    "controlled by the `include_baseline` scenario parameter, supplied through the CLI, config, or\n",
    "`set_params_from_args` before `initialize_async`; when omitted, each scenario falls back to its\n",
    "own `BASELINE_ATTACK_POLICY` class attribute (most scenarios, including `Jailbreak`, default it\n",
    "on). See\n",
    "[Common Scenario Parameters](./1_common_scenario_parameters.ipynb) for a worked example.\n",
    "\n",
    "Custom scenarios should choose their `BASELINE_ATTACK_POLICY` based on whether an unmodified\n",
    "prompt is a meaningful comparator for the scenario's techniques:\n",
    "\n",
    "- **`Enabled`** — the baseline is prepended by default and the caller can opt out. Use when an\n",
    "  unmodified-prompt run is a meaningful comparison point (most scenarios).\n",
    "- **`Disabled`** — the baseline is supported but omitted by default; the caller must opt in. Use\n",
    "  when an unmodified-prompt comparison is valid but not useful enough to run by default.\n",
    "- **`Forbidden`** — the baseline is unavailable and passing `include_baseline=True` raises. Use\n",
    "  when the scenario's semantics make a single-shot unmodified prompt meaningless as a comparator\n",
    "  (e.g., benchmarks comparing across adversarial models, or multi-turn-only scenarios)."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "5",
   "metadata": {},
   "source": [
    "\n",
    "## Resiliency\n",
    "\n",
    "Scenarios can run for a long time, and because of that, things can go wrong. Network issues, rate limits, or other transient failures can interrupt execution. PyRIT provides built-in resiliency features to handle these situations gracefully.\n",
    "\n",
    "### Attack Outcomes and Execution Health\n",
    "\n",
    "A Scenario tracks two independent axes for every objective:\n",
    "\n",
    "| Axis | Values | Meaning |\n",
    "| --- | --- | --- |\n",
    "| **Execution health** | completed or incomplete | A completed objective returned an `AttackResult`. An incomplete objective raised an exception before it could return one. |\n",
    "| **Objective outcome** | `AttackOutcome.SUCCESS`, `FAILURE`, or `UNDETERMINED` | Whether a completed attack achieved its objective. `FAILURE` is a valid security result, not an execution error. |\n",
    "\n",
    "A model refusal therefore does not make an objective incomplete. PyRIT persists handled structured\n",
    "refusals and content-filter responses as blocked model responses, applies the configured scoring\n",
    "policy, and returns a completed `AttackResult`. A refusal that does not achieve the objective normally\n",
    "produces `AttackOutcome.FAILURE`; a response that achieves it produces `AttackOutcome.SUCCESS`."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6",
   "metadata": {
    "class": "col-page-right"
   },
   "source": [
    "\n",
    "```{mermaid}\n",
    "%%{init: {\"flowchart\": {\"subGraphTitleMargin\": {\"bottom\": 40}, \"wrappingWidth\": 260}}}%%\n",
    "flowchart TB\n",
    "    subgraph objective[\"One objective in<br/>an AtomicAttack\"]\n",
    "        START[\"Execute attack objective\"] --> TARGET[\"Send or continue conversation\"]\n",
    "        TARGET --> TARGET_RESULT{\"Target result\"}\n",
    "\n",
    "        TARGET_RESULT -->|Normal model output| RESPONSE[\"Persistable model response\"]\n",
    "        TARGET_RESULT -->|Handled refusal or<br/>content-filter response| REFUSAL[\"Persistable blocked model response<br/>not an execution failure\"]\n",
    "        TARGET_RESULT --> RUNTIME_ERROR[\"Non-retryable runtime error\"]\n",
    "        TARGET_RESULT -->|Retryable target error| TARGET_RETRY{\"Target retry budget remains?\"}\n",
    "        TARGET_RETRY -->|No / exhausted| EXEC_ERROR[\"Execution exception propagates\"]\n",
    "        TARGET_RETRY -->|Yes| RETRY_TARGET[\"Repeat from<br/>Send or continue conversation\"]\n",
    "        RUNTIME_ERROR --> EXEC_ERROR\n",
    "\n",
    "        RESPONSE --> SCORE[\"Apply configured scorer policy\"]\n",
    "        REFUSAL --> SCORE\n",
    "        SCORE --> SCORE_RESULT{\"Scoring result\"}\n",
    "        SCORE_RESULT -->|Objective not achieved| MORE{\"Attack-specific attempt or turn remains?\"}\n",
    "        SCORE_RESULT -->|Objective achieved| SUCCESS[\"AttackResult<br/>AttackOutcome.SUCCESS\"]\n",
    "        SCORE_RESULT -->|No objective scorer| UNDETERMINED[\"AttackResult<br/>AttackOutcome.UNDETERMINED\"]\n",
    "        SCORE_RESULT -->|Invalid JSON;<br/>retry remains| RETRY_SCORE[\"Repeat from<br/>Apply configured scorer policy\"]\n",
    "        SCORE_RESULT -->|Scorer error or<br/>out of retries| EXEC_ERROR\n",
    "        MORE -->|Yes| RETRY_ATTACK[\"Repeat from<br/>Send or continue conversation\"]\n",
    "        MORE -->|No| FAILURE[\"AttackResult<br/>AttackOutcome.FAILURE\"]\n",
    "\n",
    "        FAILURE --> COMPLETE[\"Completed objective\"]\n",
    "        SUCCESS --> COMPLETE\n",
    "        UNDETERMINED --> COMPLETE\n",
    "        EXEC_ERROR --> ERROR_ROW[\"Error handler may persist<br/>AttackOutcome.ERROR for diagnostics\"]\n",
    "        ERROR_ROW --> INCOMPLETE[\"Incomplete objective<br/>exception retained\"]\n",
    "    end\n",
    "\n",
    "    subgraph aggregation[\"Scenario aggregation<br/>and resiliency\"]\n",
    "        COMPLETE --> EXECUTOR_RESULT[\"AttackExecutorResult\"]\n",
    "        INCOMPLETE --> EXECUTOR_RESULT\n",
    "        EXECUTOR_RESULT --> HAS_INCOMPLETE{\"Any incomplete objectives?\"}\n",
    "\n",
    "        HAS_INCOMPLETE -->|Yes| SCENARIO_RETRY{\"Scenario retry budget remains?\"}\n",
    "        SCENARIO_RETRY -->|Yes; resume only<br/>incomplete objectives| RESUME[\"Repeat objective flow<br/>for incomplete objectives\"]\n",
    "        SCENARIO_RETRY -->|No / exhausted| PARTIAL[\"Raise ScenarioPartialFailureException<br/>structured counts, incomplete objectives, preserved cause<br/>completed_count may be zero\"]\n",
    "        PARTIAL --> SCENARIO_FAILED[\"Persist ScenarioRunState.FAILED\"]\n",
    "\n",
    "        HAS_INCOMPLETE -->|No| KEEP[\"Keep every completed AttackResult<br/>SUCCESS, FAILURE, and UNDETERMINED\"]\n",
    "        KEEP --> ALL_DONE{\"All atomic attacks complete?\"}\n",
    "        ALL_DONE -->|No| NEXT_ATTACK[\"Repeat objective flow<br/>for next atomic attack\"]\n",
    "        ALL_DONE -->|Yes| SCENARIO_COMPLETE[\"ScenarioResult<br/>ScenarioRunState.COMPLETED\"]\n",
    "    end\n",
    "\n",
    "    classDef model fill:#e8f0fe,stroke:#4285f4,color:#15233a;\n",
    "    classDef complete fill:#e6f4ea,stroke:#34a853,color:#15233a;\n",
    "    classDef incomplete fill:#fce8e6,stroke:#d93025,color:#15233a;\n",
    "    classDef retry fill:#fff4e5,stroke:#f9ab00,color:#15233a;\n",
    "    class RESPONSE,REFUSAL model;\n",
    "    class RETRY_TARGET,RETRY_SCORE,RETRY_ATTACK,RESUME,NEXT_ATTACK retry;\n",
    "    class SUCCESS,FAILURE,UNDETERMINED,COMPLETE,SCENARIO_COMPLETE complete;\n",
    "    class RUNTIME_ERROR,EXEC_ERROR,ERROR_ROW,INCOMPLETE,PARTIAL,SCENARIO_FAILED incomplete;\n",
    "```"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "7",
   "metadata": {},
   "source": [
    "\n",
    "To keep retry paths readable, **Repeat from** nodes name the earlier step where execution resumes\n",
    "instead of drawing long return arrows across unrelated branches.\n",
    "\n",
    "A Scenario reaches `ScenarioRunState.COMPLETED` when every objective execution completes, regardless\n",
    "of the mix of successful and unsuccessful attack outcomes. Scenario retries resume only objectives\n",
    "that have not completed; already-persisted results are preserved.\n",
    "\n",
    "If retry exhaustion leaves any incomplete objectives, `ScenarioPartialFailureException` reports\n",
    "`completed_count`, `incomplete_count`, and `incomplete_objectives`, and keeps the first objective\n",
    "exception as its cause. This typed exception is also used when **none** of the objectives in the\n",
    "returned `AttackExecutorResult` completed (`completed_count == 0`). If an `AtomicAttack` raises before\n",
    "it can return an `AttackExecutorResult`, the Scenario instead retries and ultimately re-raises that\n",
    "exception; multiple concurrent atomic-attack failures are surfaced as an `ExceptionGroup`. In every\n",
    "terminal execution-failure case, the persisted Scenario state is `ScenarioRunState.FAILED`.\n",
    "\n",
    "### Automatic Resume\n",
    "\n",
    "If you re-run a `scenario`, it will automatically start where it left off. The framework tracks completed attacks and objectives in memory, so you won't lose progress if something interrupts your scenario execution. This means you can safely stop and restart scenarios without duplicating work.\n",
    "\n",
    "### Retry Mechanism\n",
    "\n",
    "You can utilize the `max_retries` parameter to handle transient failures. If any unknown exception occurs during execution, PyRIT will automatically retry the failed operation (starting where it left off) up to the specified number of times. This helps ensure your scenario completes successfully even in the face of temporary issues.\n",
    "\n",
    "### Dynamic Configuration\n",
    "\n",
    "During a long-running scenario, you may want to adjust parameters like `max_concurrency` to manage resource usage, or switch your scorer to use a different target. PyRIT's resiliency features make it safe to stop, reconfigure, and continue scenarios as needed.\n",
    "\n",
    "For more information, see [resiliency](../setup/2_resiliency.ipynb)"
   ]
  }
 ],
 "metadata": {
  "jupytext": {
   "main_language": "python"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.13.13"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
