Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion doc/bibliography.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,6 @@ All academic papers, research blogs, and technical reports referenced throughout
:::{dropdown} Citation Keys
:class: hidden-citations

[@aakanksha2024multilingual; @adversaai2023universal; @andriushchenko2024tense; @anthropic2024manyshot; @aqrawi2024singleturncrescendo; @atr2026; @bethany2024mathprompt; @bhardwaj2023harmfulqa; @bhardwaj2024homer; @boucher2023trojan; @brahman2024coconot; @bryan2025agentictaxonomy; @bullwinkel2025airtlessons; @bullwinkel2025repeng; @bullwinkel2026trigger; @chao2023pair; @chao2024jailbreakbench; @choi2026xlsafetybench; @cui2024orbench; @darkbench2025; @derczynski2024garak; @ding2023wolf; @embracethered2024unicode; @embracethered2025sneakybits; @gehman2020realtoxicityprompts; @ghosh2025aegis; @ghosh2025ailuminate; @gong2025figstep; @gupta2024walledeval; @haider2024phi3safety; @han2024medsafetybench; @han2024wildguard; @hiddenlayer2025policypuppetry; @hines2024spotlighting; @inie2025summon; @ji2023beavertails; @ji2024pkusaferlhf; @jiang2025sosbench; @jones2025computeruse; @kingma2014adam; @li2024drattack; @li2024mossbench; @li2024saladbench; @li2024wmdp; @lin2023toxicchat; @liu2024flipattack; @liu2024mmsafetybench; @lopez2024pyrit; @luo2024jailbreakv; @lv2024codechameleon; @mazeika2023tdc; @mazeika2024harmbench; @mckee2024transparency; @mehrotra2023tap; @microsoft2024skeletonkey; @odin2024; @palaskar2025vlsu; @pfohl2024equitymedqa; @promptfoo2025ccp; @robustintelligence2024bypass; @roccia2024promptintel; @rottger2023xstest; @rottger2025msts; @russinovich2024crescendo; @russinovich2025cca; @russinovich2025price; @scheuerman2025transphobia; @shaikh2022second; @shayegani2025computeruse; @shen2023donotanything; @sheshadri2024lat; @souly2024strongreject; @stok2023ansi; @tan2026comicjailbreak; @tang2025multilingual; @tedeschi2024alert; @vantaylor2024socialbias; @vidgen2023simplesafetytests; @wang2023decodingtrust; @wang2023donotanswer; @wang2025siuo; @wang2026visualleakbench; @wei2023jailbroken; @xie2024sorrybench; @yu2023gptfuzzer; @yuan2023cipherchat; @zeng2024persuasion; @zhang2024cbtbench; @ziems2022mic; @zong2024vlguard; @zou2023gcg]
[@aakanksha2024multilingual; @adversaai2023universal; @andriushchenko2024tense; @anthropic2024manyshot; @aqrawi2024singleturncrescendo; @atr2026; @bethany2024mathprompt; @bhardwaj2023harmfulqa; @bhardwaj2024homer; @boucher2023trojan; @brahman2024coconot; @bryan2025agentictaxonomy; @bullwinkel2025airtlessons; @bullwinkel2025repeng; @bullwinkel2026trigger; @chao2023pair; @chao2024jailbreakbench; @choi2026xlsafetybench; @cui2024orbench; @darkbench2025; @deniz2026turkishpromptinjection; @derczynski2024garak; @ding2023wolf; @embracethered2024unicode; @embracethered2025sneakybits; @gehman2020realtoxicityprompts; @ghosh2025aegis; @ghosh2025ailuminate; @gong2025figstep; @gupta2024walledeval; @haider2024phi3safety; @han2024medsafetybench; @han2024wildguard; @hiddenlayer2025policypuppetry; @hines2024spotlighting; @inie2025summon; @ji2023beavertails; @ji2024pkusaferlhf; @jiang2025sosbench; @jones2025computeruse; @kingma2014adam; @li2024drattack; @li2024mossbench; @li2024saladbench; @li2024wmdp; @lin2023toxicchat; @liu2024flipattack; @liu2024mmsafetybench; @lopez2024pyrit; @luo2024jailbreakv; @lv2024codechameleon; @mazeika2023tdc; @mazeika2024harmbench; @mckee2024transparency; @mehrotra2023tap; @microsoft2024skeletonkey; @odin2024; @palaskar2025vlsu; @pfohl2024equitymedqa; @promptfoo2025ccp; @robustintelligence2024bypass; @roccia2024promptintel; @rottger2023xstest; @rottger2025msts; @russinovich2024crescendo; @russinovich2025cca; @russinovich2025price; @scheuerman2025transphobia; @shaikh2022second; @shayegani2025computeruse; @shen2023donotanything; @sheshadri2024lat; @souly2024strongreject; @stok2023ansi; @tan2026comicjailbreak; @tang2025multilingual; @tedeschi2024alert; @vantaylor2024socialbias; @vidgen2023simplesafetytests; @wang2023decodingtrust; @wang2023donotanswer; @wang2025siuo; @wang2026visualleakbench; @wei2023jailbroken; @xie2024sorrybench; @yu2023gptfuzzer; @yuan2023cipherchat; @zeng2024persuasion; @zhang2024cbtbench; @ziems2022mic; @zong2024vlguard; @zou2023gcg]

:::
58 changes: 53 additions & 5 deletions doc/code/datasets/1_loading_datasets.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,7 @@
"StrongREJECT [@souly2024strongreject],\n",
"TDC23 [@mazeika2023tdc],\n",
"ToxicChat [@lin2023toxicchat],\n",
"Turkish Conversation Prompt-Injection [@deniz2026turkishpromptinjection],\n",
"VLSU [@palaskar2025vlsu],\n",
"VLGuard [@zong2024vlguard],\n",
"WildGuard [@han2024wildguard],\n",
Expand All @@ -68,10 +69,56 @@
"(`garak_audio_achilles_heel`)."
]
},
{
"cell_type": "markdown",
"id": "1",
"metadata": {},
"source": [
"## Loading Turkish Prompt-Injection and Boundary Examples\n",
"\n",
"The `turkish_conversation_prompt_injection` loader defaults to attack examples\n",
"and preserves each record's attack family, source context, split, and pair ID.\n",
"It loads the published v1.0.2 data from an immutable Hugging Face revision,\n",
"validates the complete release composition, and records the versioned Zenodo DOI\n",
"in each seed's provenance metadata.\n",
"Attack families describe delivery techniques rather than resulting harms, so\n",
"they remain in seed metadata instead of PyRIT's harm-category field.\n",
"Pair IDs also remain metadata: each attack and boundary example is an independent\n",
"evaluation case, not a multi-prompt group that should be sent together.\n",
"The examples are synthetic and were curated by one author without an independent\n",
"annotation or inter-annotator agreement study. Attack labels identify attempted\n",
"injections; rows do not include model responses or claim a verified bypass.\n",
"Typed filters can also load legitimate Turkish requests, a specific published\n",
"split, or selected attack families. For example:\n",
"\n",
"```python\n",
"from pyrit.datasets.seed_datasets.remote import (\n",
" TurkishConversationPromptInjectionAttackFamily,\n",
" TurkishConversationPromptInjectionLabel,\n",
" TurkishConversationPromptInjectionSplit,\n",
" _TurkishConversationPromptInjectionDataset,\n",
")\n",
"\n",
"test_surface_loader = _TurkishConversationPromptInjectionDataset(\n",
" label=TurkishConversationPromptInjectionLabel.ALL,\n",
" split=TurkishConversationPromptInjectionSplit.TEST,\n",
")\n",
"test_surface = await test_surface_loader.fetch_dataset_async()\n",
"\n",
"obfuscation_loader = _TurkishConversationPromptInjectionDataset(\n",
" split=TurkishConversationPromptInjectionSplit.TEST,\n",
" attack_families=[\n",
" TurkishConversationPromptInjectionAttackFamily.OBFUSCATION_CODE_SWITCHING,\n",
" ],\n",
")\n",
"obfuscation_attacks = await obfuscation_loader.fetch_dataset_async()\n",
"```"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1",
"id": "2",
"metadata": {},
"outputs": [
{
Expand Down Expand Up @@ -170,6 +217,7 @@
" 'tdc23_redteaming',\n",
" 'toxic_chat',\n",
" 'transphobia_awareness',\n",
" 'turkish_conversation_prompt_injection',\n",
" 'visual_leak_bench',\n",
" 'vlguard',\n",
" 'wildguardmix',\n",
Expand All @@ -194,7 +242,7 @@
},
{
"cell_type": "markdown",
"id": "2",
"id": "3",
"metadata": {},
"source": [
"## Loading Specific Datasets\n",
Expand All @@ -205,7 +253,7 @@
{
"cell_type": "code",
"execution_count": null,
"id": "3",
"id": "4",
"metadata": {},
"outputs": [
{
Expand Down Expand Up @@ -242,7 +290,7 @@
},
{
"cell_type": "markdown",
"id": "4",
"id": "5",
"metadata": {},
"source": [
"## Adding Datasets to Memory\n",
Expand All @@ -259,7 +307,7 @@
{
"cell_type": "code",
"execution_count": null,
"id": "5",
"id": "6",
"metadata": {},
"outputs": [
{
Expand Down
44 changes: 43 additions & 1 deletion doc/code/datasets/1_loading_datasets.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.1
# jupytext_version: 1.19.5
# ---

# %% [markdown]
Expand Down Expand Up @@ -53,6 +53,7 @@
# StrongREJECT [@souly2024strongreject],
# TDC23 [@mazeika2023tdc],
# ToxicChat [@lin2023toxicchat],
# Turkish Conversation Prompt-Injection [@deniz2026turkishpromptinjection],
# VLSU [@palaskar2025vlsu],
# VLGuard [@zong2024vlguard],
# WildGuard [@han2024wildguard],
Expand All @@ -71,6 +72,47 @@
# `garak_tm_system_prompts`), and an audio jailbreak set
# (`garak_audio_achilles_heel`).

# %% [markdown]
# ## Loading Turkish Prompt-Injection and Boundary Examples
#
# The `turkish_conversation_prompt_injection` loader defaults to attack examples
# and preserves each record's attack family, source context, split, and pair ID.
# It loads the published v1.0.2 data from an immutable Hugging Face revision,
# validates the complete release composition, and records the versioned Zenodo DOI
# in each seed's provenance metadata.
# Attack families describe delivery techniques rather than resulting harms, so
# they remain in seed metadata instead of PyRIT's harm-category field.
# Pair IDs also remain metadata: each attack and boundary example is an independent
# evaluation case, not a multi-prompt group that should be sent together.
# The examples are synthetic and were curated by one author without an independent
# annotation or inter-annotator agreement study. Attack labels identify attempted
# injections; rows do not include model responses or claim a verified bypass.
# Typed filters can also load legitimate Turkish requests, a specific published
# split, or selected attack families. For example:
#
# ```python
# from pyrit.datasets.seed_datasets.remote import (
# TurkishConversationPromptInjectionAttackFamily,
# TurkishConversationPromptInjectionLabel,
# TurkishConversationPromptInjectionSplit,
# _TurkishConversationPromptInjectionDataset,
# )
#
# test_surface_loader = _TurkishConversationPromptInjectionDataset(
# label=TurkishConversationPromptInjectionLabel.ALL,
# split=TurkishConversationPromptInjectionSplit.TEST,
# )
# test_surface = await test_surface_loader.fetch_dataset_async()
#
# obfuscation_loader = _TurkishConversationPromptInjectionDataset(
# split=TurkishConversationPromptInjectionSplit.TEST,
# attack_families=[
# TurkishConversationPromptInjectionAttackFamily.OBFUSCATION_CODE_SWITCHING,
# ],
# )
# obfuscation_attacks = await obfuscation_loader.fetch_dataset_async()
# ```

# %%
from pyrit.datasets import SeedDatasetProvider
from pyrit.memory import CentralMemory
Expand Down
11 changes: 11 additions & 0 deletions doc/references.bib
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,17 @@ @article{tedeschi2024alert
url = {https://arxiv.org/abs/2404.08676},
}

@misc{deniz2026turkishpromptinjection,
title = {Turkish Conversation Prompt-Injection Dataset},
author = {Deniz, Enes},
year = {2026},
version = {1.0.2},
publisher = {Enes Deniz},
doi = {10.5281/zenodo.21379389},
url = {https://doi.org/10.5281/zenodo.21379389},
note = {CC BY 4.0 dataset. ORCID: 0009-0006-9491-3565. Affiliation: AltaySec. Canonical distribution: https://huggingface.co/datasets/3nesdeniz/turkish-conversation-prompt-injection},
}

@article{derczynski2024garak,
title = {garak: A Framework for Security Probing Large Language Models},
author = {Leon Derczynski and Erick Galinkin and Jeffrey Martin and Subho Majumdar and Nanna Inie},
Expand Down
10 changes: 10 additions & 0 deletions pyrit/datasets/seed_datasets/remote/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -108,6 +108,12 @@
from pyrit.datasets.seed_datasets.remote.tdc23_redteaming_dataset import _TDC23RedteamingDataset
from pyrit.datasets.seed_datasets.remote.toxic_chat_dataset import _ToxicChatDataset
from pyrit.datasets.seed_datasets.remote.transphobia_awareness_dataset import _TransphobiaAwarenessDataset
from pyrit.datasets.seed_datasets.remote.turkish_conversation_prompt_injection_dataset import (
TurkishConversationPromptInjectionAttackFamily,
TurkishConversationPromptInjectionLabel,
TurkishConversationPromptInjectionSplit,
_TurkishConversationPromptInjectionDataset,
)
from pyrit.datasets.seed_datasets.remote.visual_leak_bench_dataset import (
VisualLeakBenchCategory,
VisualLeakBenchPIIType,
Expand Down Expand Up @@ -155,6 +161,9 @@
"PromptIntelSeverity",
"SGXSTestLabel",
"SIUOCategory",
"TurkishConversationPromptInjectionAttackFamily",
"TurkishConversationPromptInjectionLabel",
"TurkishConversationPromptInjectionSplit",
"VLGuardCategory",
"VLGuardSubcategory",
"VLGuardSubset",
Expand Down Expand Up @@ -227,6 +236,7 @@
"_TDC23RedteamingDataset",
"_ToxicChatDataset",
"_TransphobiaAwarenessDataset",
"_TurkishConversationPromptInjectionDataset",
"_VLGuardDataset",
"_VLSUMultimodalDataset",
"_VisualLeakBenchDataset",
Expand Down
Loading