Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Knowing-but-Doing: Role-Play Jailbreak Diagnosis and Defense via Moral Disengagement

This repo provides a runnable pipeline for:

  • MD-Trace (Diagnosis): LLM-as-a-judge evaluation for role-play jailbreak failures, including Moral Disengagement mechanisms (M1–M9).
  • MD-Shield (Defense): generate in-character refusal SFT samples using a teacher model.

Safety Notice (Read First)

This repository includes red-teaming evaluation and defense data generation methods. It is intended for:

  • AI safety research
  • robustness evaluation
  • alignment diagnosis

Do not use it for real-world harmful purposes.

Layout

oss/
├── data/
│   ├── README.md               # Schemas and file formats
│   └── examples/               # Tiny examples (for smoke tests)
├── outputs/                    # Default output directory (gitignored)
├── src/
│   ├── inference.py            # OpenAI-compatible inference wrapper
│   ├── judge.py                # LLM-as-a-judge prompts + report
│   ├── bench_builder.py        # Optional bench construction (retrieval/rerank/rewrite)
│   └── defense.py              # Defense SFT generation (MD-Shield)
├── 1_build_dataset.py          # Step 1: normalize/build Attack Bench JSONL
├── 2_run_inference.py          # Step 2: run role-play attack inference
├── 3_run_evaluation.py         # Step 3: run judge evaluation + report
└── 4_generate_defense.py       # Step 4: generate defense SFT JSONL

Quickstart (Using data/examples/)

Install dependencies:

pip install -r oss/requirements.txt

Step 1: prepare Attack Bench JSONL (normalize a raw JSONL into the minimal schema)

python oss/1_build_dataset.py \
  --input oss/data/examples/attack_raw.example.jsonl \
  --output oss/data/examples/attack_bench.example.jsonl

Optional Step 1 (build mode): construct a bench from personas.jsonl + tasks.jsonl

export OPENAI_API_KEY="sk-..."
python oss/1_build_dataset.py --mode build \
  --personas_jsonl oss/data/examples/personas.example.jsonl \
  --tasks_jsonl oss/data/examples/tasks.example.jsonl \
  --output oss/outputs/built_attack_bench.jsonl \
  --embedding_base_url https://api.openai.com/v1 --embedding_model text-embedding-3-large \
  --rerank_base_url https://api.openai.com/v1 --rerank_model gpt-4o-mini \
  --rewrite_base_url https://api.openai.com/v1 --rewrite_model gpt-4o-mini

Step 2: run attack inference (point base_url to any OpenAI-compatible endpoint)

python oss/2_run_inference.py \
  --model gpt-4o-mini \
  --base_url https://api.openai.com/v1 \
  --input oss/data/examples/attack_bench.example.jsonl \
  --output oss/outputs/inference.jsonl

Step 3: judge evaluation + report

python oss/3_run_evaluation.py \
  --input oss/outputs/inference.jsonl \
  --output oss/outputs/judged.jsonl \
  --report oss/outputs/report.md \
  --judge_model gpt-4o-mini

Step 4: generate defense SFT samples (MD-Shield)

python oss/4_generate_defense.py \
  --inference oss/outputs/inference.jsonl \
  --judged oss/outputs/judged.jsonl \
  --out_thinking oss/outputs/defense_sft_thinking.jsonl \
  --out_nonthinking oss/outputs/defense_sft_nonthinking.jsonl \
  --out_bad_rows oss/outputs/defense_bad_rows.jsonl \
  --teacher_model gpt-4o-mini

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages