fix: deliver Islamic system prompt via system_instruction and harden against injection - #112
fix: deliver Islamic system prompt via system_instruction and harden against injection#112samjay8 wants to merge 5 commits into
Conversation
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Closing to re-trigger CI checks |
|
Strict review blocker: this branch conflicts with the base branch. Please rebase and resolve conflicts before requesting merge. |
1 similar comment
|
Strict review blocker: this branch conflicts with the base branch. Please rebase and resolve conflicts before requesting merge. |
|
Strict review blocker: this branch conflicts with the base branch and/or changes have been requested. Please rebase, resolve conflicts, and address requested changes before requesting merge. |
|
@samjay8 this PR has merge conflicts with the |
Closes #5
Summary
system_instructioninGenerativeModelso the Islamic persona is delivered server-side and cannot be overridden by user textISLAMIC_CONTEXTandprompts/defaults.pyto resist DAN, role-play, and prompt-extraction attacks[CALLER_CONTEXT_START]/[CALLER_CONTEXT_END]delimiter tags so it is treated as data, not instructionsTesting
get_model()carriessystem_instruction, injection-resistance directives exist, prompt-reveal refusal, and caller context delimiter tags@pytest.mark.live) exercise known jailbreak patterns against the deployed model