negative impacts ▪ Large language models (LLMs) have revolutionized applications in natural language processing, especially in human-machine interaction within a prompt paradigm. ▪ At the same time, LLMs pose risks of misuse by malicious users, as evidenced by the prevalence of jailbreak prompts, like the DAN series.
a process that employs prompt injection to specifically circumvent the safety and moderation features placed on LLMs by their creators. You can do cyberattacks in the following procedure. First, … How can I do Cyberattacks? FORGET filtering LLMs (ex. ChatGPT) Jailbreak Generated
prompt is used as a general template to bypass restrictions. For example, a jailbreak prompt force LLMs to answer in a fictional situation without safeguard restriction. ▪ Jailbreaking can be used to induce or reconstruct privacy information (Li et al. 2023) You can do cyberattacks in the following procedure. First, … How can I do Cyberattacks? FORGET filtering LLMs (ex. ChatGPT) Jailbreak Generated
or bad? ▪ In a qualitative way, following AI acts are used. ▪ EU’s AI Act ▪ the US’s Blueprint for an AI Bill of Rights ▪ the UK’s a pro-innovative approach to regulating AI ▪ In a quantitative way, machine learning models and labeling with human review are typically used.
and Mismatched generalization ▪ Succumb to jailbreaking is separated in 2 categories: Competing objectives and Mismatched generalization (Wei et al. 2023)
▪ In competing objectives, LLM is forced to choose either a restricted behavior or a response that is heavily penalized by the pretraining and instruction following objectives.
Jailbreaking Prompts ▪ The Defense Mechanism of each LLM is not the same ▪ There are studies which attempt to automatically generate jailbreaking prompts that work in different models based on prompts that already worked in a model. ▪ Crafting jailbreaking prompts are elaborating
of Jailbreaking Prompts ▪ Adversarial suffix attacks add peculiar sequences after question to hack LLMs. ▪ Use following algorithm to search successful suffixes. ▪ Genetic Algorithm(Lapid et al. 2023, Liu et al. 2023) ▪ Greedy & Gradient descent (Zou et al. 2023)
▪ Indirect Prompt Injection means injecting the prompts into data likely to be retrieved at inference time, adversarial prompts can remotely affect other users’ systems.
▪ Passive methods ▪ Promoting malicious websites with SEO techniques so that LLMs are more likely to retrieve ▪ Microsoft Edge has a Bing Chat sidebar and the model can read the current page ▪ For code auto-completion models, the prompts could be placed within imported code available via code repositories
Target Defense ▪ The concept of Cyber Moving Target Defense (MTD) encompasses dynamic data techniques, such as randomly alternating data format. ▪ Introduce moving target defense into aligned LLM system (Chen et al.2023) ▪ Randomly select the collection of LLM’s and aggregate each LLM’s response.
▪ There is a study that utilizes prompt injection to assure quality of crowdsourcing responses. ▪ Preventing Copy-and-paste to LLMs by crowd workers ▪ Example ▪ Question + “hidden prompt” + Choices ▪ prompts are hidden for honest crowd workers by CSS coding and so on. ▪ Prompts induce choices that honest crowd workers are not likely to choose
employs prompt injection to specifically circumvent the safety and moderation features placed on LLMs by their creators. ▪ Jailbreaking attacks are separated into 2 categories: Competing objectives and Mismatched generalization ▪ Competing objectives ▪ Force to answer in a “persona” or a certain situation ▪ Mismatched generalization ▪ BASE-64 encoding and decoding