tuning them is hard and also costly • Fix their parameters and apply them to different tasks by different prompts • Hard prompt, prompt tuning[2], transfer learning[3] … [1] Brown, Tom, et al. "Language models are few-shot learners.“, Neurips 2020 [2] Lester, Brian, Rami Al-Rfou, and Noah Constant. "The power of scale for parameter-efficient prompt tuning “, (2021). [3] Asai, Akari, et al. "Attempt: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts.“ EMNLP 2022. Background Information
natural language specifications of a task • Manually rewrite task-inputs to the prescribed formats on a example-by-example basis[1] • Simplify complex tasks to achieve better performance in the prompting paradigm[2] [1] Mishra, Swaroop, et al. "Reframing Instructional Prompts to GPTk's Language.“ 2021. [2] Creswell, Antonia, Murray Shanahan, and Irina Higgins. "Selection-inference: Exploiting large language models for interpretable logical reasoning." 2022. Background Information
of tasks and finds the process to be brittle • Small changes to the prompt result in large performance variations[1],[2] • Significant effort is dedicated towards designing a perfect prompt for a task [1] Zhao, Zihao, et al. "Calibrate before use: Improving few-shot performance of language models." PMLR, 2021. [2] Holtzman, Ari, et al. "Surface form competition: Why the highest probability answer isn't always right." 2021. Motivation
yet imperfect prompts to improve prompting performance • Vote for the input’s true label to produce a final prediction Motivation -> Ask Me Anything : A Simple strategy for Prompting Language Models
to improvements from aggregation • Previous approach[1] focus on a single task -> focus on prompt engineering • Origin standard prompt format(hard)[2] is right? • (“John invited Mark to come watch Jurassic Park. Output True or False?”) - restrict • (“John invited Mark to come watch Jurassic _” fill-the-blank, “Park”) - cloze question • (“Where did John invite Mark?”) - open ended question [1] Wei, Jason, et al. "Chain of thought prompting elicits reasoning in large language models." 2022. [2] Brown, Tom, et al. "Language models are few-shot learners.“, Neurips 2020 Model Structure
GPT-J 6B) • For example, in WSC, • Restrictive form : “The pronoun ‘his’ refers to “Mark” in the context. True or False?” • Open-ended form : “Mark went to the park with his dog.”. Reformatting to “What does ‘his’ refer to?” • 38 % lift (50% -> 69.2 %) Model Structure
effective? • Intuitively, the task of answering open-ended questions is aligned with the next-token prediction language modeling objective • By analysis EleutherAI[1](Pile corpus[2]), • Open-ended QA structures is 1000* more frequently than the restrictive format • Large imbalances in corpus between the frequencies [1] https://www.eleuther.ai/ [2] Black, Sid, et al. "Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow” 2021. Model Structure
input to new format [1],[2] • Given input x, applying prompt-chains 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎(𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 𝑥𝑥 ) • 𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 : 𝑥𝑥 → 𝑞𝑞 − 𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔 𝑎𝑎 𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 𝑎𝑎𝑎𝑎 𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖 𝑥𝑥 − (1) • 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 : 𝑞𝑞 → 𝑎𝑎 − 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑡𝑡𝑡𝑡𝑡 𝑞𝑞 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 1 𝑡𝑡𝑡𝑡 𝑡𝑡𝑡𝑡𝑡 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 𝑜𝑜𝑜𝑜 𝑥𝑥 𝑡𝑡𝑡𝑡 𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑎𝑎 • 𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 and 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 also contains demonstration of prompts • Chains are (1) reused across inputs and (2) different pairs of functional prompts can be combined to create variety Model Structure [1] Mishra, Swaroop, et al. "Reframing Instructional Prompts to GPTk's Language." 2021 [2] Wu, Tongshuang, Michael Terry, and Carrie Jun Cai. "Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts.“ CHI `22. 2022.
input to new format [1],[2] • Given input x, applying prompt-chains 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎(𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 𝑥𝑥 ) • 𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 : 𝑥𝑥 → 𝑞𝑞 − 𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔 𝑎𝑎 𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 𝑎𝑎𝑎𝑎 𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖 𝑥𝑥 − (1) • 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 : 𝑞𝑞 → 𝑎𝑎 − 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑡𝑡𝑡𝑡𝑡 𝑞𝑞 𝑓𝑓𝑓𝑓𝑓𝑓𝑓𝑓 1 𝑡𝑡𝑡𝑡 𝑡𝑡𝑡𝑡𝑡 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 𝑜𝑜𝑜𝑜 𝑥𝑥 𝑡𝑡𝑡𝑡 𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑎𝑎 • 𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞𝑞 and 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 also contains demonstration of prompts • Chains are (1) reused across inputs and (2) different pairs of functional prompts can be combined to create variety Model Structure [1] Mishra, Swaroop, et al. "Reframing Instructional Prompts to GPTk's Language." 2021 [2] Wu, Tongshuang, Michael Terry, and Carrie Jun Cai. "Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts.“ CHI `22. 2022.
style of prompt question • Each unique prompt()-chain is a different view of the task • Each unique prompt()-chain emphasize different aspects of x Model Structure (with our running example: “Who went to the park?”, “Did John go the park?”, “Where did John go?”)
strategy in prior prompting work[1][2] • not enough for prompt dependency and varied accuracy • WS is a powerful framework that learns the accuracies and correlations for training data[3] • Smith[4] applied WS aggregating the outputs of hand-curated prompts into a labeled dataset • Prompt()-chains get varied accuracies and dependencies (Appendix A) -> Weak supervision Model Structure [1] Jiang, Zhengbao, et al. "How can we know what language models know?." TACL 2020. [2] Schick, Timo, and Hinrich Schütze. "It's not just size that matters: Small language models are also few-shot learners.“ 2020. [3] Ratner, Alexander, et al. "Snorkel: Rapid training data creation with weak supervision." 2017. [4] Smith, Ryan, et al. "Language models in the loop: Incorporating prompting into weak supervision” 2022.
measures the amount of uncertainty remaining in the true label 𝑦𝑦 given a prediction � 𝑦𝑦 • In our setting, � 𝑦𝑦 = ∅(𝑃𝑃 𝑥𝑥 ) is dependent on the two components, P and ∅ • The first term shows 𝐻𝐻 𝑦𝑦 � 𝑦𝑦 depends on the quality and quantity of the individual prompts in P(x) • The second term shows 𝐻𝐻 𝑦𝑦 � 𝑦𝑦 depends on how the aggregation step compresses the information Information Flow