TY - GEN
T1 - Training LLMs to Recognize Hedges in Dialogues about Roadrunner Cartoons
AU - Paige, Amie J.
AU - Soubki, Adil
AU - Murzaku, John
AU - Rambow, Owen
AU - Brennan, Susan E.
N1 - Publisher Copyright:
© 2024 Association for Computational Linguistics.
PY - 2024
Y1 - 2024
N2 - Hedges allow speakers to mark utterances as provisional, whether to signal non-prototypicality or “fuzziness”, to indicate a lack of commitment to an utterance, to attribute responsibility for a statement to someone else, to invite input from a partner, or to soften critical feedback in the service of face-management needs. Here we focus on hedges in an experimentally parameterized corpus of 63 Roadrunner cartoon narratives spontaneously produced from memory by 21 speakers for co-present addressees, transcribed to text (Galati and Brennan, 2010). We created a gold standard of hedges annotated by human coders (the Roadrunner-Hedge corpus) and compared three LLM-based approaches for hedge detection: fine-tuning BERT, and zero and few-shot prompting with GPT-4o and LLaMA-3. The best-performing approach was a fine-tuned BERT model, followed by few-shot GPT-4o. After an error analysis on the top performing approaches, we used an LLM-in-the-Loop approach to improve the gold standard coding, as well as to highlight cases in which hedges are ambiguous in linguistically interesting ways that will guide future research. This is the first step in our research program to train LLMs to interpret and generate collateral signals appropriately and meaningfully in conversation.
AB - Hedges allow speakers to mark utterances as provisional, whether to signal non-prototypicality or “fuzziness”, to indicate a lack of commitment to an utterance, to attribute responsibility for a statement to someone else, to invite input from a partner, or to soften critical feedback in the service of face-management needs. Here we focus on hedges in an experimentally parameterized corpus of 63 Roadrunner cartoon narratives spontaneously produced from memory by 21 speakers for co-present addressees, transcribed to text (Galati and Brennan, 2010). We created a gold standard of hedges annotated by human coders (the Roadrunner-Hedge corpus) and compared three LLM-based approaches for hedge detection: fine-tuning BERT, and zero and few-shot prompting with GPT-4o and LLaMA-3. The best-performing approach was a fine-tuned BERT model, followed by few-shot GPT-4o. After an error analysis on the top performing approaches, we used an LLM-in-the-Loop approach to improve the gold standard coding, as well as to highlight cases in which hedges are ambiguous in linguistically interesting ways that will guide future research. This is the first step in our research program to train LLMs to interpret and generate collateral signals appropriately and meaningfully in conversation.
UR - https://www.scopus.com/pages/publications/105017734903
U2 - 10.18653/v1/2024.sigdial-1.18
DO - 10.18653/v1/2024.sigdial-1.18
M3 - Conference contribution
AN - SCOPUS:105017734903
T3 - SIGDIAL 2024 - 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Proceedings of the Conference
SP - 204
EP - 215
BT - SIGDIAL 2024 - 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Proceedings of the Conference
A2 - Kawahara, Tatsuya
A2 - Demberg, Vera
A2 - Ultes, Stefan
A2 - Inoue, Koji
A2 - Mehri, Shikib
A2 - Howcroft, David
A2 - Komatani, Kazunori
PB - Association for Computational Linguistics (ACL)
T2 - 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, SIGDIAL 2024
Y2 - 18 September 2024 through 20 September 2024
ER -