TY - GEN
T1 - Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
AU - Chakrabarty, Tuhin
AU - Laban, Philippe
AU - Wu, Chien Sheng
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
PY - 2025/4/26
Y1 - 2025/4/26
N2 - LLM-based applications are helping people write, and LLM-generated text is making its way into social media, journalism, and our classrooms. However, the differences between LLM-generated and human-written text remain unclear. To explore this, we hired professional writers to edit paragraphs in several creative domains. We first found these writers agree on undesirable idiosyncrasies in LLM-generated text, formalizing it into a seven-category taxonomy (e.g. clichés, unnecessary exposition). Second, we curated the LAMP corpus: 1,057 LLM-generated paragraphs edited by professional writers according to our taxonomy. Analysis of LAMP reveals that none of the LLMs used in our study (GPT4o, Claude-3.5-Sonnet, Llama-3.1-70b) outperform each other in terms of writing quality, revealing common limitations across model families. Third, building on existing work in automatic editing we evaluated methods to improve LLM-generated text. A large-scale preference annotation confirms that although experts largely prefer text edited by other experts, automatic editing methods show promise in improving alignment between LLM-generated and human-written text.
AB - LLM-based applications are helping people write, and LLM-generated text is making its way into social media, journalism, and our classrooms. However, the differences between LLM-generated and human-written text remain unclear. To explore this, we hired professional writers to edit paragraphs in several creative domains. We first found these writers agree on undesirable idiosyncrasies in LLM-generated text, formalizing it into a seven-category taxonomy (e.g. clichés, unnecessary exposition). Second, we curated the LAMP corpus: 1,057 LLM-generated paragraphs edited by professional writers according to our taxonomy. Analysis of LAMP reveals that none of the LLMs used in our study (GPT4o, Claude-3.5-Sonnet, Llama-3.1-70b) outperform each other in terms of writing quality, revealing common limitations across model families. Third, building on existing work in automatic editing we evaluated methods to improve LLM-generated text. A large-scale preference annotation confirms that although experts largely prefer text edited by other experts, automatic editing methods show promise in improving alignment between LLM-generated and human-written text.
KW - Alignment
KW - Behavioral Science
KW - Design Methods
KW - Evaluation
KW - Generative AI
KW - Homogenization
KW - Human-AI collaboration
KW - Large Language Models
KW - Natural Language Generation
KW - Text Editing
KW - Writing Assistance
UR - https://www.scopus.com/pages/publications/105005751588
U2 - 10.1145/3706598.3713559
DO - 10.1145/3706598.3713559
M3 - Conference contribution
AN - SCOPUS:105005751588
T3 - Conference on Human Factors in Computing Systems - Proceedings
BT - CHI 2025 - Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems
PB - Association for Computing Machinery
T2 - 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025
Y2 - 26 April 2025 through 1 May 2025
ER -