TY - GEN
T1 - Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs
AU - Li, Yanhong
AU - Lan, Zixuan
AU - Zhou, Jiawei
N1 - Publisher Copyright:
©2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Large language models (LLMs) and their multimodal variants can now process visual inputs, including images of text. This raises an intriguing question: can we compress textual inputs by feeding them as images to reduce token usage while preserving performance? In this paper, we show that visual text representations are a practical and surprisingly effective form of input compression for decoder LLMs. We exploit the idea of rendering long text inputs as a single image and provide it directly to the model. This leads to dramatically reduced number of decoder tokens required, offering a new form of input compression. Through experiments on two distinct benchmarks—RULER (long-context retrieval) and CNN/DailyMail (document summarization)—we demonstrate that this text-as-image method yields substantial token savings (often nearly half) without degrading task performance.
AB - Large language models (LLMs) and their multimodal variants can now process visual inputs, including images of text. This raises an intriguing question: can we compress textual inputs by feeding them as images to reduce token usage while preserving performance? In this paper, we show that visual text representations are a practical and surprisingly effective form of input compression for decoder LLMs. We exploit the idea of rendering long text inputs as a single image and provide it directly to the model. This leads to dramatically reduced number of decoder tokens required, offering a new form of input compression. Through experiments on two distinct benchmarks—RULER (long-context retrieval) and CNN/DailyMail (document summarization)—we demonstrate that this text-as-image method yields substantial token savings (often nearly half) without degrading task performance.
UR - https://www.scopus.com/pages/publications/105028964642
U2 - 10.18653/v1/2025.findings-emnlp.558
DO - 10.18653/v1/2025.findings-emnlp.558
M3 - Conference contribution
AN - SCOPUS:105028964642
T3 - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025
SP - 10564
EP - 10578
BT - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025
A2 - Christodoulopoulos, Christos
A2 - Chakraborty, Tanmoy
A2 - Rose, Carolyn
A2 - Peng, Violet
PB - Association for Computational Linguistics (ACL)
T2 - 30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Y2 - 4 November 2025 through 9 November 2025
ER -