TY - GEN
T1 - Harnessing Language Models to Analyze Android App Permission Fidelity
AU - Tamrakar, Yunik
AU - Banerjee, Ritwik
AU - Myers, Ethan
AU - Carli, Lorenzo De
AU - Ray, Indrakshi
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Android's vast app ecosystem (over 2 million apps) poses significant privacy risks, as current methods for inferring permissions from descriptions - keyword matching, traditional natural language processing (NLP), and recurrent neural networks (RNNs) - struggle with accurate inference due to imprecise, ambiguous, or incomplete natural language descriptions. This gap undermines regulatory transparency and user trust, necessitating tools that reconcile stated functionality with actual data practices. We demonstrate that large language models like GPT-4o, applied in a zero-shot inference setting, leverage contextual reasoning to infer permissions competitively, while fine-tuned encoders (BERT, BART) surpass state-of-the-art performance when trained on minimally annotated datasets augmented with paraphrases, achieving 50-70% gains in weighted and macro F1 scores. By enabling precise permission auditing with reduced annotation costs, our work advances scalable, adaptable solutions for privacy compliance across resource-constrained and highstakes environments.
AB - Android's vast app ecosystem (over 2 million apps) poses significant privacy risks, as current methods for inferring permissions from descriptions - keyword matching, traditional natural language processing (NLP), and recurrent neural networks (RNNs) - struggle with accurate inference due to imprecise, ambiguous, or incomplete natural language descriptions. This gap undermines regulatory transparency and user trust, necessitating tools that reconcile stated functionality with actual data practices. We demonstrate that large language models like GPT-4o, applied in a zero-shot inference setting, leverage contextual reasoning to infer permissions competitively, while fine-tuned encoders (BERT, BART) surpass state-of-the-art performance when trained on minimally annotated datasets augmented with paraphrases, achieving 50-70% gains in weighted and macro F1 scores. By enabling precise permission auditing with reduced annotation costs, our work advances scalable, adaptable solutions for privacy compliance across resource-constrained and highstakes environments.
KW - Mobile applications
KW - classifier design and evaluation
KW - language models
KW - machine learning
KW - natural language processing
KW - privacy
KW - regulation
UR - https://www.scopus.com/pages/publications/105030486505
U2 - 10.1109/PST65910.2025.11268849
DO - 10.1109/PST65910.2025.11268849
M3 - Conference contribution
AN - SCOPUS:105030486505
T3 - 2025 22nd Annual International Conference on Privacy, Security, and Trust, PST 2025
BT - 2025 22nd Annual International Conference on Privacy, Security, and Trust, PST 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 22nd Annual International Conference on Privacy, Security, and Trust, PST 2025
Y2 - 26 August 2025 through 28 August 2025
ER -