Shvartzshnaider, YanLacalamita, John Pietro2026-07-242026-07-242026-03-312026-07-24https://hdl.handle.net/10315/43882Large language models (LLMs) are increasingly becoming ubiquitous in day-to-day tasks. Yet, despite the growing dependence on LLM-based systems, their sensitivity to surface-level linguistic variation has received little attention. We introduce PARSE (Prompt Alteration Response-Shift Evaluation), a modular framework that generates grammatical, typographical, and dialectal prompt variants, queries LLMs under identical conditions, and measures distributional output shifts. We apply PARSE to two case studies, film recommendation and privacy bias evaluation, across three models (GPT-4o-mini, Llama~3.2, DeepSeek-7B). Results show that output shifts scale with perturbation intensity: grammatical rewrites produce minimal effects, while typographical noise and dialect rewrites significantly alter recommendations and appropriateness ratings. Perturbations push film recommendations toward higher-rated, generic titles, and shift privacy ratings toward more restrictive values. These effects are directionally consistent across models, demonstrating that the linguistic form of a prompt systematically biases LLM outputs.Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests.Computer scienceArtificial intelligenceLinguisticsPARSE: A Framework for Evaluating Linguistic and Typographical Perturbation Bias in Large Language ModelsElectronic Thesis or Dissertation2026-07-24Large language modelsPrompt sensitivityLinguistic robustnessDialectal variationTypographical perturbationsContextual integrityRecommender systemsPrivacy biasNatural language processing evaluationModel evaluation