Prince, Enamul HoqueMahbub, Ridwan2026-07-242026-07-242026-05-212026-07-24https://hdl.handle.net/10315/43969Visual data stories have emerged as an effective medium for communicating data by integrating visualizations with coherent narratives. Data videos, a popular format of visual data stories, combine animated, chart-centric visualizations with synchronized narration. In this thesis, we introduce a novel task of data video generation for Vision Language Models (VLMs), along with a benchmark containing 328 real world data videos. We propose a multi-agent framework for generating better data videos, composed of planning and critic agents, designed to mimic the real-world process of data video generation. While these results highlight the potential of VLMs for generating coherent visual data stories, they also underscore the need to examine their robustness. Accordingly, we investigate how misleading visual designs as input can affect VLM behavior and performance, and ways to mitigate this effect. Together, these contributions not only extend the capabilities of VLMs into a new dimension of visual data storytelling, but also provide a systematic understanding of their robustness in general.Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests.Computer scienceTowards Robust and Coherent Data Visualization Generation with Vision–Language ModelsElectronic Thesis or Dissertation2026-07-24Data visualizationData video generationVision-language modelsBenchmarkingEvaluationCorpus creationMisleading visualization designsMisleading design taxonomyRobustness of VLMsMitigation