Towards Robust and Coherent Data Visualization Generation with Vision–Language Models

dc.contributor.advisorEnamul Hoque Prince
dc.contributor.authorMahbub, Ridwan
dc.date.accessioned2026-07-24T15:49:05Z
dc.date.available2026-07-24T15:49:05Z
dc.date.copyright2026-05-21
dc.date.issued2026-07-24
dc.date.updated2026-07-24T15:49:04Z
dc.degree.disciplineComputer Science
dc.degree.levelMaster's
dc.degree.nameMSc - Master of Science
dc.description.abstractVisual data stories have emerged as an effective medium for communicating data by integrating visualizations with coherent narratives. Data videos, a popular format of visual data stories, combine animated, chart-centric visualizations with synchronized narration. In this thesis, we introduce a novel task of data video generation for Vision Language Models (VLMs), along with a benchmark containing 328 real world data videos. We propose a multi-agent framework for generating better data videos, composed of planning and critic agents, designed to mimic the real-world process of data video generation. While these results highlight the potential of VLMs for generating coherent visual data stories, they also underscore the need to examine their robustness. Accordingly, we investigate how misleading visual designs as input can affect VLM behavior and performance, and ways to mitigate this effect. Together, these contributions not only extend the capabilities of VLMs into a new dimension of visual data storytelling, but also provide a systematic understanding of their robustness in general.
dc.identifier.urihttps://hdl.handle.net/10315/43969
dc.languageen
dc.rightsAuthor owns copyright, except where explicitly noted. Please contact the author directly with licensing requests.
dc.subjectComputer science
dc.subject.keywordsData visualization
dc.subject.keywordsData video generation
dc.subject.keywordsVision-language models
dc.subject.keywordsBenchmarking
dc.subject.keywordsEvaluation
dc.subject.keywordsCorpus creation
dc.subject.keywordsMisleading visualization designs
dc.subject.keywordsMisleading design taxonomy
dc.subject.keywordsRobustness of VLMs
dc.subject.keywordsMitigation
dc.titleTowards Robust and Coherent Data Visualization Generation with Vision–Language Models
dc.typeElectronic Thesis or Dissertation

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Mahbub_Ridwan_2026_MSc.pdf
Size:
13.06 MB
Format:
Adobe Portable Document Format