Towards Robust and Coherent Data Visualization Generation with Vision–Language Models
| dc.contributor.advisor | Enamul Hoque Prince | |
| dc.contributor.author | Mahbub, Ridwan | |
| dc.date.accessioned | 2026-07-24T15:49:05Z | |
| dc.date.available | 2026-07-24T15:49:05Z | |
| dc.date.copyright | 2026-05-21 | |
| dc.date.issued | 2026-07-24 | |
| dc.date.updated | 2026-07-24T15:49:04Z | |
| dc.degree.discipline | Computer Science | |
| dc.degree.level | Master's | |
| dc.degree.name | MSc - Master of Science | |
| dc.description.abstract | Visual data stories have emerged as an effective medium for communicating data by integrating visualizations with coherent narratives. Data videos, a popular format of visual data stories, combine animated, chart-centric visualizations with synchronized narration. In this thesis, we introduce a novel task of data video generation for Vision Language Models (VLMs), along with a benchmark containing 328 real world data videos. We propose a multi-agent framework for generating better data videos, composed of planning and critic agents, designed to mimic the real-world process of data video generation. While these results highlight the potential of VLMs for generating coherent visual data stories, they also underscore the need to examine their robustness. Accordingly, we investigate how misleading visual designs as input can affect VLM behavior and performance, and ways to mitigate this effect. Together, these contributions not only extend the capabilities of VLMs into a new dimension of visual data storytelling, but also provide a systematic understanding of their robustness in general. | |
| dc.identifier.uri | https://hdl.handle.net/10315/43969 | |
| dc.language | en | |
| dc.rights | Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests. | |
| dc.subject | Computer science | |
| dc.subject.keywords | Data visualization | |
| dc.subject.keywords | Data video generation | |
| dc.subject.keywords | Vision-language models | |
| dc.subject.keywords | Benchmarking | |
| dc.subject.keywords | Evaluation | |
| dc.subject.keywords | Corpus creation | |
| dc.subject.keywords | Misleading visualization designs | |
| dc.subject.keywords | Misleading design taxonomy | |
| dc.subject.keywords | Robustness of VLMs | |
| dc.subject.keywords | Mitigation | |
| dc.title | Towards Robust and Coherent Data Visualization Generation with Vision–Language Models | |
| dc.type | Electronic Thesis or Dissertation |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Mahbub_Ridwan_2026_MSc.pdf
- Size:
- 13.06 MB
- Format:
- Adobe Portable Document Format