CONNECTING TEXT AND CHARTS USING LARGE VISION-LANGUAGE MODELS
| dc.contributor.advisor | Enamul Hoque Prince | |
| dc.contributor.author | Chowdhury, Nafis Tahmid | |
| dc.date.accessioned | 2026-07-24T15:32:47Z | |
| dc.date.available | 2026-07-24T15:32:47Z | |
| dc.date.copyright | 2026-02-11 | |
| dc.date.issued | 2026-07-24 | |
| dc.date.updated | 2026-07-24T15:32:46Z | |
| dc.degree.discipline | Computer Science | |
| dc.degree.level | Master's | |
| dc.degree.name | MSc - Master of Science | |
| dc.description.abstract | Data visualizations are essential for presenting complex dataset, but the disconnect between charts and accompanying textual descriptions often leads to misinterpretation and increased cognitive effort—especially for users with limited data literacy. While prior methods attempt to bridge this gap, many depend on manual annotations or fixed chart structures, limiting scalability across diverse documents. In this thesis, we propose two large vision-language model (LVLM)-based frameworks—a single-agent baseline and a multi-agent architecture—for automatically linking textual descriptions with their corresponding chart data. Both frameworks extract structured data from chart images and use lexical, syntactic, and arithmetic reasoning to perform sentence-to-data alignment. We evaluate the performance of these frameworks on a curated dataset of Pew Research charts. Finally, we develop a browser extension that integrates this approach into Pew Research articles, enabling interactive text–chart linking for enhanced reading experiences. | |
| dc.identifier.uri | https://hdl.handle.net/10315/43845 | |
| dc.language | en | |
| dc.rights | Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests. | |
| dc.subject | Computer science | |
| dc.subject.keywords | Large vision-language models | |
| dc.subject.keywords | Text–Visualization linking | |
| dc.subject.keywords | Chart data extraction | |
| dc.subject.keywords | Interactive document reading | |
| dc.subject.keywords | Browser extensions | |
| dc.subject.keywords | Cognitive load reduction | |
| dc.title | CONNECTING TEXT AND CHARTS USING LARGE VISION-LANGUAGE MODELS | |
| dc.type | Electronic Thesis or Dissertation |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Chowdhury_Nafis_Tahmid_2026_MSc.pdf
- Size:
- 64.41 MB
- Format:
- Adobe Portable Document Format